EDBT 2026 Demo / reviewers in the wild / expert
Tomoharu Iwata
dblp:29/5953
· DBLP profile ↗
46ranked-venue papers in the field
15as first author
6since 2021 · last 2025
0000-0003-4425-1971ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 29 (9 first)Information Retrieval & Web Search · 10 (2 first)Database Systems & Data Management · 5 (4 first)Other / Interdisciplinary · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast Proximal Gradient Methods with Node Pruning for Tree-Structured Sparse Regularization
Yasutoshi Ida, Sekitoshi Kanai, Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara |
ECML/PKDD (5) | 4 |
| 2025 | Meta-learning of Class Knowledge in Zero-shot LearningabstractZero-shot learning is a promising approach to generalizing a model to categories unseen during training, and various methods have been proposed. However, they assume that class knowledge, the semantic information of the classes, is available as prior knowledge and thus fails to support domains whose class knowledge is unavailable. We propose a meta-learning method that allows us to apply the zero-shot learning method even if class knowledge is unavailable. We assume multiple zero-shot learning tasks where some classes are missing for each task, but they appear in other tasks. Our method simultaneously estimates the appropriate class knowledge and classification model using a meta-learning approach that extracts task features. We can use the trained models to perform zero-shot classification on unseen tasks without class knowledge. In experiments on datasets where true class knowledge is available, class knowledge is unavailable, and class knowledge is provided but imprecise, we show that the proposed method performs better than existing zero-shot learning methods. Yuta Nambu, Masahiro Kohjima, Tomoharu Iwata, Ryuji Yamamoto |
SDM | 3 |
| 2023 | Personal History Affects Reference Points: A Case Study of CodeforcesabstractHumans make decisions based on their internal value function, and its shape is known to be distorted and biased around a point, which the research community of behavior economics refers to as the reference point. People intensify activities that come to lie within the reach of their reference point, and abstain from acts that would incur losses once they've crossed the point. However, the impact of past experiences on decision making around the reference point has not been well studied. By analyzing a long series of user-level decisions gathered from a competitive programming website, we find that history has a clear impact on user's decision making around the reference point. Past experiences can strengthen, and sometimes weaken, the decision bias around the reference point. Experiences of past difficulties can strengthen the tendency towards loss aversion after achieving the reference point. When a person crosses a reference point for the first time, the cognitive decision bias is significant. However, repeating this crossing gradually weakens the effect. We also show the value of our insights in the task of predicting user behavior. Prediction models incorporating our insights may be used for motivating people to remain more active. Takeshi Kurashima, Tomoharu Iwata, Tomu Tominaga, Shuhei Yamamoto, Hiroyuki Toda, Kazuhisa Takemura |
ICWSM | 2 |
| 2022 | Predicting Opinion Dynamics via Sociologically-Informed Neural NetworksabstractOpinion formation and propagation are crucial phenomena in social networks and have been extensively studied across several disciplines. Traditionally, theoretical models of opinion dynamics have been proposed to describe the interactions between individuals (i.e., social interaction) and their impact on the evolution of collective opinions. Although these models can incorporate sociological and psychological knowledge on the mechanisms of social interaction, they demand extensive calibration with real data to make reliable predictions, requiring much time and effort. Recently, the widespread use of social media platforms provides new paradigms to learn deep learning models from a large volume of social media data. However, these methods ignore any scientific knowledge about the mechanism of social interaction. In this work, we present the first hybrid method called Sociologically-Informed Neural Network (SINN), which integrates theoretical models and social media data by transporting the concepts of physics-informed neural networks (PINNs) from natural science (i.e., physics) into social science (i.e., sociology and social psychology). In particular, we recast theoretical models as ordinary differential equations (ODEs). Then we train a neural network that simultaneously approximates the data and conforms to the ODEs that represent the social scientific knowledge. In addition, we extend PINNs by integrating matrix factorization and a language model to incorporate rich side information (e.g., user profiles) and structural knowledge (e.g., cluster structure of the social interaction network). Moreover, we develop an end-to-end training procedure for SINN, which involves Gumbel-Softmax approximation to include stochastic mechanisms of social interaction. Extensive experiments on real-world and synthetic datasets show SINN outperforms six baseline methods in predicting opinion dynamics. Maya Okawa, Tomoharu Iwata |
KDD | 2 |
| 2022 | Learning Optimal Priors for Task-Invariant Representations in Variational AutoencodersabstractThe variational autoencoder (VAE) is a powerful latent variable model for unsupervised representation learning. However, it does not work well in case of insufficient data points. To improve the performance in such situations, the conditional VAE (CVAE) is widely used, which aims to share task-invariant knowledge with multiple tasks through the task-invariant latent variable. In the CVAE, the posterior of the latent variable given the data point and task is regularized by the task-invariant prior, which is modeled by the standard Gaussian distribution. Although this regularization encourages independence between the latent variable and task, the latent variable remains dependent on the task. To reduce this task-dependency, the previous work introduced an additional regularizer. However, its learned representation does not work well on the target tasks. In this study, we theoretically investigate why the CVAE cannot sufficiently reduce the task-dependency and show that the simple standard Gaussian prior is one of the causes. Based on this, we propose a theoretical optimal prior for reducing the task-dependency. In addition, we theoretically show that unlike the previous work, our learned representation works well on the target tasks. Experiments on various datasets show that our approach obtains better task-invariant representations, which improves the performances of various downstream applications such as density estimation and classification. Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai, Sekitoshi Kanai, Masanori Yamada, Yuki Yamanaka, Hisashi Kashima |
KDD | 2 |
| 2021 | Dynamic Hawkes Processes for Discovering Time-evolving Communities' States behind Diffusion ProcessesabstractSequences of events including infectious disease outbreaks, social network activities, and crimes are ubiquitous and the data on such events carry essential information about the underlying diffusion processes between communities (e.g., regions, online user groups). Modeling diffusion processes and predicting future events are crucial in many applications including epidemic control, viral marketing, and predictive policing. Hawkes processes offer a central tool for modeling the diffusion processes, in which the influence from the past events is described by the triggering kernel. However, the triggering kernel parameters, which govern how each community is influenced by the past events, are assumed to be static over time. In the real world, the diffusion processes depend not only on the influences from the past, but also the current (time-evolving) states of the communities, e.g., people's awareness of the disease and people's current interests. In this paper, we propose a novel Hawkes process model that is able to capture the underlying dynamics of community states behind the diffusion processes and predict the occurrences of events based on the dynamics. Specifically, we model the latent dynamic function that encodes these hidden dynamics by a mixture of neural networks. Then we design the triggering kernel using the latent dynamic function and its integral. The proposed method, termed DHP (Dynamic Hawkes Processes), offers a flexible way to learn complex representations of the time-evolving communities' states, while at the same time it allows to computing the exact likelihood, which makes parameter learning tractable. Extensive experiments on four real-world event datasets show that DHP outperforms five widely adopted methods for event prediction. Maya Okawa, Tomoharu Iwata, Yusuke Tanaka 0002, Hiroyuki Toda, Takeshi Kurashima, Hisashi Kashima |
KDD | 2 |
| 2020 | Transfer Metric Learning for Unseen DomainsabstractAbstract We propose a transfer metric learning method to infer domain-specific data embeddings for unseen domains, from which no data are given in the training phase, by using knowledge transferred from related domains. When training and test distributions are different, the standard metric learning cannot infer appropriate data embeddings. The proposed method can infer appropriate data embeddings for the unseen domains by using latent domain vectors, which are latent representations of domains and control the property of data embeddings for each domain. This latent domain vector is inferred by using a neural network that takes the set of feature vectors in the domain as an input. The neural network is trained without the unseen domains. The proposed method can instantly infer data embeddings for the unseen domains without (re)-training once the sets of feature vectors in the domains are given. To accumulate knowledge in advance, the proposed method uses labeled and unlabeled data in multiple source domains. Labeled data, i.e., data with label information such as class labels or pair (similar/dissimilar) constraints, are used for learning data embeddings in such a way that similar data points are close and dissimilar data points are separated in the embedding space. Although unlabeled data do not have labels, they have geometric information that characterizes domains. The proposed method incorporates this information in a natural way on the basis of a probabilistic framework. The conditional distributions of the latent domain vectors, the embedded data, and the observed data are parameterized by neural networks and are optimized by maximizing the variational lower bound using stochastic gradient descent. The effectiveness of the proposed method was demonstrated through experiments using three clustering tasks. Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara |
Data Sci. Eng. | 2 |
| 2019 | Transfer Metric Learning for Unseen DomainsabstractWe propose a transfer metric learning method to infer domain-specific data embeddings for unseen domains, from which no data are given in the training phase, by using knowledge transferred from related domains. When training and test distributions are different, the standard metric learning cannot infer appropriate data embeddings. The proposed method can infer appropriate data embeddings for the unseen domains by using latent domain vectors, which are latent representations of domains and control the property of data embeddings for each domain. This latent domain vector is inferred by using a neural network that takes the set of feature vectors in the domain as an input. The neural network is trained without the unseen domains. The proposed method can instantly infer data embeddings for the unseen domains without (re)-training once the sets of feature vectors in the domains are given. To accumulate knowledge in advance, the proposed method uses labeled and unlabeled data in multiple source domains. Labeled data, i.e., data with label information such as class labels or pair (similar/dissimilar) constraints, are used for learning data embeddings in such a way that similar data points are close and dissimilar data points are separated in the embedding space. Although unlabeled data do not have labels, they have geometric information that characterizes domains. The proposed method incorporates this information in a natural way on the basis of a probabilistic framework. The conditional distributions of the latent domain vectors, the embedded data, and the observed data are parametrized by neural networks and are optimized by maximizing the variational lower bound. The effectiveness of the proposed method was demonstrated through experiments using three clustering tasks. Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara |
ICDM | 2 |
| 2019 | Deep Mixture Point Processes: Spatio-temporal Event Prediction with Rich Contextual InformationabstractPredicting when and where events will occur in cities, like taxi pick-ups, crimes, and vehicle collisions, is a challenging and important problem with many applications in fields such as urban planning, transportation optimization and location-based marketing. Though many point processes have been proposed to model events in a continuous spatio-temporal space, none of them allow for the consideration of the rich contextual factors that affect event occurrence, such as weather, social activities, geographical characteristics, and traffic. In this paper, we propose DMPP (Deep Mixture Point Processes), a point process model for predicting spatio-temporal events with the use of rich contextual information; a key advance is its incorporation of the heterogeneous and high-dimensional context available in image and text data. Specifically, we design the intensity of our point process model as a mixture of kernels, where the mixture weights are modeled by a deep neural network. This formulation allows us to automatically learn the complex nonlinear effects of the contextual factors on event occurrence. At the same time, this formulation makes analytical integration over the intensity, which is required for point process estimation, tractable. We use real-world data sets from different domains to demonstrate that DMPP has better predictive performance than existing methods. Maya Okawa, Tomoharu Iwata, Takeshi Kurashima, Yusuke Tanaka 0002, Hiroyuki Toda, Naonori Ueda |
KDD | 2 |
| 2018 | Learning Dynamics of Decision Boundaries without Additional Labeled DataabstractWe propose a method for learning the dynamics of the decision boundary to maintain classification performance without additional labeled data. In various applications, such as spam-mail classification, the decision boundary dynamically changes over time. Accordingly, the performance of classifiers deteriorates quickly unless the classifiers are retrained using additional labeled data. However, continuously preparing such data is quite expensive or impossible. The proposed method alleviates this deterioration in performance by using newly obtained unlabeled data, which are easy to prepare, as well as labeled data collected beforehand. With the proposed method, the dynamics of the decision boundary is modeled by Gaussian processes. To exploit information on the decision boundaries from unlabeled data, the low-density separation criterion, i.e., the decision boundary should not cross high-density regions, but instead lie in low-density regions, is assumed with the proposed method. We incorporate this criterion into our framework in a principled manner by introducing the entropy posterior regularization to the posterior of the classifier parameters on the basis of the generic regularized Bayesian framework. We developed an efficient inference algorithm for the model based on variational Bayesian inference. The effectiveness of the proposed method was demonstrated through experiments using two synthetic and four real-world data sets. Atsutoshi Kumagai, Tomoharu Iwata |
KDD | 2 |
| 2018 | On Reducing Dimensionality of Labeled Data Efficiently
Guoxi Zhang, Tomoharu Iwata, Hisashi Kashima |
PAKDD (3) | 2 |
| 2018 | Topic Models for Unsupervised Cluster MatchingabstractWe propose topic models for unsupervised cluster matching, which is the task of finding matching between clusters in different domains without correspondence information. For example, the proposed model finds correspondence between document clusters in English and German without alignment information, such as dictionaries and parallel sentences/documents. The proposed model assumes that documents in all languages have a common latent topic structure, and there are potentially infinite number of topic proportion vectors in a latent topic space that is shared by all languages. Each document is generated using one of the topic proportion vectors and language-specific word distributions. By inferring a topic proportion vector used for each document, we can allocate documents in different languages into common clusters, where each cluster is associated with a topic proportion vector. Documents assigned into the same cluster are considered to be matched. We develop an efficient inference procedure for the proposed model based on collapsed Gibbs sampling. The effectiveness of the proposed model is demonstrated with real data sets including multilingual corpora of Wikipedia and product reviews. Tomoharu Iwata, Tsutomu Hirao, Naonori Ueda |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Latent Dimensionality Estimation for Probabilistic Canonical Correlation Analysis Using Normalized Maximum Likelihood Code-LengthabstractDiscovering hidden common factors from multiple different but related datasets is an important task in data mining. Probabilistic canonical correlation analysis (PCCA) is successfully used for this task, where private factors, which represent independent factors that have influence on a dataset, are modeled as well as common factors. We propose a method for estimating the latent dimensionality of PCCA, which represents the numbers of common and private factors. The dimensionality estimation is indispensable for both generalization ability and interpretability. The proposed method applies the minimum description length criterion using normalized maximum likelihood coding to PCCA in a theoretically justified manner, where PCCA is transformed to a regular model by latent variable completion. We demonstrate that the proposed method surpasses conventional methods in terms of dimensionality estimation performance. Tomohiko Nakmaura, Tomoharu Iwata, Kenji Yamanishi |
DSAA | 2 |
| 2017 | Structurally Regularized Non-negative Tensor Factorization for Spatio-Temporal Pattern Discoveries
Koh Takeuchi 0001, Yoshinobu Kawahara, Tomoharu Iwata |
ECML/PKDD (1) | 3 |
| 2017 | Robust Multi-view Topic Modeling by Incorporating Detecting Anomalies
Guoxi Zhang, Tomoharu Iwata, Hisashi Kashima |
ECML/PKDD (2) | 2 |
| 2017 | Scaling Locally Linear EmbeddingabstractLocally Linear Embedding (LLE) is a popular approach to dimensionality reduction as it can effectively represent nonlinear structures of high-dimensional data. For dimensionality reduction, it computes a nearest neighbor graph from a given dataset where edge weights are obtained by applying the Lagrange multiplier method, and it then computes eigenvectors of the LLE kernel where the edge weights are used to obtain the kernel. Although LLE is used in many applications, its computation cost is significantly high. This is because, in obtaining edge weights, its computation cost is cubic in the number of edges to each data point. In addition, the computation cost in obtaining the eigenvectors of the LLE kernel is cubic in the number of data points. Our approach, Ripple, is based on two ideas: (1) it incrementally updates the edge weights by exploiting the Woodbury formula and (2) it efficiently computes eigenvectors of the LLE kernel by exploiting the LU decomposition-based inverse power method. Experiments show that Ripple is significantly faster than the original approach of LLE by guaranteeing the same results of dimensionality reduction. Yasuhiro Fujiwara, Naoki Marumo, Mathieu Blondel, Koh Takeuchi 0001, Hideaki Kim, Tomoharu Iwata, Naonori Ueda |
SIGMOD Conference | 6 |
| 2017 | Robust unsupervised cluster matching for network data
Tomoharu Iwata, Katsuhiko Ishiguro |
Data Min. Knowl. Discov. | 1 |
| 2017 | Unsupervised group matching with application to cross-lingual topic matching without alignment information
Tomoharu Iwata, Motonobu Kanagawa, Tsutomu Hirao, Kenji Fukumizu |
Data Min. Knowl. Discov. | 1 |
| 2016 | Inferring Latent Triggers of Purchases with Consideration of Social Effects and Media AdvertisementsabstractThis paper proposes a method for inferring from single-source data the factors that trigger purchases. Here, single-source data are the histories of item purchases and media advertisement views for each individual. We assume a sequence of purchase events to be a stochastic process incorporating the following three factors: (a) user preference, (b) social effects received from other users, and (c) media advertising effects. As our user-purchase model incorporates the latent relationships between users and advertisers, it can infer the latent triggers of purchases. Experiments on real single-source data show that our model can (a) achieve high prediction accuracy for purchases, (b) discover the key information, i.e., popular items, influential users, and influential advertisers, (c) estimate the relative impact of the three factors on purchases, and (d) find user segments according to the estimated factors. Yusuke Tanaka 0002, Takeshi Kurashima, Yasuhiro Fujiwara, Tomoharu Iwata, Hiroshi Sawada |
WSDM | 4 |
| 2016 | Probabilistic latent variable models for unsupervised many-to-many object matching
Tomoharu Iwata, Tsutomu Hirao, Naonori Ueda |
Inf. Process. Manag. | 1 |
| 2015 | Higher Order Fused Regularization for Supervised Learning with Grouped Parameters
Koh Takeuchi 0001, Yoshinobu Kawahara, Tomoharu Iwata |
ECML/PKDD (1) | 3 |
| 2014 | Probabilistic latent network visualization: inferring and embedding diffusion networksabstractThe diffusion of information, rumors, and diseases are assumed to be probabilistic processes over some network structure. An event starts at one node of the network, and then spreads to the edges of the network. In most cases, the underlying network structure that generates the diffusion process is unobserved, and we only observe the times at which each node is altered/influenced by the process. This paper proposes a probabilistic model for inferring the diffusion network, which we call Probabilistic Latent Network Visualization (PLNV); it is based on cascade data, a record of observed times of node influence. An important characteristic of our approach is to infer the network by embedding it into a low-dimensional visualization space. We assume that each node in the network has latent coordinates in the visualization space, and diffusion is more likely to occur between nodes that are placed close together. Our model uses maximum a posteriori estimation to learn the latent coordinates of nodes that best explain the observed cascade data. The latent coordinates of nodes in the visualization space can 1) enable the system to suggest network layouts most suitable for browsing, and 2) lead to high accuracy in inferring the underlying network when analyzing the diffusion process of new or rare information, rumors, and disease. Takeshi Kurashima, Tomoharu Iwata, Noriko Takaya, Hiroshi Sawada |
KDD | 2 |
| 2013 | A Probabilistic Model for Diversifying Recommendation Lists
Yutaka Kabutoya, Tomoharu Iwata, Hiroyuki Toda, Hiroyuki Kitagawa |
APWeb | 2 |
| 2013 | Clustering-based anomaly detection in multi-view dataabstractThis paper proposes a simple yet effective anomaly detection method for multi-view data. The proposed approach detects anomalies by comparing the neighborhoods in different views. Specifically, clustering is performed separately in the different views and affinity vectors are derived for each object from the clustering results. Then, the anomalies are detected by comparing affinity vectors in the multiple views. An advantage of the proposed method over existing methods is that the tuning parameters can be determined effectively from the given data. Through experiments on synthetic and benchmark datasets, we show that the proposed method outperforms existing methods. Alejandro Marcos Alvarez, Makoto Yamada, Akisato Kimura, Tomoharu Iwata |
CIKM | 4 |
| 2013 | A Probabilistic Behavior Model for Discovering Unrecognized KnowledgeabstractDiscovering interesting behavior patterns and profiles of users as they interact with E-commerce (EC) sites is an important task for site managers. We propose a probabilistic behavior model for extracting latent classes of items that impact the users' item selections but cannot be inferred from the current knowledge of the managers. The proposed model assumes that the current knowledge is represented by categories of items that are defined in the EC site, and a user selects items depending on both of their categories and latent classes. By estimating latent classes, each of which shows items accessed by users with common interests, we can find interesting factors for explaining user behavior. We evaluate our proposed model using item-access log data observed in an EC site. The results show that our model can accurately predict users' item selection, and actually discover latent classes of items having similar latent characteristic such as "colored design" and "impression" by using item categories such as "coat" and "hat" as the current knowledge of the managers. Takeshi Kurashima, Tomoharu Iwata, Noriko Takaya, Hiroshi Sawada |
ICDM | 2 |
| 2013 | Discovering latent influence in online social activities via shared cascade poisson processesabstractMany people share their activities with others through online communities. These shared activities have an impact on other users' activities. For example, users are likely to become interested in items that are adopted (e.g. liked, bought and shared) by their friends. In this paper, we propose a probabilistic model for discovering latent influence from sequences of item adoption events. An inhomogeneous Poisson process is used for modeling a sequence, in which adoption by a user triggers the subsequent adoption of the same item by other users. For modeling adoption of multiple items, we employ multiple inhomogeneous Poisson processes, which share parameters, such as influence for each user and relations between users. The proposed model can be used for finding influential users, discovering relations between users and predicting item popularity in the future. We present an efficient Bayesian inference procedure of the proposed model based on the stochastic EM algorithm. The effectiveness of the proposed model is demonstrated by using real data sets in a social bookmark sharing service. Tomoharu Iwata, Amar Shah 0001, Zoubin Ghahramani |
KDD | 1 |
| 2013 | Geo topic model: joint modeling of user's activity area and interests for location recommendationabstractThis paper proposes a method that analyzes the location log data of multiple users to recommend locations to be visited. The method uses our new topic model, called Geo Topic Model, that can jointly estimate both the user's interests and activity area hosting the user's home, office and other personal places. By explicitly modeling geographical features of locations and users, the user's interests in other features of locations, which we call latent topics, can be inferred effectively. The topic interests estimated by our model 1) lead to high accuracy in predicting visit behavior as driven by personal interests, 2) make possible the generation of recommendations when the user is in an unfamiliar area (e.g. sightseeing), and 3) enable the recommender system to suggest an interpretable representation of the user profile that can be customized by the user. Experiments are conducted using real location logs of landmark and restaurant visits to evaluate the recommendation performance of the proposed method in terms of the accuracy of predicting visit selections. We also show that our model can estimate latent features of locations such as art, nature and atmosphere as latent topics, and describe each user's preference based on them. Takeshi Kurashima, Tomoharu Iwata, Takahide Hoshide, Noriko Takaya, Ko Fujimura |
WSDM | 2 |
| 2013 | Topic model for analyzing purchase data with price information
Tomoharu Iwata, Hiroshi Sawada |
Data Min. Knowl. Discov. | 1 |
| 2013 | Travel route recommendation using geotagged photos
Takeshi Kurashima, Tomoharu Iwata, Go Irie, Ko Fujimura |
Knowl. Inf. Syst. | 2 |
| 2013 | Modeling Noisy Annotated Data with Application to Social AnnotationabstractWe propose a probabilistic topic model for analyzing and extracting content-related annotations from noisy annotated discrete data such as webpages stored using social bookmarking services. With these services, because users can attach annotations freely, some annotations do not describe the semantics of the content, thus they are noisy, i.e., not content related. The extraction of content-related annotations can be used as a prepossessing step in machine learning tasks such as text classification and image recognition, or can improve information retrieval performance. The proposed model is a generative model for content and annotations, in which the annotations are assumed to originate either from topics that generated the content or from a general distribution unrelated to the content. We demonstrate the effectiveness of the proposed method by using synthetic data and real social annotation data for text and images. Tomoharu Iwata, Takeshi Yamada, Naonori Ueda |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Creating Stories: Social Curation of Twitter Messages
Kevin Duh, Tsutomu Hirao, Akisato Kimura, Katsuhiko Ishiguro, Tomoharu Iwata, Ching-man Au Yeung |
ICWSM | 5 |
| 2012 | Fast mining and forecasting of complex time-stamped eventsabstractGiven huge collections of time-evolving events such as web-click logs, which consist of multiple attributes (e.g., URL, userID, times- tamp), how do we find patterns and trends? How do we go about capturing daily patterns and forecasting future events? We need two properties: (a) effectiveness, that is, the patterns should help us understand the data, discover groups, and enable forecasting, and (b) scalability, that is, the method should be linear with the data size. We introduce TriMine, which performs three-way mining for all three attributes, namely, URLs, users, and time. Specifically TriMine discovers hidden topics, groups of URLs, and groups of users, simultaneously. Thanks to its concise but effective summarization, it makes it possible to accomplish the most challenging and important task, namely, to forecast future events. Extensive experiments on real datasets demonstrate that TriMine discovers meaningful topics and makes long-range forecasts, which are notoriously difficult to achieve. In fact, TriMine consistently outperforms the best state-of-the-art existing methods in terms of accuracy and execution speed (up to 74x faster). Yasuko Matsubara, Yasushi Sakurai, Christos Faloutsos, Tomoharu Iwata, Masatoshi Yoshikawa |
KDD | 4 |
| 2012 | Bidirectional Semi-supervised Learning with Graphs
Tomoharu Iwata, Kevin Duh |
ECML/PKDD (2) | 1 |
| 2012 | A Topic Model for Recommending Movies via Linked Open DataabstractWe propose an algorithm for recommending both well-watched old movies and unwatched new ones. To recommend both old favourites and new releases, hybrids of collaborative and content-based filtering are the most suitable methods. However, hybrid movie recommenders have two issues. First, it is necessary to acquire content-descriptive metadata, which is not always easily available. Second, the metadata, once acquired, may be noisy, which can damage recommendation accuracy. In our algorithm, we address the first issue by automatically drawing movie metadata from Linked Open Data, and the second by modeling the relevance of the collected metadata to the transaction history before using the relationship between them to make recommendations. We experimentally demonstrate that our method can effectively collect metadata from LOD, and that our method outperforms conventional hybrid methods found in the literature in both well-watched and unwatched movie recommendation using the noisy collected movie metadata. Yutaka Kabutoya, Róbert Sumi, Tomoharu Iwata, Toshio Uchiyama, Tadasu Uchiyama |
Web Intelligence | 3 |
| 2012 | Sequential Modeling of Topic Dynamics with Multiple TimescalesabstractWe propose an online topic model for sequentially analyzing the time evolution of topics in document collections. Topics naturally evolve with multiple timescales. For example, some words may be used consistently over one hundred years, while other words emerge and disappear over periods of a few days. Thus, in the proposed model, current topic-specific distributions over words are assumed to be generated based on the multiscale word distributions of the previous epoch. Considering both the long- and short-timescale dependency yields a more robust model. We derive efficient online inference procedures based on a stochastic EM algorithm, in which the model is sequentially updated using newly obtained data; this means that past data are not required to make the inference. We demonstrate the effectiveness of the proposed method in terms of predictive performance and computational efficiency by examining collections of real documents with timestamps. Tomoharu Iwata, Takeshi Yamada, Yasushi Sakurai, Naonori Ueda |
ACM Trans. Knowl. Discov. Data | 1 |
| 2011 | Extracting multi-dimensional relations: a generative model of groups of entities in a corpusabstractExtracting relations among different entities from various data sources has been an important topic in data mining. While many methods focus only on a single type of relations, real world entities maintain relations that contain much richer information. We propose a hierarchical Bayesian model for extracting multi-dimensional relations among entities from a text corpus. Using data from Wikipedia, we show that our model can accurately predict the relevance of an entity given the topic of the document as well as the set of entities that are already mentioned in that document. Ching-man Au Yeung, Tomoharu Iwata |
CIKM | 2 |
| 2011 | Strength of social influence in trust networks in product review sitesabstractSome popular product review sites such as Epinions allow users to establish a trust network among themselves, indicating who they trust in providing product reviews and ratings. While trust relations have been found to be useful in generating personalised recommendations, the relations between trust and product ratings has so far been overlooked. In this paper, we examine large datasets collected from Epinions and Ciao, two popular product review sites. We discover that in general users who trust each other tend to have smaller differences in their ratings as time passes, giving support to the theories of homophily and social influence. However, we also discover that this does not hold true across all trusted users. A trust relation does not guarantee that two users have similar preferences, implying that personalised recommendations based on trust relations do not necessarily produce more accurate predictions. We propose a method to estimate the strengths of trust relations so as to estimate the true influence among the trusted users. Our method extends the popular matrix factorisation technique for collaborative filtering, which allow us to generate more accurate rating predictions at the same time. We also show that the estimated strengths of trust relations correlate with the similarity among the users. Our work contributes to the understanding of the interplay between trust relations and product ratings, and suggests that trust networks may serve as a more general socialising venue than only an indication of similarity in user preferences. Ching-man Au Yeung, Tomoharu Iwata |
WSDM | 2 |
| 2011 | Improving Classifier Performance Using Data with Different TaxonomiesabstractWe propose a framework for improving classifier performance by effectively using auxiliary samples. The auxiliary samples are labeled not in terms of the target taxonomy according to which we wish to classify samples, but according to classification schemes or taxonomies that are different from the target taxonomy. Our method finds a classifier by minimizing a weighted error over the target and auxiliary samples. The weights are defined so that the weighted error approximates the expected error when samples are classified into the target taxonomy. Experiments using synthetic and text data show that our method significantly improves the classifier performance in most cases compared to conventional data augmentation methods. Tomoharu Iwata, Toshiyuki Tanaka 0003, Takeshi Yamada, Naonori Ueda |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2010 | Travel route recommendation using geotags in photo sharing sitesabstractThe ability to create geotagged photos enables people to share their personal experiences as tourists at specific locations and times. Assuming that the collection of each photographer's geotagged photos is a sequence of visited locations, photo-sharing sites are important sources for gathering the location histories of tourists. By following their location sequences, we can find representative and diverse travel routes that link key landmarks. In this paper, we propose a travel route recommendation method that makes use of the photographers' histories as held by Flickr. Recommendations are performed by our photographer behavior model, which estimates the probability of a photographer visiting a landmark. We incorporate user preference and present location information into the probabilistic behavior model by combining topic models and Markov models. We demonstrate the effectiveness of the proposed method using a real-life dataset holding information from 71,718 photographers taken in the United States in terms of the prediction accuracy of travel behavior. Takeshi Kurashima, Tomoharu Iwata, Go Irie, Ko Fujimura |
CIKM | 2 |
| 2010 | Effective Question Recommendation Based on Multiple Features for Question Answering Communities
Yutaka Kabutoya, Tomoharu Iwata, Hisako Shiohara, Ko Fujimura |
ICWSM | 2 |
| 2010 | Online multiscale dynamic topic modelsabstractWe propose an online topic model for sequentially analyzing the time evolution of topics in document collections. Topics naturally evolve with multiple timescales. For example, some words may be used consistently over one hundred years, while other words emerge and disappear over periods of a few days. Thus, in the proposed model, current topic-specific distributions over words are assumed to be generated based on the multiscale word distributions of the previous epoch. Considering both the long-timescale dependency as well as the short-timescale dependency yields a more robust model. We derive efficient online inference procedures based on a stochastic EM algorithm, in which the model is sequentially updated using newly obtained data; this means that past data are not required to make the inference. We demonstrate the effectiveness of the proposed method in terms of predictive performance and computational efficiency by examining collections of real documents with timestamps. Tomoharu Iwata, Takeshi Yamada, Yasushi Sakurai, Naonori Ueda |
KDD | 1 |
| 2010 | Modeling Multiple Users' Purchase over a Single Account for Collaborative Filtering
Yutaka Kabutoya, Tomoharu Iwata, Ko Fujimura |
WISE | 2 |
| 2008 | Probabilistic latent semantic visualization: topic model for visualizing documentsabstractWe propose a visualization method based on a topic model for discrete data such as documents. Unlike conventional visualization methods based on pairwise distances such as multi-dimensional scaling, we consider a mapping from the visualization space into the space of documents as a generative process of documents. In the model, both documents and topics are assumed to have latent coordinates in a two- or three-dimensional Euclidean space, or visualization space. The topic proportions of a document are determined by the distances between the document and the topics in the visualization space, and each word is drawn from one of the topics according to its topic proportions. A visualization, i.e. latent coordinates of documents, can be obtained by fitting the model to a given set of documents using the EM algorithm, resulting in documents with similar topics being embedded close together. We demonstrate the effectiveness of the proposed model by visualizing document and movie data sets, and quantitatively compare it with conventional visualization methods. Tomoharu Iwata, Takeshi Yamada, Naonori Ueda |
KDD | 1 |
| 2008 | Recommendation Method for Improving Customer Lifetime ValueabstractIt is important for online stores to improve customer lifetime value (LTV) if they are to increase their profits. Conventional recommendation methods suggest items that best coincide with user's interests to maximize the purchase probability, and this does not necessarily help improve LTV. We present a novel recommendation method that maximizes the probability of the LTV being improved, which can apply to both measured and subscription services. Our method finds frequent purchase patterns among high-LTV users and recommends items for a new user that simulate the found patterns. Using survival analysis techniques, we efficiently find the patterns from log data. Furthermore, we infer a user's interests from the purchase history based on maximum entropy models and use the interests to improve recommendation. Since a higher LTV is the result of greater user satisfaction, our method benefits users as well as online stores. We evaluate our method using two sets of real log data for measured and subscription services. Tomoharu Iwata, Kazumi Saito, Takeshi Yamada |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2007 | Modeling user behavior in recommender systems based on maximum entropyabstractWe propose a model for user purchase behavior in online stores that provide recommendation services. We model the purchase probability given recommendations for each user based on the maximum entropy principle using features that deal with recommendations and user interests. The proposed model enable us to measure the effect of recommendations on user purchase behavior, and the effect can be used to evaluate recommender systems. We show the validity of our model using the log data of an online cartoon distribution service, and measure the recommendation effects for evaluating the recommender system. Tomoharu Iwata, Kazumi Saito, Takeshi Yamada |
WWW | 1 |
| 2006 | Recommendation method for extending subscription periodsabstractOnline stores providing subscription services need to extend user subscription periods as long as possible to increase their profits. Conventional recommendation methods recommend items that best coincide with user's interests to maximize the purchase probability, which does not necessarily contribute to extend subscription periods. We present a novel recommendation method for subscription services that maximizes the probability of the subscription period being extended. Our method finds frequent purchase patterns in the long subscription period users, and recommends items for a new user to simulate the found patterns. Using survival analysis techniques, we efficiently extract information from the log data for finding the patterns. Furthermore, we infer user's interests from purchase histories based on maximum entropy models, and use the interests to improve the recommendations. Since a longer subscription period is the result of greater user satisfaction, our method benefits users as well as online stores. We evaluate our method using the real log data of an online cartoon distribution service for cell-phone in Japan. Tomoharu Iwata, Kazumi Saito, Takeshi Yamada |
KDD | 1 |