Hiroyuki Toda

dblp:32/4046 · DBLP profile ↗
← Back
57ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 29 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Aggregated Multi-output Gaussian Processes with Knowledge Transfer Across Domains
Yusuke Tanaka 0002, Toshiyuki Tanaka 0003, Tomoharu Iwata, Takeshi Kurashima, Maya Okawa, Yasunori Akagi, Hiroyuki Toda
Mach. Learn.7
2025 Estimating Impact of Behavior Change Messages Using Large Language Models
Takuya Okada, Yoshiaki Takimoto, Takeshi Kurashima, Hiroyuki Toda
PAKDD (7)4
2024 Exploring of STGNN for Traffic Forecasting at Expanding Traffic Network
Tomoki Kawabata, Hiroyuki Toda
DEXA (2)2
2024 Understanding Human Mobility Characteristics Through Behavior and Corresponding Environmental Information
Ryuichi Sudo, Hiroyuki Toda
DEXA (2)2
2023 Personal History Affects Reference Points: A Case Study of Codeforces
abstract
Humans make decisions based on their internal value function, and its shape is known to be distorted and biased around a point, which the research community of behavior economics refers to as the reference point. People intensify activities that come to lie within the reach of their reference point, and abstain from acts that would incur losses once they've crossed the point. However, the impact of past experiences on decision making around the reference point has not been well studied. By analyzing a long series of user-level decisions gathered from a competitive programming website, we find that history has a clear impact on user's decision making around the reference point. Past experiences can strengthen, and sometimes weaken, the decision bias around the reference point. Experiences of past difficulties can strengthen the tendency towards loss aversion after achieving the reference point. When a person crosses a reference point for the first time, the cognitive decision bias is significant. However, repeating this crossing gradually weakens the effect. We also show the value of our insights in the task of predicting user behavior. Prediction models incorporating our insights may be used for motivating people to remain more active.
Takeshi Kurashima, Tomoharu Iwata, Tomu Tominaga, Shuhei Yamamoto, Hiroyuki Toda, Kazuhisa Takemura
ICWSM5
2023 Inverse Problem of Censored Markov Chain: Estimating Markov Chain Parameters from Censored Transition Data
Masahiro Kohjima, Takeshi Kurashima, Hiroyuki Toda
PAKDD (2)3
2023 MAP inference algorithms without approximation for collective graphical models on path graphs via discrete difference of convex algorithm
abstract
Abstract Collective graphical model (CGM) is a probabilistic model that provides a framework for analyzing aggregated count data. Maximum a posteriori (MAP) inference of unobserved variables under given observations is one of the essential operations in CGM. Because the MAP inference problem is known to be NP-hard in general, the current mainstream approach is to solve an alternative problem obtained by approximating the objective function and applying continuous relaxation. However, this approach has two significant drawbacks. First, the quality of the solution deteriorates when the values in the count data are negligible due to the inaccuracy of Stirling’s approximation. Second, the application of continuous relaxation causes the violation of integrality constraints. This paper proposes novel algorithms for MAP inference in CGMs on path graphs to overcome these problems. Our method is based on the discrete difference of convex algorithm (DCA); DCA is a general framework to minimize the sum of a convex function and a concave function by repeatedly minimizing surrogate functions. Utilizing the particular structure of path graphs, we efficiently solve the surrogate function minimization by minimum convex cost flow algorithms. Furthermore, our approach also leads to a new method of solving another important task; MAP inference of the sample size in CGM on path graphs. Our method is naturally applicable to this task, allowing us to design very efficient algorithms. Experimental results on synthetic and real-world datasets show the effectiveness of the proposed algorithms.
Yasunori Akagi, Naoki Marumo, Hideaki Kim, Takeshi Kurashima, Hiroyuki Toda
Mach. Learn.5
2022 Joint Modeling of Multi-Sample and Subband Signals for Fast Neural Vocoding on CPU
Hiroki Kanagawa, Yusuke Ijima, Hiroyuki Toda
INTERSPEECH3
2022 Fast Bayesian Estimation of Point Process Intensity as Function of Covariates
abstract
In this paper, we tackle the Bayesian estimation of point process intensity as a function of covariates. We propose a novel augmentation of permanental process called augmented permanental process, a doubly-stochastic point process that uses a Gaussian process on covariate space to describe the Bayesian a priori uncertainty present in the square root of intensity, and derive a fast Bayesian estimation algorithm that scales linearly with data size without relying on either domain discretization or Markov Chain Monte Carlo computation. The proposed algorithm is based on a non-trivial finding that the representer theorem, one of the most desirable mathematical property for machine learning problems, holds for the augmented permanental process, which provides us with many significant computational advantages. We evaluate our algorithm on synthetic and real-world data, and show that it outperforms state-of-the-art methods in terms of predictive accuracy while being substantially faster than a conventional Bayesian method.
Hideaki Kim, Taichi Asami, Hiroyuki Toda
NeurIPS3
2022 Context-aware spatio-temporal event prediction via convolutional Hawkes processes
abstract
Massive spatio-temporal event data sets are now available that cover events such as disease outbreaks, armed conflicts and crimes. Predicting such events and revealing the underlying triggering patterns are a crucial task for many applications, ranging from disease control to global politics. Traditional event prediction models based on Hawkes processes capture the spatio-temporal relationships between events, but cannot incorporate complex and heterogeneous external features, including population distribution, weather and terrain. This paper proposes an event prediction method that effectively utilizes the rich external information present in sets of unstructured data (e.g., map images, satellite images and weather map). Specifically, we extend a convolutional neural network (CNN) by combining it with continuous kernel convolution; and design the conditional intensity of Hawkes process based on the extended neural network model that accepts images as its input. Our approach of using the continuous convolution kernel provides a flexible way to discover the complex effect of external factors on the triggering process, as well as yielding tractable optimization algorithms. We use real-world event data from different domains (i.e., disease outbreaks, armed conflicts and protests) to demonstrate that the proposed method has better prediction performance than existing methods.
Maya Okawa, Tomoharu Iwata, Yusuke Tanaka 0002, Takeshi Kurashima, Hiroyuki Toda, Hisashi Kashima
Mach. Learn.5
2021 Integrated Optimization of Bipartite Matching and Its Stochastic Behavior: New Formulation and Approximation Algorithm via Min-cost Flow Optimization
abstract
The research field of stochastic matching has yielded many developments for various applications. In most stochastic matching problems, the probability distributions inherent in the nodes and edges are set a priori, and are not controllable. However, many matching services have options, which we call control variables, that affect the probability distributions and thus what constitutes an optimum matching. Although several methods for optimizing the values of the control variables have been developed, their optimization in consideration of the matching problem is still in its infancy. In this paper, we formulate an optimization problem for determining the values of the control variables so as to maximize the expected value of matching weights. Since this problem involves hard to evaluate objective values and is non-convex, we construct an approximation algorithm via a minimum-cost flow algorithm that can find 3-approximation solutions rapidly. Simulations on real data from a ride-hailing platform and a crowd-sourcing market show that the proposed method can find solutions with high profits of the service provider in practical time.
Yuya Hikima, Yasunori Akagi, Hideaki Kim, Masahiro Kohjima, Takeshi Kurashima, Hiroyuki Toda
AAAI6
2021 Asterisk-Shaped Features for Tabular Data
abstract
Data often accumulates in tabular format with many attribute items, and prediction using machine learning adds value to data for business. However, studies on machine learning for tabular data only input attribute values, which reduces accuracy. Therefore, we propose an inference method that inputs attribute values and values from aggregated tabular data that has varying attribute values for each attribute item. In an experiment, we compared our proposed method with AutoGluon-Tabular using AutoML benchmark datasets. Our proposed method achieved the highest accuracy for 21 out of 39 datasets.
Yuki Kurauchi, Yoshiaki Takimoto, Shuhei Yamamoto, Shunichi Seko, Hiroyuki Toda
CIKM5
2021 Effects of Personal Characteristics on Temporal Response Patterns in Ecological Momentary Assessments
Tomu Tominaga, Shuhei Yamamoto, Takeshi Kurashima, Hiroyuki Toda
INTERACT (5)4
2021 Dynamic Hawkes Processes for Discovering Time-evolving Communities' States behind Diffusion Processes
abstract
Sequences of events including infectious disease outbreaks, social network activities, and crimes are ubiquitous and the data on such events carry essential information about the underlying diffusion processes between communities (e.g., regions, online user groups). Modeling diffusion processes and predicting future events are crucial in many applications including epidemic control, viral marketing, and predictive policing. Hawkes processes offer a central tool for modeling the diffusion processes, in which the influence from the past events is described by the triggering kernel. However, the triggering kernel parameters, which govern how each community is influenced by the past events, are assumed to be static over time. In the real world, the diffusion processes depend not only on the influences from the past, but also the current (time-evolving) states of the communities, e.g., people's awareness of the disease and people's current interests. In this paper, we propose a novel Hawkes process model that is able to capture the underlying dynamics of community states behind the diffusion processes and predict the occurrences of events based on the dynamics. Specifically, we model the latent dynamic function that encodes these hidden dynamics by a mixture of neural networks. Then we design the triggering kernel using the latent dynamic function and its integral. The proposed method, termed DHP (Dynamic Hawkes Processes), offers a flexible way to learn complex representations of the time-evolving communities' states, while at the same time it allows to computing the exact likelihood, which makes parameter learning tractable. Extensive experiments on four real-world event datasets show that DHP outperforms five widely adopted methods for event prediction.
Maya Okawa, Tomoharu Iwata, Yusuke Tanaka 0002, Hiroyuki Toda, Takeshi Kurashima, Hisashi Kashima
KDD4
2021 Non-approximate Inference for Collective Graphical Models on Path Graphs via Discrete Difference of Convex Algorithm
abstract
The importance of aggregated count data, which is calculated from the data of multiple individuals, continues to increase. Collective Graphical Model (CGM) is a probabilistic approach to the analysis of aggregated data. One of the most important operations in CGM is maximum a posteriori (MAP) inference of unobserved variables under given observations. Because the MAP inference problem for general CGMs has been shown to be NP-hard, an approach that solves an approximate problem has been proposed. However, this approach has two major drawbacks. First, the quality of the solution deteriorates when the values in the count tables are small, because the approximation becomes inaccurate. Second, since continuous relaxation is applied, the integrality constraints of the output are violated. To resolve these problems, this paper proposes a new method for MAP inference for CGMs on path graphs. Our method is based on the Difference of Convex Algorithm (DCA), which is a general methodology to minimize a function represented as the sum of a convex function and a concave function. In our algorithm, important subroutines in DCA can be efficiently calculated by minimum convex cost flow algorithms. Experiments show that the proposed method outputs higher quality solutions than the conventional approach.
Yasunori Akagi, Naoki Marumo, Hideaki Kim, Takeshi Kurashima, Hiroyuki Toda
NeurIPS5
2021 Price and Time Optimization for Utility-Aware Taxi Dispatching
Yuya Hikima, Masahiro Kohjima, Yasunori Akagi, Takeshi Kurashima, Hiroyuki Toda
PRICAI (1)5
2021 Time-delayed collective flow diffusion models for inferring latent people flow from aggregated data at limited locations
abstract
The rapid adoption of wireless sensor devices has made it easier to record location information of people in a variety of spaces (e.g., exhibition halls). Location information is often aggregated due to privacy and/or cost concerns. The aggregated data we use as input consist of the numbers of incoming and outgoing people at each location and at each time step. Since the aggregated data lack tracking information of individuals, determining the flow of people between locations is not straightforward. In this article, we address the problem of inferring latent people flows, that is, transition populations between locations, from just aggregated population data gathered from observed locations. Existing models assume that everyone is always in one of the observed locations at every time step; this, however, is an unrealistic assumption, because we do not always have a large enough number of sensor devices to cover the large-scale spaces targeted. To overcome this drawback, we propose a probabilistic model with flow conservation constraints that incorporate travel duration distributions between observed locations. To handle noisy settings, we adopt noisy observation models for the numbers of incoming and outgoing people, where the noise is regarded as a factor that may disturb flow conservation, e.g., people may appear in or disappear from the predefined space of interest. We develop an approximate expectation-maximization (EM) algorithm that simultaneously estimates transition populations and model parameters. Our experiments demonstrate the effectiveness of the proposed model on real-world datasets of pedestrian data in exhibition halls, bike trip data and taxi trip data in New York City.
Yusuke Tanaka 0002, Tomoharu Iwata, Takeshi Kurashima, Hiroyuki Toda, Naonori Ueda, Toshiyuki Tanaka 0003
Artif. Intell.4
2020 Exact and Efficient Inference for Collective Flow Diffusion Model via Minimum Convex Cost Flow Algorithm
abstract
Collective Flow Diffusion Model (CFDM) is a general framework to find the hidden movements underlying aggregated population data. The key procedure in CFDM analysis is MAP inference of hidden variables. Unfortunately, existing approaches fail to offer exact MAP inferences, only approximate versions, and take a lot of computation time when applied to large scale problems. In this paper, we propose an exact and efficient method for MAP inference in CFDM. Our key idea is formulating the MAP inference problem as a combinatorial optimization problem called Minimum Convex Cost Flow Problem (C-MCFP) with no approximation or continuous relaxation. On the basis of this formulation, we propose an efficient inference method that employs the C-MCFP algorithm as a subroutine. Our experiments on synthetic and real datasets show that the proposed method is effective both in single MAP inference and people flow estimation with EM algorithm.
Yasunori Akagi, Takuya Nishimura 0002, Yusuke Tanaka 0002, Takeshi Kurashima, Hiroyuki Toda
AAAI5
2020 Can Reinforcement Learning Lead to Healthy Life?: Simulation Study Based on User Activity Logs
abstract
The importance of developing an application based on intervention technology that leads to a healthier life is widely recognized. A challenging part of realizing the application is the need for planning, i.e., considering a user's health goal (e.g., sleep at 10:00 p.m. to get enough sleep), providing intervention at the appropriate timing to help the user achieve the goal. The reinforcement learning (RL) approach is well suited to this type of problem since it is a methodology for planning; RL finds the optimal strategy as that which maximizes future expected profit. The purpose of this study is to clarify the effects of intervention based on RL to support healthy daily life. Therefore, we (i) collect real daily activity data from participants, (ii) generate a user model that imitates the user's response to system interventions, (iii) examine valuable goals and design them as rewards in RL and (iv) obtain optimal intervention strategies by RL via simulations given a user model and goals. We evaluate a generated user model and verify by simulations whether our method could successfully achieve the goal. In addition, we analyze the cases that demonstrated higher probability of achieving the goal and report the features.
Masami Takahashi, Masahiro Kohjima, Takeshi Kurashima, Hiroyuki Toda
ICPR4
2020 Learning with Labeled and Unlabeled Multi-Step Transition Data for Recovering Markov Chain from Incomplete Transition Data
abstract
Due to the difficulty of comprehensive data collection, created by factors such as privacy protection and sensor device limitations, we often need to analyze incomplete transition data where some information is missing from the ideal (complete) transition data. In this paper, we propose a new method that can estimate, in a unified manner, Markov chain parameters from incomplete transition data that consist of hidden transition data (data from which visited state information is partially hidden) and dropped transition data (data from which some state visits are dropped). A key to developing the method is regarding the hidden and dropped transition data as labeled and unlabeled multi-step transition data, where the labels represent the number of steps required for each transition. This allows us to describe the generative process of multi-step transition data, and thus develop a new probabilistic model. We confirm the effectiveness of the proposal by experiments on synthetic and real data.
Masahiro Kohjima, Takeshi Kurashima, Hiroyuki Toda
IJCAI3
2020 Identifying Near-Miss Traffic Incidents in Event Recorder Data
Shuhei Yamamoto, Takeshi Kurashima, Hiroyuki Toda
PAKDD (2)3
2019 Refining Coarse-Grained Spatial Data Using Auxiliary Spatial Data Sets with Various Granularities
abstract
We propose a probabilistic model for refining coarse-grained spatial data by utilizing auxiliary spatial data sets. Existing methods require that the spatial granularities of the auxiliary data sets are the same as the desired granularity of target data. The proposed model can effectively make use of auxiliary data sets with various granularities by hierarchically incorporating Gaussian processes. With the proposed model, a distribution for each auxiliary data set on the continuous space is modeled using a Gaussian process, where the representation of uncertainty considers the levels of granularity. The finegrained target data are modeled by another Gaussian process that considers both the spatial correlation and the auxiliary data sets with their uncertainty. We integrate the Gaussian process with a spatial aggregation process that transforms the fine-grained target data into the coarse-grained target data, by which we can infer the fine-grained target Gaussian process from the coarse-grained data. Our model is designed such that the inference of model parameters based on the exact marginal likelihood is possible, in which the variables of finegrained target and auxiliary data are analytically integrated out. Our experiments on real-world spatial data sets demonstrate the effectiveness of the proposed model.
Yusuke Tanaka 0002, Tomoharu Iwata, Toshiyuki Tanaka 0003, Takeshi Kurashima, Maya Okawa, Hiroyuki Toda
AAAI6
2019 Exemplar Based Mixture Models with Censored Data
abstract
In this paper, we propose a method that can handle censored data, data collected under the condition that the exact value is recorded only when the value is within a certain range, abbreviated information is recorded otherwise. It is known that existing methods that use mixture models with censored data suffer from (i) the existence of local optimum solutions and (ii) the need to compute the statistics of truncated distributions for parameter estimation. Our proposal, exemplar based censored mixture model (EBCM), overcomes these two difficulties at once by adopting the exemplar based model approach. The effectiveness of EBCM is confirmed by experiments on synthetic and real world dat sets.
Masahiro Kohjima, Tatsushi Matsubayashi, Hiroyuki Toda
ACML3
2019 Generalized Interval Valued Nonnegative Matrix Factorization
abstract
In this paper, we propose a probabilistic model for analyzing the generalized interval valued matrix, a matrix that has scalar valued elements and bounded/unbounded interval valued elements. We derive a majorization minimization algorithm for parameter estimation and prove that the objective function is monotonically decreasing by the parameter update. An experiment shows that the proposed model well handles interval- valued elements and offers improved performance.
Masahiro Kohjima, Tatsushi Matsubayashi, Hiroyuki Toda
ICASSP3
2019 Deep Mixture Point Processes: Spatio-temporal Event Prediction with Rich Contextual Information
abstract
Predicting when and where events will occur in cities, like taxi pick-ups, crimes, and vehicle collisions, is a challenging and important problem with many applications in fields such as urban planning, transportation optimization and location-based marketing. Though many point processes have been proposed to model events in a continuous spatio-temporal space, none of them allow for the consideration of the rich contextual factors that affect event occurrence, such as weather, social activities, geographical characteristics, and traffic. In this paper, we propose DMPP (Deep Mixture Point Processes), a point process model for predicting spatio-temporal events with the use of rich contextual information; a key advance is its incorporation of the heterogeneous and high-dimensional context available in image and text data. Specifically, we design the intensity of our point process model as a mixture of kernels, where the mixture weights are modeled by a deep neural network. This formulation allows us to automatically learn the complex nonlinear effects of the contextual factors on event occurrence. At the same time, this formulation makes analytical integration over the intensity, which is required for point process estimation, tractable. We use real-world data sets from different domains to demonstrate that DMPP has better predictive performance than existing methods.
Maya Okawa, Tomoharu Iwata, Takeshi Kurashima, Yusuke Tanaka 0002, Hiroyuki Toda, Naonori Ueda
KDD5
2019 Spatially Aggregated Gaussian Processes with Multivariate Areal Outputs
abstract
We propose a probabilistic model for inferring the multivariate function from multiple areal data sets with various granularities. Here, the areal data are observed not at location points but at regions. Existing regression-based models can only utilize the sufficiently fine-grained auxiliary data sets on the same domain (e.g., a city). With the proposed model, the functions for respective areal data sets are assumed to be a multivariate dependent Gaussian process (GP) that is modeled as a linear mixing of independent latent GPs. Sharing of latent GPs across multiple areal data sets allows us to effectively estimate the spatial correlation for each areal data set; moreover it can easily be extended to transfer learning across multiple domains. To handle the multivariate areal data, we design an observation model with a spatial aggregation process for each areal data set, which is an integral of the mixed GP over the corresponding region. By deriving the posterior GP, we can predict the data value at any location point by considering the spatial correlations and the dependences between areal data sets, simultaneously. Our experiments on real-world data sets demonstrate that our model can 1) accurately refine coarse-grained areal data, and 2) offer performance improvements by using the areal data sets from multiple domains.
Yusuke Tanaka 0002, Toshiyuki Tanaka 0003, Tomoharu Iwata, Takeshi Kurashima, Maya Okawa, Yasunori Akagi, Hiroyuki Toda
NeurIPS7
2018 Feasibility Study for Estimation of Depression Severity using Voice Analysis
Mitsuteru Nakamura, Shuji Shinohara, Yasuhiro Omiya, Masakazu Higuchi, Shunji Mitsuyoshi, Takeshi Takano, Hiroyuki Toda, Taku Saito, Masaaki Tanichi, Aihide Yoshino, Shinichi Tokuno
BIBM7
2018 A Fast and Accurate Method for Estimating People Flow from Spatiotemporal Population Data
abstract
Real-time spatiotemporal population data is attracting a great deal of attention for understanding crowd movements in cities.The data is the aggregation of personal location information and consists of just areas and the number of people in each area at certain time instants. Accordingly, it does not explicitly represent crowd movement. This paper proposes a probabilistic model based on collective graphical models that can estimate crowd movement from spatiotemporal population data. There are two technical challenges: (i) poor estimation accuracy as the traditional approach means the model would have too many degrees of freedom, (ii) excessive computation cost. Our key idea for overcoming these two difficulties is to model the transition probability between grid cells (cells hereafter) in a geospatial grid space by using three factors: departure probability of cells, gathering score of cells, and geographical distance between cells. These advances enable us to reduce the degrees of freedom of the model appropriately and derive an efficient estimation algorithm. To evaluate the performance of our method, we conduct experiments using real-world spatiotemporal population data. The results confirm the effectiveness of our method, both in estimation accuracy and computation cost.
Yasunori Akagi, Takuya Nishimura 0002, Takeshi Kurashima, Hiroyuki Toda
IJCAI4
2018 Estimating Latent People Flow without Tracking Individuals
abstract
Analyzing people flows is important for better navigation and location-based advertising. Since the location information of people is often aggregated for protecting privacy, it is not straightforward to estimate transition populations between locations from aggregated data. Here, aggregated data are incoming and outgoing people counts at each location; they do not contain tracking information of individuals. This paper proposes a probabilistic model for estimating unobserved transition populations between locations from only aggregated data. With the proposed model, temporal dynamics of people flows are assumed to be probabilistic diffusion processes over a network, where nodes are locations and edges are paths between locations. By maximizing the likelihood with flow conservation constraints that incorporate travel duration distributions between locations, our model can robustly estimate transition populations between locations. The statistically significant improvement of our model is demonstrated using real-world datasets of pedestrian data in exhibition halls, bike trip data and taxi trip data in New York City.
Yusuke Tanaka 0002, Tomoharu Iwata, Takeshi Kurashima, Hiroyuki Toda, Naonori Ueda
IJCAI4
2018 Variational Bayes for Mixture Models with Censored Data
Masahiro Kohjima, Tatsushi Matsubayashi, Hiroyuki Toda
ECML/PKDD (2)3
2018 Multi Agent Flow Estimation Based on Bayesian Optimization with Time Delay and Low Dimensional Parameter Conversion
Hiroshi Kiyotake, Masahiro Kohjima, Tatsushi Matsubayashi, Hiroyuki Toda
PRIMA4
2018 Time-Series Predictions for People-Flow with Simulation Data
Hengjin Tang, Tatsushi Matsubayashi, Hiroyuki Toda
PRIMA4
2017 Online Traffic Flow Prediction Using Convolved Bilinear Poisson Regression
abstract
Predicting traffic flows on multiple inter-city roads play a critical role in traffic management. Temporal patterns of traffic flow can dynamically change over time as a result of traffic management measures such as the construction of new roads. Given this possibility, it is sensible to use only recent data, instead of all past data. In this study, by incorporating the latent factor model into a conventional approach, we construct a novel traffic prediction method for multiple inter-city roads based on the most recent training data. This formulation leads to a reduction in the number of model parameters, since it assumes that there is a set of patterns underlying the data (in our case, the periodic patterns shared across roads in traffic flow data), which inherently offers robustness against sparse observations. In addition, we adopt stochastic variational Bayes method to solve the optimization problem, which allows us to update model parameters online. By continually updating the model based on the most recent training data, we can instantly provide accurate predictions of inter-city traffic flow that have intermittently changing patterns. Using a real-world traffic flow dataset collected in the Greater Tokyo Area via GPS-equipped mobile phones, we evaluate the predictive performance of the proposed method, and confirm that it performs better than existing methods in the sparse domain.
Maya Okawa, Hideaki Kim, Hiroyuki Toda
MDM3
2017 Predicting Destinations from Partial Trajectories Using Recurrent Neural Network
Yuki Endo 0003, Kyosuke Nishida, Hiroyuki Toda, Hiroshi Sawada
PAKDD (1)3
2016 How fashionable is each street?: Quantifying road characteristics using social media
abstract
Determining routes that provide opportunities to satisfy the various demands of users is still an open problem. This is because it is virtually impossible to manually quantify the characteristics of each road and there are few resources describing roads directly such that we meet any demand that may arise. The goal of this study is to automatically quantify the characteristics of roads for demands that can be described using keywords such as “fashionable”. To achieve this goal, we propose a two-stage method that analyzes social media and road networks. First, our method estimates the topic distribution (i.e., the characteristics) of each point-of-interest (POI) by analyzing geotagged texts with the Latent Dirichlet Allocation model. Next, it uses a Markov random field model to estimate the characteristics of each road on the basis of those of POIs and the road networks associated with the POIs. Experiments on real datasets demonstrate that our method achieves statistically significant improvements over baseline methods in terms of ranking quality in the information retrieval for roads in three areas given 25 keywords.
Takuya Nishimura 0002, Kyosuke Nishida, Hiroyuki Toda, Hiroshi Sawada
ASONAM3
2016 Deep Feature Extraction from Trajectories for Transportation Mode Estimation
Yuki Endo 0003, Hiroyuki Toda, Kyosuke Nishida, Akihisa Kawanobe
PAKDD (2)2
2016 Semantic sensitive tensor factorization
abstract
The ability to predict the activities of users is an important one for recommender systems and analyses of social media. User activities can be represented in terms of relationships involving three or more things (e.g. when a user tags items on a webpage or tweets about a location he or she visited). Such relationships can be represented as a tensor, and tensor factorization is becoming an increasingly important means for predicting users' possible activities. However, the prediction accuracy of factorization is poor for ambiguous and/or sparsely observed objects. Our solution, Semantic Sensitive Tensor Factorization (SSTF), incorporates the semantics expressed by an object vocabulary or taxonomy into the tensor factorization. SSTF first links objects to classes in the vocabulary (taxonomy) and resolves the ambiguities of objects that may have several meanings. Next, it lifts sparsely observed objects to their classes to create augmented tensors. Then, it factorizes the original tensor and augmented tensors simultaneously. Since it shares semantic knowledge during the factorization, it can resolve the sparsity problem. Furthermore, as a result of the natural use of semantic information in tensor factorization, SSTF can combine heterogeneous and unbalanced datasets from different Linked Open Data sources. We implemented SSTF in the Bayesian probabilistic tensor factorization framework. Experiments on publicly available large-scale datasets using vocabularies from linked open data and a taxonomy from WordNet show that SSTF has up to 12% higher accuracy in comparison with state-of-the-art tensor factorization methods.
Makoto Nakatsuji, Hiroyuki Toda, Hiroshi Sawada, Jinguang Zheng, James A. Hendler
Artif. Intell.2
2015 Assigning Tasks to Workers by Referring to Their Schedules in Mobile Crowdsourcing
abstract
This paper focuses on task assignments to workers in mobile crowdsoucing systems. The current method does not work so well since it considers only workers who are ready to work at the time of optimization. Our method handles workers' day-long schedules, creates a `time-extended' worker-task graph that expresses the relationships between workers and tasks over a time period and finds the best set of worker-task-time triples. Our evaluation using real world visiting logs shows it increases the rate of assigned tasks by more than 8.2% compared with a state-of-the-art assignment method.
Mayumi Hadano, Makoto Nakatsuji, Hiroyuki Toda, Yoshimasa Koike
HCOMP3
2014 Semantic Data Representation for Improving Tensor Factorization
abstract
Predicting human activities is important for improving recommender systems or analyzing social relationships among users. Those human activities are usually repre- sented as multi-object relationships (e.g. user’s tagging activities for items or user’s tweeting activities at some locations). Since multi-object relationships are naturally represented as a tensor, tensor factorization is becom- ing more important for predicting users’ possible ac- tivities. However, its prediction accuracy is weak for ambiguous and/or sparsely observed objects. Our so- lution, Semantic data Representation for Tensor Fac- torization (SRTF), tackles these problems by incorpo- rating semantics into tensor factorization based on the following ideas: (1) It first links objects to vocabu- laries/taxonomies and resolves the ambiguity caused by objects that can be used for multiple purposes. (2) It next links objects to composite classes that merge classes in different kinds of vocabularies/taxonomies (e.g. classes in vocabularies for movie genres and those for directors) to avoid low prediction accuracy caused by rough-grained semantics. (3) It then lifts sparsely observed objects into their classes to solve the sparsity problem for rarely observed objects. To the best of our knowledge, this is the first study that leverages seman- tics to inject expert knowledge into tensor factorization. Experiments show that SRTF achieves up to 10% higher accuracy than state-of-the-art methods.
Makoto Nakatsuji, Yasuhiro Fujiwara, Hiroyuki Toda, Hiroshi Sawada, Jinguang Zheng, James A. Hendler
AAAI3
2014 Motivation System Using Purpose-for-Action
Noriko Yokoyama, Kaname Funakoshi, Hiroyuki Toda, Yoshimasa Koike
DEXA (1)3
2014 Probabilistic identification of visited point-of-interest for personalized automatic check-in
abstract
Automatic check-in, which is to identify a user's visited points of interest (POIs) from his or her trajectories, is still an open problem because of positioning errors and the high POI density in small areas. In this study, we propose a probabilistic visited-POI identification method. The method uses a new hierarchical Bayesian model for identifying the latent visited-POI label of stay points, which are automatically extracted from trajectories. This model learns from labeled and unlabeled stay point data (i.e., semi-supervised learning) and takes into account personal preferences, stay locations including positioning errors, stay times for each category, and prior knowledge about typical user preferences and stay times. Experimental results with real user trajectories and POIs of Foursquare demonstrated that our method achieved statistically significant improvements in precision at 1 and recall at 3 over the nearest neighbor method and a conventional method that uses a supervised learning-to-rank algorithm.
Kyosuke Nishida, Hiroyuki Toda, Takeshi Kurashima, Yoshihiko Suhara
UbiComp2
2013 A Probabilistic Model for Diversifying Recommendation Lists
Yutaka Kabutoya, Tomoharu Iwata, Hiroyuki Toda, Hiroyuki Kitagawa
APWeb3
2013 What is he/she like?: estimating Twitter user attributes from contents and social neighbors
abstract
We propose a new method for estimating user attributes (gender, age, occupation, and interests) of a Twitter user from the user's contents (profile document and tweets) and social neighbors, i.e. those whom the user has mentioned. Our labeling method is able to collect a large amount of training data automatically by using Twitter users associated with a blog account. Furthermore, we experiment estimation methods using social neighbors with three adjustable levels of its information and show that our method, which uses the target user's profile document and tweets and the neighbors' profile documents (not including tweets), achieves the best accuracy.
Jun Ito, Takahide Hoshide, Hiroyuki Toda, Tadasu Uchiyama, Kyosuke Nishida
ASONAM3
2012 Collaborative Filtering by Analyzing Dynamic User Interests Modeled by Taxonomy
Makoto Nakatsuji, Yasuhiro Fujiwara, Toshio Uchiyama, Hiroyuki Toda
ISWC (1)4
2010 Detecting periodic changes in search intentions in a search engine
abstract
Information needs expressed by using the same query for a search engine might be totally different, whether on week days or weekends, or during the day or at night. For queries having no temporal changes in search intentions, the same search results ranking may be returned regardless of the time, but for those with temporal changes the ranking must be suitably altered depending on the time of input. To achieve time-dependent search results rankings, we focus on the temporal changes in the search intentions. We present the results obtained by analyzing a commercial search engine log and propose a method of detecting queries showing periodic changes in the search intentions.
Masaya Murata, Hiroyuki Toda, Yumiko Matsuura, Ryoji Kataoka, Takayoshi Mochizuki
CIKM2
2010 Proposal of a method to analyze 3D deformation/fracture characteristics inside materials based on a stratified matching approach
Mitsuru Nakazawa, Masakazu Kobayashi, Hiroyuki Toda, Yoshimitsu Aoki
Mach. Vis. Appl.3
2009 Geographic information retrieval to suit immediate surroundings
abstract
This paper proposes a highly effective geographic information retrieval method. It assesses the extent implied by place names in documents and then emphasizes place names that are highly specific in terms of identifying locations. Furthermore, the method also assesses the proximity between place names and keywords in each document and adjusts the document score based on the proximity between place name associated with user's geographic intention and keywords associated with user's query. Evaluation results show that the two methods proposed herein offer improved performance according to some TREC-style evaluation metrics.
Hiroyuki Toda, Norihito Yasuda, Yumiko Matsuura, Ryoji Kataoka
GIS1
2009 Access concentration detection in click logs to improve mobile Web-IR
Masaya Murata, Hiroyuki Toda, Yumiko Matsuura, Ryoji Kataoka
Inf. Sci.2
2008 Incorporating place name extents into geo-ir ranking
abstract
This paper proposes a novel Geo-IR ranking method that realizes effective searches that emphasize the user's immediate surroundings. It assesses the extent implied by place names in documents and then emphasizes place names that are highly specific in terms of identifying locations.
Hiroyuki Toda, Norihito Yasuda, Yumiko Matsuura, Ryoji Kataoka
CIKM1
2008 Ranking Entities Using Comparative Relations
Takeshi Kurashima, Katsuji Bessho, Hiroyuki Toda, Toshio Uchiyama, Ryoji Kataoka
DEXA3
2008 3D image analysis for evaluating internal deformation/fracture characteristics of materials
abstract
In the past, D/F characteristics, load-deformation relationships until the materials are fractured, have been analyzed on the surface. The D/F characteristics are affected by more than ten thousand micro-scale internal structures like air bubbles (pores), cracks and particles; therefore, it is required to analyze nano-scale D/F characteristics inside materials. In this paper, we propose a method that automatically obtains the corresponding relations of the particles from nano-order 3DCT images at each deformation stage. The particles are deformation-proof and may have different geometries. First of all, some big particles are considered as landmarks and matched between pre- and post-deformation. The results of landmark matching make it easy to match many remaining particles and pores.
Mitsuru Nakazawa, Yoshimitsu Aoki, Masakazu Kobayashi, Hiroyuki Toda
ICPR4
2008 Improving Mobile Web-IR Using Access Concentration Sites in Search Results
Masaya Murata, Hiroyuki Toda, Yumiko Matsuura, Ryoji Kataoka
WISE2
2007 Creating Personal Histories from the Web Using Namesake Disambiguation and Event Extraction
Rui Kimura, Satoshi Oyama, Hiroyuki Toda, Katsumi Tanaka
ICWE3
2007 Event mining from the Blogosphere using topic words
Yoshihiko Suhara, Hiroyuki Toda, Akito Sakurai
ICWSM2
2007 Search Result Clustering Using Informatively Named Entities
abstract
Clustering the results of a search helps the user to review the information gathered. In this article, we regard the clustering task as indexing the search results. Here, an index means a structured label list that can make it easier for the user to comprehend the labels and search results. To realize this goal, we make three proposals. The first is to use Named Entity Extraction for term extraction. The second is to create a new label-selecting criterion based on importance in the search result and the relation between terms and search queries. The third is a label categorization using category information of labels, which is generated by named entity extraction. We implement a prototype system based on these proposals and find that it offers a much higher performance than existing methods; we focus on news articles in this article, but the system is not topic specific.
Hiroyuki Toda, Ryoji Kataoka, Masahiro Oku
Int. J. Hum. Comput. Interact.1
2006 Topic Structure Mining for Document Sets Using Graph-Based Analysis
Hiroyuki Toda, Ryoji Kataoka, Hiroyuki Kitagawa
DEXA1
2001 Goal-Oriented Information Retrieval Using Feedback from Users
Hiroyuki Toda, Toshifumi Enomoto, Tetsuji Satoh
WAIM1