Ryota Tomioka

dblp:50/2945 · DBLP profile ↗
← Back
42ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0002-8092-6553ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 10Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
26 papers
Trustworthy machine learning · 25% Representation and self-supervised learning · 20% Probabilistic and Bayesian machine learning · 18%
Theoretical computer science
9 papers
Mathematical optimization · 62% Algorithms and data structures · 21% Combinatorics and discrete mathematics · 11%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 80% Medical and health informatics · 20%
Databases, data mining, and information retrieval
3 papers
Data mining · 50% Web and social media mining · 48% Information retrieval · 2%

Topics — the 30 heaviest of 81, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization › continuous optimization
convex optimization
0.962015
Interpolating Convex and Non-Convex Tensor Decompositions via the Subspace Norm · NIPS 2015
Multitask learning meets tensor factorization: task imputation via convex optimization · NIPS 2014
Convex Tensor Decomposition via Structured Schatten Norm Regularization · NIPS 2013
Machine learning › Generative modeling
variational autoencoder
0.722019
Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders · NeurIPS 2019
Multi-Level Variational Autoencoder: Learning Disentangled Representations From Grouped Observations · AAAI 2018
Algorithms and data structures › numerical linear algebra › matrix and tensor decomposition
tensor decomposition
0.742015
Interpolating Convex and Non-Convex Tensor Decompositions via the Subspace Norm · NIPS 2015
Multitask learning meets tensor factorization: task imputation via convex optimization · NIPS 2014
Convex Tensor Decomposition via Structured Schatten Norm Regularization · NIPS 2013
Computational science and engineering › statistical computing
boltzmann distribution sampling
0.712023
Timewarp: Transferable Acceleration of Molecular Dynamics by Learning Time-Coarsened Dynamics · NeurIPS 2023
Computational science and engineering › computational chemistry › molecular simulation › molecular dynamics
enhanced sampling
0.712023
Timewarp: Transferable Acceleration of Molecular Dynamics by Learning Time-Coarsened Dynamics · NeurIPS 2023
Computational science and engineering › computational chemistry › molecular simulation
molecular dynamics
0.712023
Timewarp: Transferable Acceleration of Molecular Dynamics by Learning Time-Coarsened Dynamics · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
domain generalization
0.512021
An Information-theoretic Approach to Distribution Shifts · NeurIPS 2021
Machine learning › Trustworthy machine learning
fairness
0.512021
An Information-theoretic Approach to Distribution Shifts · NeurIPS 2021
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.512021
An Information-theoretic Approach to Distribution Shifts · NeurIPS 2021
Machine learning › Trustworthy machine learning › dataset bias
selection bias
0.512021
An Information-theoretic Approach to Distribution Shifts · NeurIPS 2021
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412020
On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them · NeurIPS 2020
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.412020
On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them · NeurIPS 2020
Machine learning › Deep learning architectures and training
loss landscape
0.412020
On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
prior network
0.412020
Conservative Uncertainty Estimation By Fitting Prior Networks · ICLR 2020
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412020
Conservative Uncertainty Estimation By Fitting Prior Networks · ICLR 2020
Machine learning › Trustworthy machine learning
robustness
0.422019
On Certifying Non-Uniform Bounds against Adversarial Attacks · ICML 2019
Invariant Common Spatial Patterns: Alleviating Nonstationarities in Brain-Computer Interfacing · NIPS 2007
Machine learning › Trustworthy machine learning › robustness
certified robustness
0.412019
On Certifying Non-Uniform Bounds against Adversarial Attacks · ICML 2019
Machine learning › Representation and self-supervised learning
hierarchical representation
0.412019
Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
0.412019
Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.422015
Condition for perfect dimensionality recovery by variational Bayesian PCA · J. Mach. Learn. Res. 2015
Perfect Dimensionality Recovery by Variational Bayesian PCA · NIPS 2012
Combinatorics and discrete mathematics
algebraic combinatorics
0.422015
The algebraic combinatorial approach for low-rank matrix completion · J. Mach. Learn. Res. 2015
A Combinatorial Algebraic Approach for the Identifiability of Low-Rank Matrix Completion · ICML 2012
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery › matrix completion
low-rank matrix completion
0.422015
The algebraic combinatorial approach for low-rank matrix completion · J. Mach. Learn. Res. 2015
A Combinatorial Algebraic Approach for the Identifiability of Low-Rank Matrix Completion · ICML 2012
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery
matrix completion
0.422015
The algebraic combinatorial approach for low-rank matrix completion · J. Mach. Learn. Res. 2015
A Combinatorial Algebraic Approach for the Identifiability of Low-Rank Matrix Completion · ICML 2012
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.312018
Multi-Level Variational Autoencoder: Learning Disentangled Representations From Grouped Observations · AAAI 2018
Machine learning › Learning paradigms
multi-task learning
0.322014
Multitask learning meets tensor factorization: task imputation via convex optimization · NIPS 2014
A New Multi-task Learning Method for Personalized Activity Recognition · ICDM 2011
Web and social media mining › event detection
emerging topic detection
0.322014
Discovering Emerging Topics in Social Streams via Link-Anomaly Detection · IEEE Trans. Knowl. Data Eng. 2014
Discovering Emerging Topics in Social Streams via Link Anomaly Detection · ICDM 2011
Data mining › anomaly detection › spam detection
link fraud detection
0.322014
Discovering Emerging Topics in Social Streams via Link-Anomaly Detection · IEEE Trans. Knowl. Data Eng. 2014
Discovering Emerging Topics in Social Streams via Link Anomaly Detection · ICDM 2011
Web and social media mining
social network analysis
0.322014
Discovering Emerging Topics in Social Streams via Link-Anomaly Detection · IEEE Trans. Knowl. Data Eng. 2014
Discovering Emerging Topics in Social Streams via Link Anomaly Detection · ICDM 2011
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.312017
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding · NIPS 2017
Machine learning › Efficient and distributed learning › communication compression
gradient quantization
0.312017
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding · NIPS 2017

Methods — techniques the papers use, named apart from their topics

convex optimization · 1.0information bottleneck · 0.8dimensionality reduction · 0.8normalizing flow · 0.7markov chain monte carlo · 0.7variational bayes · 0.5information-theoretic analysis · 0.5uncertainty estimation · 0.4prior networks · 0.4periodic adversarial scheduling · 0.4adversarial training · 0.4augmented lagrangian method · 0.4probabilistic model · 0.3subspace norm · 0.2nuclear norm minimization · 0.2kronecker product · 0.2algebraic combinatorial approach · 0.2anomaly scoring · 0.2
YearPublicationVenuePosition
2024 Latent Representation and Simulation of Markov Processes via Time-Lagged Information Bottleneck
abstract
Markov processes are widely used mathematical models for describing dynamic systems in various fields. However, accurately simulating large-scale systems at long time scales is computationally expensive due to the short time steps required for accurate integration. In this paper, we introduce an inference process that maps complex systems into a simplified representational space and models large jumps in time. To achieve this, we propose Time-lagged Information Bottleneck (T-IB), a principled objective rooted in information theory, which aims to capture relevant temporal features while discarding high-frequency information to simplify the simulation task and minimize the inference error. Our experiments demonstrate that T-IB learns information-optimal representations for accurately modeling the statistical properties and dynamics of the original process at a selected time lag, outperforming existing time-lagged dimensionality reduction methods.
Marco Federici, Patrick Forré, Ryota Tomioka, Bastiaan S. Veeling
ICLR3
2023 Timewarp: Transferable Acceleration of Molecular Dynamics by Learning Time-Coarsened Dynamics
abstract
*Molecular dynamics* (MD) simulation is a widely used technique to simulate molecular systems, most commonly at the all-atom resolution where equations of motion are integrated with timesteps on the order of femtoseconds ($1\textrm{fs}=10^{-15}\textrm{s}$). MD is often used to compute equilibrium properties, which requires sampling from an equilibrium distribution such as the Boltzmann distribution. However, many important processes, such as binding and folding, occur over timescales of milliseconds or beyond, and cannot be efficiently sampled with conventional MD. Furthermore, new MD simulations need to be performed for each molecular system studied. We present *Timewarp*, an enhanced sampling method which uses a normalising flow as a proposal distribution in a Markov chain Monte Carlo method targeting the Boltzmann distribution. The flow is trained offline on MD trajectories and learns to make large steps in time, simulating the molecular dynamics of $10^{5} - 10^{6} \textrm{fs}$. Crucially, Timewarp is *transferable* between molecular systems: once trained, we show that it generalises to unseen small peptides (2-4 amino acids) at all-atom resolution, exploring their metastable states and providing wall-clock acceleration of sampling compared to standard MD. Our method constitutes an important step towards general, transferable algorithms for accelerating MD.
Leon Klein, Andrew Y. K. Foong, Tor Erlend Fjelde, Bruno Mlodozeniec, Marc Brockschmidt, Sebastian Nowozin, Frank Noé, Ryota Tomioka
NeurIPS8
2021 Regularized Policies are Reward Robust
abstract
Entropic regularization of policies in Reinforcement Learning (RL) is a commonly used heuristic to ensure that the learned policy explores the state-space sufficiently before overfitting to a local optimal policy. The primary motivation for using entropy is for exploration and disambiguating optimal policies; however, the theoretical effects are not entirely understood. In this work, we study the more general regularized RL objective and using Fenchel duality; we derive the dual problem which takes the form of an adversarial reward problem. In particular, we find that the optimal policy found by a regularized objective is precisely an optimal policy of a reinforcement learning problem under a worst-case adversarial reward. Our result allows us to reinterpret the popular entropic regularization scheme as a form of robustification. Furthermore, due to the generality of our results, we apply to other existing regularization schemes. Our results thus give insights into the effects of regularization of policies and deepen our understanding of exploration through robust rewards at large.
Hisham Husain, Kamil Ciosek, Ryota Tomioka
AISTATS3
2021 An Information-theoretic Approach to Distribution Shifts
abstract
Safely deploying machine learning models to the real world is often a challenging process. For example, models trained with data obtained from a specific geographic location tend to fail when queried with data obtained elsewhere, agents trained in a simulation can struggle to adapt when deployed in the real world or novel environments, and neural networks that are fit to a subset of the population might carry some selection bias into their decision process.In this work, we describe the problem of data shift from an information-theoretic perspective by (i) identifying and describing the different sources of error, (ii) comparing some of the most promising objectives explored in the recent domain generalization and fair classification literature. From our theoretical analysis and empirical evaluation, we conclude that the model selection procedure needs to be guided by careful considerations regarding the observed data, the factors used for correction, and the structure of the data-generating process.
Marco Federici, Ryota Tomioka, Patrick Forré
NeurIPS2
2020 Conservative Uncertainty Estimation By Fitting Prior Networks
Kamil Ciosek, Vincent Fortuin, Ryota Tomioka, Katja Hofmann, Richard E. Turner
ICLR3
2020 On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them
abstract
We analyze the influence of adversarial training on the loss landscape of machine learning models. To this end, we first provide analytical studies of the properties of adversarial loss functions under different adversarial budgets. We then demonstrate that the adversarial loss landscape is less favorable to optimization, due to increased curvature and more scattered gradients. Our conclusions are validated by numerical analyses, which show that training under large adversarial budgets impede the escape from suboptimal random initialization, cause non-vanishing gradients and make the models' minima found sharper. Based on these observations, we show that a periodic adversarial scheduling (PAS) strategy can effectively overcome these challenges, yielding better results than vanilla adversarial training while being much less sensitive to the choice of learning rate.
Chen Liu 0027, Mathieu Salzmann, Tao Lin 0004, Ryota Tomioka, Sabine Süsstrunk
NeurIPS4
2019 On Certifying Non-Uniform Bounds against Adversarial Attacks
abstract
This work studies the robustness certification problem of neural network models, which aims to find certified adversary-free regions as large as possible around data points. In contrast to the existing approaches that seek regions bounded uniformly along all input features, we consider non-uniform bounds and use it to study the decision boundary of neural network models. We formulate our target as an optimization problem with nonlinear constraints. Then, a framework applicable for general feedforward neural networks is proposed to bound the output logits so that the relaxed problem can be solved by the augmented Lagrangian method. Our experiments show the non-uniform bounds have larger volumes than uniform ones. Compared with normal models, the robust models have even larger non-uniform bounds and better interpretability. Further, the geometric similarity of the non-uniform bounds gives a quantitative, data-agnostic metric of input features’ robustness.
Chen Liu 0027, Ryota Tomioka, Volkan Cevher
ICML2
2019 Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders
abstract
The Variational Auto-Encoder (VAE) is a popular method for learning a generative model and embeddings of the data. Many real datasets are hierarchically structured. However, traditional VAEs map data in a Euclidean latent space which cannot efficiently embed tree-like structures. Hyperbolic spaces with negative curvature can. We therefore endow VAEs with a Poincaré ball model of hyperbolic geometry as a latent space and rigorously derive the necessary methods to work with two main Gaussian generalisations on that space. We empirically show better generalisation to unseen data than the Euclidean counterpart, and can qualitatively and quantitatively better recover hierarchical structures.
Emile Mathieu, Charline Le Lan, Chris J. Maddison, Ryota Tomioka, Yee Whye Teh
NeurIPS4
2018 Multi-Level Variational Autoencoder: Learning Disentangled Representations From Grouped Observations
abstract
We would like to learn a representation of the data that reflects the semantics behind a specific grouping of the data, where within a group the samples share a common factor of variation. For example, consider a set of face images grouped by identity. We wish to anchor the semantics of the grouping into a disentangled representation that we can exploit. However, existing deep probabilistic models often assume that the samples are independent and identically distributed, thereby disregard the grouping information. We present the Multi-Level Variational Autoencoder (ML-VAE), a new deep probabilistic model for learning a disentangled representation of grouped data. The ML-VAE separates the latent representation into semantically relevant parts by working both at the group level and the observation level, while retaining efficient test-time inference. We experimentally show that our model (i) learns a semantically meaningful disentanglement, (ii) enables control over the latent representation, and (iii) generalises to unseen groups.
Diane Bouchacourt, Ryota Tomioka, Sebastian Nowozin
AAAI2
2017 Batch Policy Gradient Methods for Improving Neural Conversation Models
Kirthevasan Kandasamy, Yoram Bachrach, Ryota Tomioka, Daniel Tarlow, David Carter
ICLR (Poster)3
2017 QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
abstract
Parallel implementations of stochastic gradient descent (SGD) have received significant research attention, thanks to its excellent scalability properties. A fundamental barrier when parallelizing SGD is the high bandwidth cost of communicating gradient updates between nodes; consequently, several lossy compresion heuristics have been proposed, by which nodes only communicate quantized gradients. Although effective in practice, these heuristics do not always guarantee convergence, and it is not clear whether they can be improved. In this paper, we propose Quantized SGD (QSGD), a family of compression schemes for gradient updates which provides convergence guarantees. QSGD allows the user to smoothly trade off \emph{communication bandwidth} and \emph{convergence time}: nodes can adjust the number of bits sent per iteration, at the cost of possibly higher variance. We show that this trade-off is inherent, in the sense that improving it past some threshold would violate information-theoretic lower bounds. QSGD guarantees convergence for convex and non-convex objectives, under asynchrony, and can be extended to stochastic variance-reduced techniques. When applied to training deep neural networks for image classification and automated speech recognition, QSGD leads to significant reductions in end-to-end training time. For example, on 16GPUs, we can train the ResNet152 network to full accuracy on ImageNet 1.8x faster than the full-precision variant.
Dan Alistarh, Demjan Grubic, Jerry Li 0001, Ryota Tomioka, Milan Vojnovic
NIPS4
2016 f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization
abstract
Generative neural networks are probabilistic models that implement sampling using feedforward neural networks: they take a random input vector and produce a sample from a probability distribution defined by the network weights. These models are expressive and allow efficient computation of samples and derivatives, but cannot be used for computing likelihoods or for marginalization. The generative-adversarial training method allows to train such models through the use of an auxiliary discriminative neural network. We show that the generative-adversarial approach is a special case of an existing more general variational divergence estimation approach. We show that any $f$-divergence can be used for training generative neural networks. We discuss the benefits of various choices of divergence functions on training complexity and the quality of the obtained generative models.
Sebastian Nowozin, Botond Cseke, Ryota Tomioka
NIPS3
2016 Theoretical and Experimental Analyses of Tensor-Based Regression and Classification
abstract
We theoretically and experimentally investigate tensor-based regression and classification. Our focus is regularization with various tensor norms, including the overlapped trace norm, the latent trace norm, and the scaled latent trace norm. We first give dual optimization methods using the alternating direction method of multipliers, which is computationally efficient when the number of training samples is moderate. We then theoretically derive an excess risk bound for each tensor norm and clarify their behavior. Finally, we perform extensive experiments using simulated and real data and demonstrate the superiority of tensor-based learning methods over vector- and matrix-based learning methods.
Kishan Wimalawarne, Ryota Tomioka, Masashi Sugiyama
Neural Comput.2
2015 Norm-Based Capacity Control in Neural Networks
abstract
We investigate the capacity, convexity and characterization of a general family of norm-constrained feed-forward networks.
Behnam Neyshabur, Ryota Tomioka, Nathan Srebro
COLT2
2015 Interpolating Convex and Non-Convex Tensor Decompositions via the Subspace Norm
abstract
We consider the problem of recovering a low-rank tensor from its noisy observation. Previous work has shown a recovery guarantee with signal to noise ratio $O(n^{\ceil{K/2}/2})$ for recovering a $K$th order rank one tensor of size $n\times \cdots \times n$ by recursive unfolding. In this paper, we first improve this bound to $O(n^{K/4})$ by a much simpler approach, but with a more careful analysis. Then we propose a new norm called the \textit{subspace} norm, which is based on the Kronecker products of factors obtained by the proposed simple estimator. The imposed Kronecker structure allows us to show a nearly ideal $O(\sqrt{n}+\sqrt{H^{K-1}})$ bound, in which the parameter $H$ controls the blend from the non-convex estimator to mode-wise nuclear norm minimization. Furthermore, we empirically demonstrate that the subspace norm achieves the nearly ideal denoising performance even with $H=O(1)$.
Qinqing Zheng, Ryota Tomioka
NIPS2
2015 The algebraic combinatorial approach for low-rank matrix completion
Franz J. Király, Louis Theran, Ryota Tomioka
J. Mach. Learn. Res.3
2015 Condition for perfect dimensionality recovery by variational Bayesian PCA
Shinichi Nakajima, Ryota Tomioka, Masashi Sugiyama, S. Derin Babacan
J. Mach. Learn. Res.2
2014 Early detection of persistent topics in social networks
abstract
In social networking services (SNSs), persistent topics are extremely rare and valuable. In this paper, we propose an algorithm for the detection of persistent topics in SNSs based on Topic Graph. A topic graph is a subgraph of the ordinary social network graph that consists of the users who shared a certain topic up to some time point. Based on the assumption that the time-evolutions of the topic graphs associated with a persistent and non-persistent topics are different, we propose to detect persistent topics by performing anomaly detection on the feature values extracted from the time-evolution of the topic graph. For anomaly detection, we use principal component analysis to capture the subspace spanned by normal (non-persistent) topics. We demonstrate our technique on a real data set we gathered from Twitter and show that it performs significantly better than a base-line method based on power law curve fitting and the linear influence model.
Shota Saito, Ryota Tomioka, Kenji Yamanishi
ASONAM2
2014 Multitask learning meets tensor factorization: task imputation via convex optimization
Kishan Wimalawarne, Masashi Sugiyama, Ryota Tomioka
NIPS3
2014 Discovering Emerging Topics in Social Streams via Link-Anomaly Detection
abstract
Detection of emerging topics is now receiving renewed interest motivated by the rapid growth of social networks. Conventional-term-frequency-based approaches may not be appropriate in this context, because the information exchanged in social-network posts include not only text but also images, URLs, and videos. We focus on emergence of topics signaled by social aspects of theses networks. Specifically, we focus on mentions of users--links between users that are generated dynamically (intentionally or unintentionally) through replies, mentions, and retweets. We propose a probability model of the mentioning behavior of a social network user, and propose to detect the emergence of a new topic from the anomalies measured through the model. Aggregating anomaly scores from hundreds of users, we show that we can detect emerging topics only based on the reply/mention relationships in social-network posts. We demonstrate our technique in several real data sets we gathered from Twitter. The experiments show that the proposed mention-anomaly-based approaches can detect new topics at least as early as text-anomaly-based approaches, and in some cases much earlier when the topic is poorly identified by the textual contents in posts.
Toshimitsu Takahashi, Ryota Tomioka, Kenji Yamanishi
IEEE Trans. Knowl. Data Eng.2
2013 Quantitative Prediction of Glaucomatous Visual Field Loss from Few Measurements
abstract
We propose database-aware regression methods for extrapolation from few measurements in the context of quantitative prognosis. The idea is to leverage a database of patients with similar conditions to increase the effective number of samples when we train a predictive model. Applying the proposed method to a database of glaucoma patients, we were able to predict the disease condition at a future time point significantly more accurately than the conventional patient-wise linear regression approach. In fact, our prediction was 50% more accurate than the conventional approach when three or less measurements were available and with only two measurements at least as accurate as the conventional approach with six measurements. Moreover, the proposed method can provide spatially localized prediction and also the (localized) speed of progression, which are valuable for doctors in making decisions.
Zenghan Liang, Ryota Tomioka, Hiroshi Murata, Ryo Asaoka, Kenji Yamanishi
ICDM2
2013 Non-negative Multiple Tensor Factorization
abstract
Non-negative Tensor Factorization (NTF) is a widely used technique for decomposing a non-negative value tensor into sparse and reasonably interpretable factors. However, NTF performs poorly when the tensor is extremely sparse, which is often the case with real-world data and higher-order tensors. In this paper, we propose Non-negative Multiple Tensor Factorization (NMTF), which factorizes the target tensor and auxiliary tensors simultaneously. Auxiliary data tensors compensate for the sparseness of the target data tensor. The factors of the auxiliary tensors also allow us to examine the target data from several different aspects. We experimentally confirm that NMTF performs better than NTF in terms of reconstructing the given data. Furthermore, we demonstrate that the proposed NMTF can successfully extract spatio-temporal patterns of people's daily life such as leisure, drinking, and shopping activity by analyzing several tensors extracted from online review data sets.
Koh Takeuchi 0001, Ryota Tomioka, Katsuhiko Ishiguro, Akisato Kimura, Hiroshi Sawada
ICDM2
2013 Infinite Positive Semidefinite Tensor Factorization for Source Separation of Mixture Signals
abstract
This paper presents a new class of tensor factorization called positive semidefinite tensor factorization (PSDTF) that decomposes a set of positive semidefinite (PSD) matrices into the convex combinations of fewer PSD basis matrices. PSDTF can be viewed as a natural extension of nonnegative matrix factorization. One of the main problems of PSDTF is that an appropriate number of bases should be given in advance. To solve this problem, we propose a nonparametric Bayesian model based on a gamma process that can instantiate only a limited number of necessary bases from the infinitely many bases assumed to exist. We derive a variational Bayesian algorithm for closed-form posterior inference and a multiplicative update rule for maximum-likelihood estimation. We evaluated PSDTF on both synthetic data and real music recordings to show its superiority.
Kazuyoshi Yoshii, Ryota Tomioka, Daichi Mochihashi, Masataka Goto
ICML (3)2
2013 Convex Tensor Decomposition via Structured Schatten Norm Regularization
abstract
We propose a new class of structured Schatten norms for tensors that includes two recently proposed norms (overlapped'' and "latent'') for convex-optimization-based tensor decomposition. Based on the properties of the structured Schatten norms, we mathematically analyze the performance of "latent'' approach for tensor decomposition, which was empirically found to perform better than the "overlapped'' approach in some settings. We show theoretically that this is indeed the case. In particular, when the unknown true tensor is low-rank in a specific mode, this approach performs as well as knowing the mode with the smallest rank. Along the way, we show a novel duality result for structures Schatten norms, which is also interesting in the general context of structured sparsity. We confirm through numerical simulations that our theory can precisely predict the scaling behaviour of the mean squared error. "
Ryota Tomioka, Taiji Suzuki
NIPS1
2013 Global analytic solution of fully-observed variational Bayesian matrix factorization
Shinichi Nakajima, Masashi Sugiyama, S. Derin Babacan, Ryota Tomioka
J. Mach. Learn. Res.4
2012 A Combinatorial Algebraic Approach for the Identifiability of Low-Rank Matrix Completion
Franz J. Király, Ryota Tomioka
ICML2
2012 Perfect Dimensionality Recovery by Variational Bayesian PCA
abstract
The variational Bayesian (VB) approach is one of the best tractable approximations to the Bayesian estimation, and it was demonstrated to perform well in many applications. However, its good performance was not fully understood theoretically. For example, VB sometimes produces a sparse solution, which is regarded as a practical advantage of VB, but such sparsity is hardly observed in the rigorous Bayesian estimation. In this paper, we focus on probabilistic PCA and give more theoretical insight into the empirical success of VB. More specifically, for the situation where the noise variance is unknown, we derive a sufficient condition for perfect recovery of the true PCA dimensionality in the large-scale limit when the size of an observed matrix goes to infinity. In our analysis, we obtain bounds for a noise variance estimator and simple closed-form solutions for other parameters, which themselves are actually very useful for better implementation of VB-PCA.
Shinichi Nakajima, Ryota Tomioka, Masashi Sugiyama, S. Derin Babacan
NIPS2
2012 Tensor factorization using auxiliary information
abstract
Most of the existing analysis methods for tensors (or multi-way arrays) only assume that tensors to be completed are of low rank. However, for example, when they are applied to tensor completion problems, their prediction accuracy tends to be significantly worse when only a limited number of entries are observed. In this paper, we propose to use relationships among data as auxiliary information in addition to the low-rank assumption to improve the quality of tensor decomposition. We introduce two regularization approaches using graph Laplacians induced from the relationships, one for moderately sparse cases and the other for extremely sparse cases. We also give present two kinds of iterative algorithms for approximate solutions: one based on an EM-like algorithms which is stable but not so scalable, and the other based on gradient-based optimization which is applicable to large scale datasets. Numerical experiments on tensor completion using synthetic and benchmark datasets show that the use of auxiliary information improves completion accuracy over the existing methods based only on the low-rank assumption, especially when observations are sparse.
Atsuhiro Narita, Kohei Hayashi, Ryota Tomioka, Hisashi Kashima
Data Min. Knowl. Discov.3
2011 A New Multi-task Learning Method for Personalized Activity Recognition
abstract
Personalized activity recognition usually faces the problem of data sparseness. We aim at improving accuracy of personalized activity recognition by incorporating the information from other persons. We propose a new online multi-task learning method for personalized activity recognition. The proposed online multi-task learning method automatically learns the ``transfer-factors" (similarities) among different tasks (i.e., among different persons in our case). Experiments demonstrate that the proposed method significantly outperforms existing methods. The novelty of this paper is twofold: (1) A new multi-task learning framework, which can naturally learn similarities among tasks, (2) To our knowledge, this is the first study of large-scale personalized activity recognition.
Xu Sun 0001, Hisashi Kashima, Ryota Tomioka, Naonori Ueda, Ping Li 0001
ICDM3
2011 Discovering Emerging Topics in Social Streams via Link Anomaly Detection
abstract
Detection of emerging topics are now receiving renewed interest motivated by the rapid growth of social networks. Conventional term-frequency-based approaches may not be appropriate in this context, because the information exchanged are not only texts but also images, URLs, and videos. We focus on the social aspects of theses networks. That is, the links between users that are generated dynamically intentionally or unintentionally through replies, mentions, and retweets. We propose a probability model of the mentioning behaviour of a social network user, and propose to detect the emergence of a new topic from the anomaly measured through the model. We combine the proposed mention anomaly score with a recently proposed change-point detection technique based on the Sequentially Discounting Normalized Maximum Likelihood (SDNML), or with Kleinberg's burst model. Aggregating anomaly scores from hundreds of users, we show that we can detect emerging topics only based on the reply/mention relationships in social network posts. We demonstrate our technique in a number of real data sets we gathered from Twitter. The experiments show that the proposed mention-anomaly-based approaches can detect new topics at least as early as the conventional term-frequency-based approach, and sometimes much earlier when the keyword is ill-defined.
Toshimitsu Takahashi, Ryota Tomioka, Kenji Yamanishi
ICDM2
2011 Statistical Performance of Convex Tensor Decomposition
abstract
We analyze the statistical performance of a recently proposed convex tensor decomposition algorithm. Conventionally tensor decomposition has been formulated as non-convex optimization problems, which hindered the analysis of their performance. We show under some conditions that the mean squared error of the convex method scales linearly with the quantity we call the normalized rank of the true tensor. The current analysis naturally extends the analysis of convex low-rank matrix estimation to tensors. Furthermore, we show through numerical experiments that our theory can precisely predict the scaling behaviour in practice.
Ryota Tomioka, Taiji Suzuki, Kohei Hayashi, Hisashi Kashima
NIPS1
2011 Large Scale Real-Life Action Recognition Using Conditional Random Fields with Stochastic Training
Xu Sun 0001, Hisashi Kashima, Ryota Tomioka, Naonori Ueda
PAKDD (2)3
2011 Real-Time Change-Point Detection Using Sequentially Discounting Normalized Maximum Likelihood Coding
Yasuhiro Urabe, Kenji Yamanishi, Ryota Tomioka, Hiroki Iwai
PAKDD (2)3
2011 Tensor Factorization Using Auxiliary Information
Atsuhiro Narita, Kohei Hayashi, Ryota Tomioka, Hisashi Kashima
ECML/PKDD (2)3
2011 Super-Linear Convergence of Dual Augmented Lagrangian Algorithm for Sparsity Regularized Estimation
Ryota Tomioka, Taiji Suzuki, Masashi Sugiyama
J. Mach. Learn. Res.1
2011 SpicyMKL: a fast algorithm for Multiple Kernel Learning with thousands of kernels
Taiji Suzuki, Ryota Tomioka
Mach. Learn.2
2010 A Fast Augmented Lagrangian Algorithm for Learning Low-Rank Matrices
Ryota Tomioka, Taiji Suzuki, Masashi Sugiyama, Hisashi Kashima
ICML1
2010 Global Analytic Solution for Variational Bayesian Matrix Factorization
abstract
Bayesian methods of matrix factorization (MF) have been actively explored recently as promising alternatives to classical singular value decomposition. In this paper, we show that, despite the fact that the optimization problem is non-convex, the global optimal solution of variational Bayesian (VB) MF can be computed analytically by solving a quartic equation. This is highly advantageous over a popular VBMF algorithm based on iterated conditional modes since it can only find a local optimal solution after iterations. We further show that the global optimal solution of empirical VBMF (hyperparameters are also learned from data) can also be analytically computed. We illustrate the usefulness of our results through experiments.
Shinichi Nakajima, Masashi Sugiyama, Ryota Tomioka
NIPS3
2009 Dual-Augmented Lagrangian Method for Efficient Sparse Reconstruction
abstract
We propose an efficient algorithm for sparse signal reconstruction problems. The proposed algorithm is an augmented Lagrangian method based on the dual problem. It is efficient when the number of unknown variables is much larger than the number of observations because of the dual formulation. Moreover, the primal variable is explicitly updated and the sparsity in the solution is exploited. Numerical comparison with the state-of-the-art algorithms shows that the proposed algorithm is favorable when the design matrix is poorly conditioned or dense and very large.
Ryota Tomioka, Masashi Sugiyama
IEEE Signal Process. Lett.1
2007 Classifying matrices with a spectral regularization
abstract
We propose a method for the classification of matrices. We use a linear classifier with a novel regularization scheme based on the spectral l1-norm of its coefficient matrix. The spectral regularization not only provides a principled way of complexity control but also enables automatic determination of the rank of the coefficient matrix. Using the Linear Matrix Inequality technique, we formulate the inference task as a single convex optimization problem. We apply our method to the motor-imagery EEG classification problem. The method not only improves upon conventional methods in the classification performance but also determines a subspace in the signal that concentrates discriminative information without any additional feature extraction step. The method can be easily generalized to regression problems by changing the loss function. Connections to other methods are also discussed.
Ryota Tomioka, Kazuyuki Aihara
ICML1
2007 Invariant Common Spatial Patterns: Alleviating Nonstationarities in Brain-Computer Interfacing
abstract
Brain-Computer Interfaces can suffer from a large variance of the subject condi- tions within and across sessions. For example vigilance fluctuations in the indi- vidual, variable task involvement, workload etc. alter the characteristics of EEG signals and thus challenge a stable BCI operation. In the present work we aim to define features based on a variant of the common spatial patterns (CSP) algorithm that are constructed invariant with respect to such nonstationarities. We enforce invariance properties by adding terms to the denominator of a Rayleigh coefficient representation of CSP such as disturbance covariance matrices from fluctuations in visual processing. In this manner physiological prior knowledge can be used to shape the classification engine for BCI. As a proof of concept we present a BCI classifier that is robust to changes in the level of parietal a -activity. In other words, the EEG decoding still works when there are lapses in vigilance.
Benjamin Blankertz, Motoaki Kawanabe, Ryota Tomioka, Friederike U. Hohlefeld, Vadim V. Nikulin, Klaus-Robert Müller
NIPS3
2006 Logistic Regression for Single Trial EEG Classification
abstract
We propose a novel framework for the classification of single trial ElectroEncephaloGraphy (EEG), based on regularized logistic regression. Framed in this robust statistical framework no prior feature extraction or outlier removal is required. We present two variations of parameterizing the regression function: (a) with a full rank symmetric matrix coefficient and (b) as a difference of two rank=1 matrices. In the first case, the problem is convex and the logistic regression is optimal under a generative model. The latter case is shown to be related to the Common Spatial Pattern (CSP) algorithm, which is a popular technique in Brain Computer Interfacing. The regression coefficients can also be topographically mapped onto the scalp similarly to CSP pro jections, which allows neuro-physiological interpretation. Simulations on 162 BCI datasets demonstrate that classification accuracy and robustness compares favorably against conventional CSP based classifiers.
Ryota Tomioka, Kazuyuki Aihara, Klaus-Robert Müller
NIPS1