Shinichi Nakajima

dblp:97/6115 · DBLP profile ↗
← Back
53ranked-venue papers
17as first author
16since 2021 · last 2026
0000-0003-3970-4569ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 16 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Brain Source Reconstruction by Initializing 3-D Neural Networks With Physical Inverse Solutions
abstract
Reconstructing brain sources is a fundamental challenge in neuroscience, crucial for understanding brain function and dysfunction. Electroencephalography (EEG) signals have a high temporal resolution. However, identifying the correct spatial location of brain sources from these signals remains difficult due to the ill-posed structure of the problem. Traditional methods predominantly rely on manually crafted priors, missing the flexibility of data-driven learning, while recent deep learning approaches focus on end-to-end learning, typically using the physical information of the forward model only for generating training data. We propose the novel hybrid method 3D-PIUNet for EEG source localization that effectively integrates the strengths of traditional and deep learning techniques. 3D-PIUNet starts from an initial physics-informed estimate by using the pseudo inverse to map from measurements to source space. Secondly, by viewing the brain as a 3D volume, we use a 3D convolutional U-Net to capture spatial dependencies and refine the solution according to the learned data prior. Training the model relies on simulated pseudo-realistic brain source data, covering different source distributions. Trained on this data, our model significantly improves spatial accuracy, demonstrating superior performance over both traditional and end-to-end data-driven methods. Additionally, we validate our findings with real EEG data from a visual task, where 3D-PIUNet successfully identifies the visual cortex and reconstructs the expected temporal behavior, thereby showcasing its practical applicability.
Marco Morik, Ali Hashemi 0002, Klaus-Robert Müller, Stefan Haufe, Shinichi Nakajima
IEEE Trans. Medical Imaging5
2025 Multilevel Generative Samplers for Investigating Critical Phenomena
abstract
Investigating critical phenomena or phase transitions is of high interest in physics and chemistry, for which Monte Carlo (MC) simulations, a crucial tool for numerically analyzing macroscopic properties of given systems, are often hindered by an emerging divergence of correlation length---known as scale invariance at criticality (SIC) in the renormalization group theory. SIC causes the system to behave the same at any length scale, from which many existing sampling methods suffer: long-range correlations cause critical slowing down in Markov chain Monte Carlo (MCMC), and require intractably large receptive fields for generative samplers. In this paper, we propose a Renormalization-informed Generative Critical Sampler (RiGCS)---a novel sampler specialized for near-critical systems, where SIC is leveraged as an advantage rather than a nuisance. Specifically, RiGCS builds on MultiLevel Monte Carlo (MLMC) with Heat Bath (HB) algorithms, which perform ancestral sampling from low-resolution to high-resolution lattice configurations with site wise-independent conditional HB sampling. Although MLMC-HB is highly efficient under exact SIC, it suffers from a low acceptance rate under slight SIC violation. Notably, SIC violation always occurs in finite-size systems, and may induce long-range and higher-order interactions in the renormalized distributions, which are not considered by independent HB samplers. RiGCS enhances MLMC-HB by replacing a part of the conditional HB sampler with generative models that capture those residual interactions and improve the sampling efficiency. Our experiments show that the effective sample size of RiGCS is a few orders of magnitude higher than state-of-the-art generative model baselines in sampling configurations for $128 \times 128$ two-dimensional Ising systems. SIC also allows us to adopt a specialized sequential training protocol with model transfer, which significantly accelerates training.
Ankur Singha, Elia Cellini, Kim Nicoli, Karl Jansen, Stefan Kühn, Shinichi Nakajima
ICLR6
2024 Adaptive Observation Cost Control for Variational Quantum Eigensolvers
abstract
The objective to be minimized in the variational quantum eigensolver (VQE) has a restricted form, which allows a specialized sequential minimal optimization (SMO) that requires only a few observations in each iteration. However, the SMO iteration is still costly due to the observation noise—one observation at a point typically requires averaging over hundreds to thousands of repeated quantum measurement shots for achieving a reasonable noise level. In this paper, we propose an adaptive cost control method, named subspace in confident region (SubsCoRe), for SMO. SubsCoRe uses the Gaussian process (GP) surrogate, and requires it to have low uncertainty over the subspace being updated, so that optimization in each iteration is performed with guaranteed accuracy. Adaptive cost control is performed by setting the required accuracy according to the progress of the optimization, and identifying the minimum number of measurement shots, as well as their distribution, satisfying the SubsCoRe requirement.
Christopher J. Anders, Kim Nicoli, Bingting Wu, Naima Elosegui, Samuele Pedrielli, Lena Funcke, Karl Jansen, Stefan Kühn, Shinichi Nakajima
ICML9
2024 Federated Learning over Connected Modes
abstract
Statistical heterogeneity in federated learning poses two major challenges: slow global training due to conflicting gradient signals, and the need of personalization for local distributions. In this work, we tackle both challenges by leveraging recent advances in \emph{linear mode connectivity} --- identifying a linearly connected low-loss region in the parameter space of neural networks, which we call solution simplex. We propose federated learning over connected modes (\textsc{Floco}), where clients are assigned local subregions in this simplex based on their gradient signals, and together learn the shared global solution simplex. This allows personalization of the client models to fit their local distributions within the degrees of freedom in the solution simplex and homogenizes the update signals for the global simplex training. Our experiments show that \textsc{Floco} accelerates the global training process, and significantly improves the local accuracy with minimal computational overhead in cross-silo federated learning settings.
Dennis Grinwald, Philipp Wiesner, Shinichi Nakajima
NeurIPS3
2024 Generative Fractional Diffusion Models
abstract
We introduce the first continuous-time score-based generative model that leverages fractional diffusion processes for its underlying dynamics. Although diffusion models have excelled at capturing data distributions, they still suffer from various limitations such as slow convergence, mode-collapse on imbalanced data, and lack of diversity. These issues are partially linked to the use of light-tailed Brownian motion (BM) with independent increments. In this paper, we replace BM with an approximation of its non-Markovian counterpart, fractional Brownian motion (fBM), characterized by correlated increments and Hurst index $H \in (0,1)$, where $H=0.5$ recovers the classical BM. To ensure tractable inference and learning, we employ a recently popularized Markov approximation of fBM (MA-fBM) and derive its reverse-time model, resulting in *generative fractional diffusion models* (GFDM). We characterize the forward dynamics using a continuous reparameterization trick and propose *augmented score matching* to efficiently learn the score function, which is partly known in closed form, at minimal added cost. The ability to drive our diffusion model via MA-fBM offers flexibility and control. $H \leq 0.5$ enters the regime of *rough paths* whereas $H>0.5$ regularizes diffusion paths and invokes long-term memory. The Markov approximation allows added control by varying the number of Markov processes linearly combined to approximate fBM. Our evaluations on real image datasets demonstrate that GFDM achieves greater pixel-wise diversity and enhanced image quality, as indicated by a lower FID, offering a promising alternative to traditional diffusion models
Gabriel Nobis, Maximilian Springenberg, Marco Aversa, Michael Detzel, Rembert Daems, Roderick Murray-Smith, Shinichi Nakajima, Sebastian Lapuschkin, Stefano Ermon, Tolga Birdal, Manfred Opper, Christoph Knochenhauer, Luis Oala, Wojciech Samek
NeurIPS7
2023 Relevant Walk Search for Explaining Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have become important machine learning tools for graph analysis, and its explainability is crucial for safety, fairness, and robustness. Layer-wise relevance propagation for GNNs (GNN-LRP) evaluates the relevance of walks to reveal important information flows in the network, and provides higher-order explanations, which have been shown to be superior to the lower-order, i.e., node-/edge-level, explanations. However, identifying relevant walks by GNN-LRP requires exponential computational complexity with respect to the network depth, which we will remedy in this paper. Specifically, we propose polynomial-time algorithms for finding top-$K$ relevant walks, which drastically reduces the computation and thus increases the applicability of GNN-LRP to large-scale problems. Our proposed algorithms are based on the max-product algorithm—a common tool for finding the maximum likelihood configurations in probabilistic graphical models—and can find the most relevant walks exactly at the neuron level and approximately at the node level. Our experiments demonstrate the performance of our algorithms at scale and their utility across application domains, i.e., on epidemiology, molecular, and natural language benchmarks. We provide our codes under github.com/xiong-ping/rel_walk_gnnlrp.
Ping Xiong 0002, Thomas Schnake, Michael Gastegger, Grégoire Montavon, Klaus-Robert Müller, Shinichi Nakajima
ICML6
2023 Labeling Neural Representations with Inverse Recognition
abstract
Deep Neural Networks (DNNs) demonstrate remarkable capabilities in learning complex hierarchical data representations, but the nature of these representations remains largely unknown. Existing global explainability methods, such as Network Dissection, face limitations such as reliance on segmentation masks, lack of statistical significance testing, and high computational demands. We propose Inverse Recognition (INVERT), a scalable approach for connecting learned representations with human-understandable concepts by leveraging their capacity to discriminate between these concepts. In contrast to prior work, INVERT is capable of handling diverse types of neurons, exhibits less computational complexity, and does not rely on the availability of segmentation masks. Moreover, INVERT provides an interpretable metric assessing the alignment between the representation and its corresponding explanation and delivering a measure of statistical significance. We demonstrate the applicability of INVERT in various scenarios, including the identification of representations affected by spurious correlations, and the interpretation of the hierarchical structure of decision-making within the models.
Kirill Bykov, Laura Kopf, Shinichi Nakajima, Marius Kloft, Marina M.-C. Höhne
NeurIPS3
2023 Physics-Informed Bayesian Optimization of Variational Quantum Circuits
abstract
In this paper, we propose a novel and powerful method to harness Bayesian optimization for variational quantum eigensolvers (VQEs) - a hybrid quantum-classical protocol used to approximate the ground state of a quantum Hamiltonian. Specifically, we derive a *VQE-kernel* which incorporates important prior information about quantum circuits: the kernel feature map of the VQE-kernel exactly matches the known functional form of the VQE's objective function and thereby significantly reduces the posterior uncertainty. Moreover, we propose a novel acquisition function for Bayesian optimization called \emph{Expected Maximum Improvement over Confident Regions} (EMICoRe) which can actively exploit the inductive bias of the VQE-kernel by treating regions with low predictive uncertainty as indirectly "observed". As a result, observations at as few as three points in the search domain are sufficient to determine the complete objective function along an entire one-dimensional subspace of the optimization landscape. Our numerical experiments demonstrate that our approach improves over state-of-the-art baselines.
Kim Nicoli, Christopher J. Anders, Lena Funcke, Tobias Hartung, Karl Jansen, Stefan Kühn, Klaus-Robert Müller, Paolo Stornati, Pan Kessel, Shinichi Nakajima
NeurIPS10
2023 Domain-Specific Word Embeddings with Structure Prediction
abstract
Abstract Complementary to finding good general word embeddings, an important question for representation learning is to find dynamic word embeddings, for example, across time or domain. Current methods do not offer a way to use or predict information on structure between sub-corpora, time or domain and dynamic embeddings can only be compared after post-alignment. We propose novel word embedding methods that provide general word representations for the whole corpus, domain- specific representations for each sub-corpus, sub-corpus structure, and embedding alignment simultaneously. We present an empirical evaluation on New York Times articles and two English Wikipedia datasets with articles on science and philosophy. Our method, called Word2Vec with Structure Prediction (W2VPred), provides better performance than baselines in terms of the general analogy tests, domain-specific analogy tests, and multiple specific word embedding evaluations as well as structure prediction performance when no structure is given a priori. As a use case in the field of Digital Humanities we demonstrate how to raise novel research questions for high literature from the German Text Archive.
David Lassner, Stephanie Brandl, Anne Baillot, Shinichi Nakajima
Trans. Assoc. Comput. Linguistics4
2023 Langevin Cooling for Unsupervised Domain Translation
abstract
Domain translation is the task of finding correspondence between two domains. Several deep neural network (DNN) models, e.g., CycleGAN and cross-lingual language models, have shown remarkable successes on this task under the unsupervised setting-the mappings between the domains are learned from two independent sets of training data in both domains (without paired samples). However, those methods typically do not perform well on a significant proportion of test samples. In this article, we hypothesize that many of such unsuccessful samples lie at the fringe-relatively low-density areas-of data distribution, where the DNN was not trained very well, and propose to perform the Langevin dynamics to bring such fringe samples toward high-density areas. We demonstrate qualitatively and quantitatively that our strategy, called Langevin cooling (L-Cool), enhances state-of-the-art methods in image translation and language translation tasks.
Vignesh Srinivasan, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima
IEEE Trans. Neural Networks Learn. Syst.4
2022 NoiseGrad - Enhancing Explanations by Introducing Stochasticity to Model Weights
abstract
Many efforts have been made for revealing the decision-making process of black-box learning machines such as deep neural networks, resulting in useful local and global explanation methods. For local explanation, stochasticity is known to help: a simple method, called SmoothGrad, has improved the visual quality of gradient-based attribution by adding noise to the input space and averaging the explanations of the noisy inputs. In this paper, we extend this idea and propose NoiseGrad that enhances both local and global explanation methods. Specifically, NoiseGrad introduces stochasticity in the weight parameter space, such that the decision boundary is perturbed. NoiseGrad is expected to enhance the local explanation, similarly to SmoothGrad, due to the dual relationship between the input perturbation and the decision boundary perturbation. We evaluate NoiseGrad and its fusion with SmoothGrad - FusionGrad - qualitatively and quantitatively with several evaluation criteria, and show that our novel approach significantly outperforms the baseline methods. Both NoiseGrad and FusionGrad are method-agnostic and as handy as SmoothGrad using a simple heuristic for the choice of the hyperparameter setting without the need of fine-tuning.
Kirill Bykov, Anna Hedström, Shinichi Nakajima, Marina M.-C. Höhne
AAAI3
2022 Path-Gradient Estimators for Continuous Normalizing Flows
abstract
Recent work has established a path-gradient estimator for simple variational Gaussian distributions and has argued that the path-gradient is particularly beneficial in the regime in which the variational distribution approaches the exact target distribution. In many applications, this regime can however not be reached by a simple Gaussian variational distribution. In this work, we overcome this crucial limitation by proposing a path-gradient estimator for the considerably more expressive variational family of continuous normalizing flows. We outline an efficient algorithm to calculate this estimator and establish its superior performance empirically.
Lorenz Vaitl, Kim Nicoli, Shinichi Nakajima, Pan Kessel
ICML3
2022 Efficient Computation of Higher-Order Subgraph Attribution via Message Passing
abstract
Explaining graph neural networks (GNNs) has become more and more important recently. Higher-order interpretation schemes, such as GNN-LRP (layer-wise relevance propagation for GNN), emerged as powerful tools for unraveling how different features interact thereby contributing to explaining GNNs. GNN-LRP gives a relevance attribution of walks between nodes at each layer, and the subgraph attribution is expressed as a sum over exponentially many such walks. In this work, we demonstrate that such exponential complexity can be avoided. In particular, we propose novel algorithms that enable to attribute subgraphs with GNN-LRP in linear-time (w.r.t. the network depth). Our algorithms are derived via message passing techniques that make use of the distributive property, thereby directly computing quantities for higher-order explanations. We further adapt our efficient algorithms to compute a generalization of subgraph attributions that also takes into account the neighboring graph features. Experimental results show the significant acceleration of the proposed algorithms and demonstrate the high usefulness and scalability of our novel generalized subgraph attribution method.
Ping Xiong 0002, Thomas Schnake, Grégoire Montavon, Klaus-Robert Müller, Shinichi Nakajima
ICML5
2022 Higher-Order Explanations of Graph Neural Networks via Relevant Walks
abstract
Graph Neural Networks (GNNs) are a popular approach for predicting graph structured data. As GNNs tightly entangle the input graph into the neural network structure, common explainable AI approaches are not applicable. To a large extent, GNNs have remained black-boxes for the user so far. In this paper, we show that GNNs can in fact be naturally explained using higher-order expansions, i.e., by identifying groups of edges that jointly contribute to the prediction. Practically, we find that such explanations can be extracted using a nested attribution scheme, where existing techniques such as layer-wise relevance propagation (LRP) can be applied at each step. The output is a collection of walks into the input graph that are relevant for the prediction. Our novel explanation method, which we denote by GNN-LRP, is applicable to a broad range of graph neural networks and lets us extract practically relevant insights on sentiment analysis of text data, structure-property relationships in quantum chemistry, and image classification.
Thomas Schnake, Oliver Eberle, Jonas Lederer, Shinichi Nakajima, Kristof Schütt, Klaus-Robert Müller, Grégoire Montavon
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 ML-Based QoE Estimation in 5G Networks Using Different Regression Techniques
abstract
Monitoring and providing customers with a satisfying Quality of Experience (QoE) is a crucial business incentive for mobile network operators (MNOs). While the MNO is capable of monitoring a vast amount of network-related key performance indicators (KPIs), it typically does not have access to application-specific performance metrics. Among others, this is due to practical obstacles, such as missing standardized interfaces between the network and the application. Existing QoE models allow to map collected KPIs to the user-perceived quality. However, they are not dynamic, cumbersome to obtain, and often rely on application-level information, such as the stalling duration in the case of video streaming. The 5G networking architecture provides new features which can potentially overcome current limitations of in-network QoE monitoring. More specifically, the Application Function (AF) provides a standardized interface for communicating between 5G systems and third parties, such as application providers. The Network Data Analytics Function (NWDAF) is capable of collecting a vast number of network statistics from other 5G network functions and is dedicated to training and deploying Machine Learning (ML) models. This opens new possibilities, unimaginable for earlier mobile network generations, to dynamically learn the relationship between network KPIs and QoE by utilizing ML. Besides elaborating on how the new capabilities introduced with 5G can support an ML-based QoE estimation, we perform a simulation-based feasibility study which evaluates the estimation accuracy of different state-of-the-art regression techniques. In addition, we discuss them with respect to various qualitative aspects from an MNO’s point of view.
Susanna Schwarzmann, Clarissa Cassales Marquezan, Riccardo Trivisonno, Shinichi Nakajima, Vincent Barriac, Thomas Zinner
IEEE Trans. Netw. Serv. Manag.4
2021 Robustifying models against adversarial attacks by Langevin dynamics
Vignesh Srinivasan, Csaba Rohrer, Arturo Marbán, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima
Neural Networks6
2020 Benign Examples: Imperceptible Changes Can Enhance Image Translation Performance
Vignesh Srinivasan, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima
AAAI4
2020 Accuracy vs. Cost Trade-off for Machine Learning Based QoE Estimation in 5G Networks
abstract
Since their first release, 5G systems have been enhanced with Network Data Analytics Functionalities (NWDAF) as well as with the ability to interact with 3rd parties' Application Functions (AFs). Such capabilities enable a variety of potentials, unimaginable for earlier generation networks, notable examples being 5G built-in Machine Learning (ML) mechanisms for QoE estimation, subject of this paper. In this work, an ML-based mechanism for video streaming QoE estimation in 5G networks is presented and evaluated. The mechanism relies on an ML algorithm embedded in NWDAF, the collection of 5G network KPIs, and the collection of QoE information from video streaming service provider, i.e., the 3rd party AF. The mechanism has been evaluated in terms of QoE estimation accuracy against the cost in terms of required input sources and data for the estimation, and its performance has been compared to alternative methodologies not making use of ML. The evaluation, via simulation activity, clearly highlights the benefits of the proposed mechanism. Based on the derived results, the required input sources are ranked with respect to their importance.
Susanna Schwarzmann, Clarissa Cassales Marquezan, Riccardo Trivisonno, Shinichi Nakajima, Thomas Zinner
ICC4
2020 Towards Best Practice in Explaining Neural Network Decisions with LRP
abstract
Within the last decade, neural network based predictors have demonstrated impressive - and at times superhuman - capabilities. This performance is often paid for with an intransparent prediction process and thus has sparked numerous contributions in the novel field of explainable artificial intelligence (XAI). In this paper, we focus on a popular and widely used method of XAI, the Layer-wise Relevance Propagation (LRP). Since its initial proposition LRP has evolved as a method, and a best practice for applying the method has tacitly emerged, based however on humanly observed evidence alone. In this paper we investigate - and for the first time quantify - the effect of this current best practice on feedforward neural networks in a visual object detection setting. The results verify that the layer-dependent approach to LRP applied in recent literature better represents the model's reasoning, and at the same time increases the object localization and class discriminativity of LRP.
Maximilian Kohlbrenner, Alexander Bauer 0001, Shinichi Nakajima, Alexander Binder, Wojciech Samek, Sebastian Lapuschkin
IJCNN3
2020 Optimizing for Measure of Performance in Max-Margin Parsing
abstract
Many learning tasks in the field of natural language processing including sequence tagging, sequence segmentation, and syntactic parsing have been successfully approached by means of structured prediction methods. An appealing property of the corresponding training algorithms is their ability to integrate the loss function of interest into the optimization process improving the final results according to the chosen measure of performance. Here, we focus on the task of constituency parsing and show how to optimize the model for the F1-score in the max-margin framework of a structural support vector machine (SVM). For reasons of computational efficiency, it is a common approach to binarize the corresponding grammar before training. Unfortunately, this introduces a bias during the training procedure as the corresponding loss function is evaluated on the binary representation, while the resulting performance is measured on the original unbinarized trees. Here, we address this problem by extending the inference procedure presented by Bauer et al. Specifically, we propose an algorithmic modification that allows evaluating the loss on the unbinarized trees. The new approach properly models the loss function of interest resulting in better prediction accuracy and still benefits from the computational efficiency due to binarized representation. The presented idea can be easily transferred to other structured loss functions.
Alexander Bauer 0001, Shinichi Nakajima, Nico Görnitz, Klaus-Robert Müller
IEEE Trans. Neural Networks Learn. Syst.2
2019 Partial Optimality of Dual Decomposition for MAP Inference in Pairwise MRFs
abstract
Markov random fields (MRFs) are a powerful tool for modelling statistical dependencies for a set of random variables using a graphical representation. An important computational problem related to MRFs, called maximum a posteriori (MAP) inference, is finding a joint variable assignment with the maximal probability. It is well known that the two popular optimisation techniques for this task, linear programming (LP) relaxation and dual decomposition (DD), have a strong connection both providing an optimal solution to the MAP problem when a corresponding LP relaxation is tight. However, less is known about their relationship in the opposite and more realistic case. In this paper, we explain how the fully integral assignments obtained via DD partially agree with the optimal fractional assignments via LP relaxation when the latter is not tight. In particular, for binary pairwise MRFs the corresponding result suggests that both methods share the partial optimality property of their solutions.
Alexander Bauer 0001, Shinichi Nakajima, Nico Görnitz, Klaus-Robert Müller
AISTATS2
2018 Support Vector Data Descriptions and k-Means Clustering: One Class?
abstract
We present ClusterSVDD, a methodology that unifies support vector data descriptions (SVDDs) and $k$ -means clustering into a single formulation. This allows both methods to benefit from one another, i.e., by adding flexibility using multiple spheres for SVDDs and increasing anomaly resistance and flexibility through kernels to $k$ -means. In particular, our approach leads to a new interpretation of $k$ -means as a regularized mode seeking algorithm. The unifying formulation further allows for deriving new algorithms by transferring knowledge from one-class learning settings to clustering settings and vice versa. As a showcase, we derive a clustering method for structured data based on a one-class learning scenario. Additionally, our formulation can be solved via a particularly simple optimization scheme. We evaluate our approach empirically to highlight some of the proposed benefits on artificially generated data, as well as on real-world problems, and provide a Python software package comprising various implementations of primal and dual SVDD as well as our proposed ClusterSVDD.
Nico Görnitz, Luiz Alberto Lima, Klaus-Robert Müller, Marius Kloft, Shinichi Nakajima
IEEE Trans. Neural Networks Learn. Syst.5
2018 Transductive Regression for Data With Latent Dependence Structure
abstract
Analyzing data with latent spatial and/or temporal structure is a challenge for machine learning. In this paper, we propose a novel nonlinear model for studying data with latent dependence structure. It successfully combines the concepts of Markov random fields, transductive learning, and regression, making heavy use of the notion of joint feature maps. Our transductive conditional random field regression model is able to infer the latent states by combining limited labeled data of high precision with unlabeled data containing measurement uncertainty. In this manner, we can propagate accurate information and greatly reduce uncertainty. We demonstrate the usefulness of our novel framework on generated time series data with the known temporal structure and successfully validate it on synthetic as well as real-world offshore data with the spatial structure from the oil industry to predict rock porosities from acoustic impedance data.
Nico Görnitz, Luiz Alberto Lima, Luiz Eduardo Varella, Klaus-Robert Müller, Shinichi Nakajima
IEEE Trans. Neural Networks Learn. Syst.5
2017 Minimizing Trust Leaks for Robust Sybil Detection
abstract
Sybil detection is a crucial task to protect online social networks (OSNs) against intruders who try to manipulate automatic services provided by OSNs to their customers. In this paper, we first discuss the robustness of graph-based Sybil detectors SybilRank and Integro and refine theoretically their security guarantees towards more realistic assumptions. After that, we formally introduce adversarial settings for the graph-based Sybil detection problem and derive a corresponding optimal attacking strategy by exploitation of trust leaks. Based on our analysis, we propose transductive Sybil ranking (TSR), a robust extension to SybilRank and Integro that directly minimizes trust leaks. Our empirical evaluation shows significant advantages of TSR over state-of-the-art competitors on a variety of attacking scenarios on artificially generated data and real-world datasets.
János Höner, Shinichi Nakajima, Alexander Bauer 0001, Klaus-Robert Müller, Nico Görnitz
ICML2
2017 Sparse probit linear mixed model
Stephan Mandt, Florian Wenzel, Shinichi Nakajima, John P. Cunningham, Christoph Lippert, Marius Kloft
Mach. Learn.3
2017 Efficient Exact Inference With Loss Augmented Objective in Structured Learning
abstract
Structural support vector machine (SVM) is an elegant approach for building complex and accurate models with structured outputs. However, its applicability relies on the availability of efficient inference algorithms--the state-of-the-art training algorithms repeatedly perform inference to compute a subgradient or to find the most violating configuration. In this paper, we propose an exact inference algorithm for maximizing nondecomposable objectives due to special type of a high-order potential having a decomposable internal structure. As an important application, our method covers the loss augmented inference, which enables the slack and margin scaling formulations of structural SVM with a variety of dissimilarity measures, e.g., Hamming loss, precision and recall, Fβ-loss, intersection over union, and many other functions that can be efficiently computed from the contingency table. We demonstrate the advantages of our approach in natural language parsing and sequence segmentation applications.
Alexander Bauer 0001, Shinichi Nakajima, Klaus-Robert Müller
IEEE Trans. Neural Networks Learn. Syst.2
2015 Condition for perfect dimensionality recovery by variational Bayesian PCA
Shinichi Nakajima, Ryota Tomioka, Masashi Sugiyama, S. Derin Babacan
J. Mach. Learn. Res.1
2014 Analysis of Empirical MAP and Empirical Partially Bayes: Can They be Alternatives to Variational Bayes?
abstract
Variational Bayesian (VB) learning is known to be a promising approximation to Bayesian learning with computational efficiency. However, in some applications, e.g., large-scale collaborative filtering and tensor factorization, VB is still computationally too costly. In such cases, looser approximations such as MAP estimation and partially Bayesian (PB) learning, where a part of the parameters are point-estimated, seem attractive. In this paper, we theoretically investigate the behavior of the MAP and the PB solutions of matrix factorization. A notable finding is that the global solutions of MAP and PB in the empirical Bayesian scenario, where the hyperparameters are also estimated from observation, are trivial and useless, while their local solutions behave similarly to the global solution of VB. This suggests that empirical MAP and empirical PB with local search can be alternatives to empirical VB equipped with the useful automatic relevance determination property. Experiments support our theory.
Shinichi Nakajima, Masashi Sugiyama
AISTATS1
2014 Analysis of Variational Bayesian Latent Dirichlet Allocation: Weaker Sparsity Than MAP
Shinichi Nakajima, Issei Sato, Masashi Sugiyama, Kazuho Watanabe, Hiroko Kobayashi
NIPS1
2013 Global Solver and Its Efficient Approximation for Variational Bayesian Low-rank Subspace Clustering
abstract
When a probabilistic model and its prior are given, Bayesian learning offers inference with automatic parameter tuning. However, Bayesian learning is often obstructed by computational difficulty: the rigorous Bayesian learning is intractable in many models, and its variational Bayesian (VB) approximation is prone to suffer from local minima. In this paper, we overcome this difficulty for low-rank subspace clustering (LRSC) by providing an exact global solver and its efficient approximation. LRSC extracts a low-dimensional structure of data by embedding samples into the union of low-dimensional subspaces, and its variational Bayesian variant has shown good performance. We first prove a key property that the VB-LRSC model is highly redundant. Thanks to this property, the optimization problem of VB-LRSC can be separated into small subproblems, each of which has only a small number of unknown variables. Our exact global solver relies on another key property that the stationary condition of each subproblem is written as a set of polynomial equations, which is solvable with the homotopy method. For further computational efficiency, we also propose an efficient approximate variant, of which the stationary condition can be written as a polynomial equation with a single variable. Experimental results show the usefulness of our approach.
Shinichi Nakajima, Akiko Takeda, S. Derin Babacan, Masashi Sugiyama, Ichiro Takeuchi
NIPS1
2013 Parametric Task Learning
abstract
We introduce a novel formulation of multi-task learning (MTL) called parametric task learning (PTL) that can systematically handle infinitely many tasks parameterized by a continuous parameter. Our key finding is that, for a certain class of PTL problems, the path of optimal task-wise solutions can be represented as piecewise-linear functions of the continuous task parameter. Based on this fact, we employ a parametric programming technique to obtain the common shared representation across all the continuously parameterized tasks efficiently. We show that our PTL formulation is useful in various scenarios such as learning under non-stationarity, cost-sensitive learning, and quantile regression, and demonstrate the usefulness of the proposed method experimentally in these scenarios.
Ichiro Takeuchi, Tatsuya Hongo, Masashi Sugiyama, Shinichi Nakajima
NIPS4
2013 Global analytic solution of fully-observed variational Bayesian matrix factorization
Shinichi Nakajima, Masashi Sugiyama, S. Derin Babacan, Ryota Tomioka
J. Mach. Learn. Res.1
2013 Variational Bayesian sparse additive matrix factorization
Shinichi Nakajima, Masashi Sugiyama, S. Derin Babacan
Mach. Learn.1
2012 Probabilistic Low-Rank Subspace Clustering
abstract
In this paper, we consider the problem of clustering data points into low-dimensional subspaces in the presence of outliers. We pose the problem using a density estimation formulation with an associated generative model. Based on this probability model, we first develop an iterative expectation-maximization (EM) algorithm and then derive its global solution. In addition, we develop two Bayesian methods based on variational Bayesian (VB) approximation, which are capable of automatic dimensionality selection. While the first method is based on an alternating optimization scheme for all unknowns, the second method makes use of recent results in VB matrix factorization leading to fast and effective estimation. Both methods are extended to handle sparse outliers for robustness and can handle missing values. Experimental results suggest that proposed methods are very effective in clustering and identifying outliers.
S. Derin Babacan, Shinichi Nakajima, Minh N. Do
NIPS2
2012 Perfect Dimensionality Recovery by Variational Bayesian PCA
abstract
The variational Bayesian (VB) approach is one of the best tractable approximations to the Bayesian estimation, and it was demonstrated to perform well in many applications. However, its good performance was not fully understood theoretically. For example, VB sometimes produces a sparse solution, which is regarded as a practical advantage of VB, but such sparsity is hardly observed in the rigorous Bayesian estimation. In this paper, we focus on probabilistic PCA and give more theoretical insight into the empirical success of VB. More specifically, for the situation where the noise variance is unknown, we derive a sufficient condition for perfect recovery of the true PCA dimensionality in the large-scale limit when the size of an observed matrix goes to infinity. In our analysis, we obtain bounds for a noise variance estimator and simple closed-form solutions for other parameters, which themselves are actually very useful for better implementation of VB-PCA.
Shinichi Nakajima, Ryota Tomioka, Masashi Sugiyama, S. Derin Babacan
NIPS1
2011 On Bayesian PCA: Automatic Dimensionality Selection and Analytic Solution
Shinichi Nakajima, Masashi Sugiyama, S. Derin Babacan
ICML1
2011 Global Solution of Fully-Observed Variational Bayesian Matrix Factorization is Column-Wise Independent
abstract
Variational Bayesian matrix factorization (VBMF) efficiently approximates the posterior distribution of factorized matrices by assuming matrix-wise independence of the two factors. A recent study on fully-observed VBMF showed that, under a stronger assumption that the two factorized matrices are column-wise independent, the global optimal solution can be analytically computed. However, it was not clear how restrictive the column-wise independence assumption is. In this paper, we prove that the global solution under matrix-wise independence is actually column-wise independent, implying that the column-wise independence assumption is harmless. A practical consequence of our theoretical finding is that the global solution under matrix-wise independence (which is a standard setup) can be obtained analytically in a computationally very efficient way without any iterative algorithms. We experimentally illustrate advantages of using our analytic solution in probabilistic principal component analysis.
Shinichi Nakajima, Masashi Sugiyama, S. Derin Babacan
NIPS1
2011 Theoretical Analysis of Bayesian Matrix Factorization
Shinichi Nakajima, Masashi Sugiyama
J. Mach. Learn. Res.1
2010 Implicit Regularization in Variational Bayesian Matrix Factorization
Shinichi Nakajima, Masashi Sugiyama
ICML1
2010 Global Analytic Solution for Variational Bayesian Matrix Factorization
abstract
Bayesian methods of matrix factorization (MF) have been actively explored recently as promising alternatives to classical singular value decomposition. In this paper, we show that, despite the fact that the optimization problem is non-convex, the global optimal solution of variational Bayesian (VB) MF can be computed analytically by solving a quartic equation. This is highly advantageous over a popular VBMF algorithm based on iterated conditional modes since it can only find a local optimal solution after iterations. We further show that the global optimal solution of empirical VBMF (hyperparameters are also learned from data) can also be analytically computed. We illustrate the usefulness of our results through experiments.
Shinichi Nakajima, Masashi Sugiyama, Ryota Tomioka
NIPS1
2010 Semi-supervised local Fisher discriminant analysis for dimensionality reduction
Masashi Sugiyama, Tsuyoshi Idé, Shinichi Nakajima, Jun Sese
Mach. Learn.3
2009 Multi-class image segmentation using conditional random fields and global classification
abstract
A key aspect of semantic image segmentation is to integrate local and global features for the prediction of local segment labels. We present an approach to multi-class segmentation which combines two methods for this integration: a Conditional Random Field (CRF) which couples to local image features and an image classification method which considers global features. The CRF follows the approach of Reynolds & Murphy (2007) and is based on an unsupervised multi scale pre-segmentation of the image into patches, where patch labels correspond to the random variables of the CRF. The output of the classifier is used to constraint this CRF. We demonstrate and compare the approach on a standard semantic segmentation data set.
Nils Plath, Marc Toussaint, Shinichi Nakajima
ICML3
2009 Analysis of Variational Bayesian Matrix Factorization
Shinichi Nakajima, Masashi Sugiyama
PAKDD1
2009 Feature Selection for Density Level-Sets
Marius Kloft, Shinichi Nakajima, Ulf Brefeld
ECML/PKDD (1)2
2009 Pool-based active learning in approximate linear regression
Masashi Sugiyama, Shinichi Nakajima
Mach. Learn.2
2008 Semi-Supervised Local Fisher Discriminant Analysis for Dimensionality Reduction
Masashi Sugiyama, Tsuyoshi Idé, Shinichi Nakajima, Jun Sese
PAKDD3
2008 Pool-Based Agnostic Experiment Design in Linear Regression
Masashi Sugiyama, Shinichi Nakajima
ECML/PKDD (2)2
2007 Generalization Error of Automatic Relevance Determination
Shinichi Nakajima, Sumio Watanabe
ICANN (1)1
2007 Direct Importance Estimation with Model Selection and Its Application to Covariate Shift Adaptation
abstract
When training and test samples follow different input distributions (i.e., the situation called \emph{covariate shift}), the maximum likelihood estimator is known to lose its consistency. For regaining consistency, the log-likelihood terms need to be weighted according to the \emph{importance} (i.e., the ratio of test and training input densities). Thus, accurately estimating the importance is one of the key tasks in covariate shift adaptation. A naive approach is to first estimate training and test input densities and then estimate the importance by the ratio of the density estimates. However, since density estimation is a hard problem, this approach tends to perform poorly especially in high dimensional cases. In this paper, we propose a direct importance estimation method that does not require the input density estimates. Our method is equipped with a natural model selection procedure so tuning parameters such as the kernel width can be objectively optimized. This is an advantage over a recently developed method of direct importance estimation. Simulations illustrate the usefulness of our approach.
Masashi Sugiyama, Shinichi Nakajima, Hisashi Kashima, Paul von Bünau, Motoaki Kawanabe
NIPS2
2007 Variational Bayes Solution of Linear Neural Networks and Its Generalization Performance
abstract
It is well known that in unidentifiable models, the Bayes estimation provides much better generalization performance than the maximum likelihood (ML) estimation. However, its accurate approximation by Markov chain Monte Carlo methods requires huge computational costs. As an alternative, a tractable approximation method, called the variational Bayes (VB) approach, has recently been proposed and has been attracting attention. Its advantage over the expectation maximization (EM) algorithm, often used for realizing the ML estimation, has been experimentally shown in many applications; nevertheless, it has not yet been theoretically shown. In this letter, through analysis of the simplest unidentifiable models, we theoretically show some properties of the VB approach. We first prove that in three-layer linear neural networks, the VB approach is asymptotically equivalent to a positive-part James-Stein type shrinkage estimation. Then we theoretically clarify its free energy, generalization error, and training error. Comparing them with those of the ML estimation and the Bayes estimation, we discuss the advantage of the VB approach. We also show that unlike in the Bayes estimation, the free energy and the generalization error are less simply related with each other and that in typical cases, the VB free energy well approximates the Bayes one, while the VB generalization error significantly differs from the Bayes one.
Shinichi Nakajima, Sumio Watanabe
Neural Comput.1
2006 Analytic Solution of Hierarchical Variational Bayes in Linear Inverse Problem
Shinichi Nakajima, Sumio Watanabe
ICANN (2)1
2006 Localized Bayes Estimation for Non-identifiable Models
Shingo Takamatsu, Shinichi Nakajima, Sumio Watanabe
ICONIP (1)2
2005 Generalization Error of Linear Neural Networks in an Empirical Bayes Approach
Shinichi Nakajima, Sumio Watanabe
IJCAI1