EDBT 2026 Demo / reviewers in the wild / expert
Hideitsu Hino
dblp:49/5462
· DBLP profile ↗
66ranked-venue papers
18as first author
19since 2021 · last 2026
0000-0002-6405-4361ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 15 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Certified Robustness in Few-Shot Classification with Contrastive Loss and Defensive Noise in Fine-Tuning
Hiroya Kato, Seira Hidano, Takao Murakami, Hideitsu Hino |
ICISSP (2) | 4 |
| 2025 | A Family of Distributions of Random Subsets for Controlling Positive and Negative DependenceabstractPositive and negative dependence are fundamental concepts that characterize the attractive and repulsive behavior of random subsets. Although some probabilistic models are known to exhibit positive or negative dependence, it is challenging to seamlessly bridge them with a practicable probabilistic model. In this study, we introduce a new family of distributions, named the discrete kernel point process (DKPP), which includes determinantal point processes and parts of Boltzmann machines. We also develop some computational methods for probabilistic operations and inference with DKPPs, such as calculating marginal and conditional probabilities and learning the parameters. Our numerical experiments demonstrate the controllability of positive and negative dependence and the effectiveness of the computational methods for DKPPs. Takahiro Kawashima, Hideitsu Hino |
AISTATS | 2 |
| 2025 | Difference-of-submodular Bregman DivergenceabstractThe Bregman divergence, which is generated from a convex function, is commonly used as a pseudo-distance for comparing vectors or functions in continuous spaces. In contrast, defining an analog of the Bregman divergence for discrete spaces is nontrivial. Iyer & Bilmes (2012b) considered Bregman divergences on discrete domains using submodular functions as generating functions, the discrete analogs of convex functions. In this paper, we further generalize this framework to cases where the generating function is neither submodular nor supermodular, thus increasing the flexibility and representational capacity of the resulting divergence, which we term the difference-of-submodular Bregman divergence. Additionally, we introduce a learnable form of this divergence using permutation-invariant neural networks (NNs) and demonstrate through experiments that it effectively captures key structural properties in discrete data. As a result, the proposed method significantly improves the performance of existing methods on tasks such as clustering and set retrieval problems. This work addresses the challenge of defining meaningful divergences in discrete settings and provides a new tool for tasks requiring structure-preserving distance measures. Masanari Kimura, Takahiro Kawashima, Tasuku Soma, Hideitsu Hino |
ICLR | 4 |
| 2025 | Scalable Sobolev IPM for Probability Measures on a GraphabstractWe investigate the Sobolev IPM problem for probability measures supported on a graph metric space. Sobolev IPM is an important instance of integral probability metrics (IPM), and is obtained by constraining a critic function within a unit ball defined by the Sobolev norm. In particular, it has been used to compare probability measures and is crucial for several theoretical works in machine learning. However, to our knowledge, there are no efficient algorithmic approaches to compute Sobolev IPM effectively, which hinders its practical applications. In this work, we establish a relation between Sobolev norm and weighted $L^p$-norm, and leverage it to propose a novel regularization for Sobolev IPM. By exploiting the graph structure, we demonstrate that the regularized Sobolev IPM provides a closed-form expression for fast computation. This advancement addresses long-standing computational challenges, and paves the way to apply Sobolev IPM for practical applications, even in large-scale settings. Additionally, the regularized Sobolev IPM is negative definite. Utilizing this property, we design positive-definite kernels upon the regularized Sobolev IPM, and provide preliminary evidences of their advantages for comparing probability measures on a given graph for document classification and topological data analysis. Tam Le, Truyen Nguyen, Hideitsu Hino, Kenji Fukumizu |
ICML | 3 |
| 2025 | An Efficient Orlicz-Sobolev Approach for Transporting Unbalanced Measures on a GraphabstractWe investigate optimal transport (OT) for measures on graph metric spaces with different total masses. To mitigate the limitations of traditional $L^p$ geometry, Orlicz-Wasserstein (OW) and generalized Sobolev transport (GST) employ \emph{Orlicz geometric structure}, leveraging convex functions to capture nuanced geometric relationships and remarkably contribute to advance certain machine learning approaches. However, both OW and GST are restricted to measures with equal total mass, limiting their applicability to real-world scenarios where mass variation is common, and input measures may have noisy supports, or outliers. To address unbalanced measures, OW can either incorporate mass constraints or marginal discrepancy penalization, but this leads to a more complex two-level optimization problem. Additionally, GST provides a scalable yet rigid framework, which poses significant challenges to extend GST to accommodate nonnegative measures. To tackle these challenges, in this work we revisit the entropy partial transport (EPT) problem. By exploiting Caffarelli \& McCann's insights, we develop a novel variant of EPT endowed with Orlicz geometric structure, called \emph{Orlicz-EPT}. We establish theoretical background to solve Orlicz-EPT using a binary search algorithmic approach. Especially, by leveraging the dual EPT and the underlying graph structure, we formulate a novel regularization approach that leads to the proposed \emph{Orlicz-Sobolev transport} (OST). Notably, we demonstrate that OST can be efficiently computed by simply solving a univariate optimization problem, in stark contrast to the intensive computation needed for Orlicz-EPT. Building on this, we derive geometric structures for OST and draw its connections to other transport distances. We empirically illustrate that OST is several-order faster than Orlicz-EPT. Furthermore, we show preliminary evidence on the advantages of OST for measures on a graph in document classification and topological data analysis. Tam Le, Truyen Nguyen, Hideitsu Hino, Kenji Fukumizu |
NeurIPS | 3 |
| 2025 | An (ε ,δ )-Accurate Level Set Estimation with a Stopping Criterion
Hideaki Ishibashi, Kota Matsui, Kentaro Kutsukake, Hideitsu Hino |
ECML/PKDD (2) | 4 |
| 2025 | Gradual Domain Adaptation via Normalizing FlowsabstractStandard domain adaptation methods do not work well when a large gap exists between the source and target domains. Gradual domain adaptation is one of the approaches used to address the problem. It involves leveraging the intermediate domain, which gradually shifts from the source domain to the target domain. In previous work, it is assumed that the number of intermediate domains is large and the distance between adjacent domains is small; hence, the gradual domain adaptation algorithm, involving self-training with unlabeled data sets, is applicable. In practice, however, gradual self-training will fail because the number of intermediate domains is limited and the distance between adjacent domains is large. We propose the use of normalizing flows to deal with this problem while maintaining the framework of unsupervised domain adaptation. The proposed method learns a transformation from the distribution of the target domains to the gaussian mixture distribution via the source domain. We evaluate our proposed method by experiments using real-world data sets and confirm that it mitigates the problem we have explained and improves the classification performance. Shogo Sagawa, Hideitsu Hino |
Neural Comput. | 2 |
| 2023 | A stopping criterion for Bayesian optimization by the gap of expected minimum simple regretsabstractBayesian optimization (BO) improves the efficiency of black-box optimization; however, the associated computational cost and power consumption remain dominant in the application of machine learning methods. This paper proposes a method of determining the stopping time in BO. The proposed criterion is based on the difference between the expectation of the minimum of a variant of the simple regrets before and after evaluating the objective function with a new parameter setting. Unlike existing stopping criteria, the proposed criterion is guaranteed to converge to the theoretically optimal stopping criterion for any choices of arbitrary acquisition functions and threshold values. Moreover, the threshold for the stopping criterion can be determined automatically and adaptively. We experimentally demonstrate that the proposed stopping criterion finds reasonable timing to stop a BO with a small number of evaluations of the objective function. Hideaki Ishibashi, Masayuki Karasuyama, Ichiro Takeuchi, Hideitsu Hino |
AISTATS | 4 |
| 2023 | Gaussian Process Koopman Mode DecompositionabstractWe propose a nonlinear probabilistic generative model of Koopman mode decomposition based on an unsupervised gaussian process. Existing data-driven methods for Koopman mode decomposition have focused on estimating the quantities specified by Koopman mode decomposition: eigenvalues, eigenfunctions, and modes. Our model enables the simultaneous estimation of these quantities and latent variables governed by an unknown dynamical system. Furthermore, we introduce an efficient strategy to estimate the parameters of our model by low-rank approximations of covariance matrices. Applying the proposed model to both synthetic data and a real-world epidemiological data set, we show that various analyses are available using the estimated parameters. Takahiro Kawashima, Hideitsu Hino |
Neural Comput. | 2 |
| 2023 | Cost-effective framework for gradual domain adaptation with multifidelityabstractIn domain adaptation, when there is a large distance between the source and target domains, the prediction performance will degrade. Gradual domain adaptation is one of the solutions to such an issue, assuming that we have access to intermediate domains, which shift gradually from the source to the target domain. In previous works, it was assumed that the number of samples in the intermediate domains was sufficiently large; hence, self-training was possible without the need for labeled data. If the number of accessible intermediate domains is restricted, the distances between domains become large, and self-training will fail. Practically, the cost of samples in intermediate domains will vary, and it is natural to consider that the closer an intermediate domain is to the target domain, the higher the cost of obtaining samples from the intermediate domain is. To solve the trade-off between cost and accuracy, we propose a framework that combines multifidelity and active domain adaptation. The effectiveness of the proposed method is evaluated by experiments with real-world datasets. Shogo Sagawa, Hideitsu Hino |
Neural Networks | 2 |
| 2023 | ATNAS: Automatic Termination for Neural Architecture SearchabstractNeural architecture search (NAS) is a framework for automating the design process of a neural network structure. While the recent one-shot approaches have reduced the search cost, there still exists an inherent trade-off between cost and performance. It is important to appropriately stop the search and further reduce the high cost of NAS. Meanwhile, the differentiable architecture search (DARTS), a typical one-shot approach, is known to suffer from overfitting. Heuristic early-stopping strategies have been proposed to overcome such performance degradation. In this paper, we propose a more versatile and principled early-stopping criterion on the basis of the evaluation of a gap between expectation values of generalisation errors of the previous and current search steps with respect to the architecture parameters. The stopping threshold is automatically determined at each search epoch without cost. In numerical experiments, we demonstrate the effectiveness of the proposed method. We stop the one-shot NAS algorithms and evaluate the acquired architectures on the benchmark datasets: NAS-Bench-201 and NATS-Bench. Our algorithm is shown to reduce the cost of the search process while maintaining a high performance. Kotaro Sakamoto, Hideaki Ishibashi, Rei Sato, Shinichi Shirakawa, Youhei Akimoto, Hideitsu Hino |
Neural Networks | 6 |
| 2022 | One-bit Submission for Locally Private Quasi-MLE: Its Asymptotic Normality and LimitationabstractLocal differential privacy (LDP) is an information-theoretic privacy definition suitable for statistical surveys that involve an untrusted data curator. An LDP version of quasi-maximum likelihood estimator (QMLE) has been developed, but the existing method to build LDP QMLE is difficult to implement for a large-scale survey system in the real world due to long waiting time, expensive communication cost, and the boundedness assumption of derivative of a log-likelihood function. We provided alternative LDP protocols without those issues, which are potentially much easily deployable to a large-scale survey. We also provided sufficient conditions for the consistency and asymptotic normality and limitations of our protocol. Our protocol is less burdensome for the users, and the theoretical guarantees cover more realistic cases than those for the existing method. Hajime Ono, Kazuhiro Minami, Hideitsu Hino |
AISTATS | 3 |
| 2022 | Domain Adaptation with Optimal Transport for Extended Variable SpaceabstractDomain adaptation aims to transfer knowledge of labeled instances obtained from a source domain to a target domain to fill the gap between the domains. Most domain adaptation methods assume that the source and target domains have the same dimensionality. Methods that are applicable when the number of features for each sample is different in each domain have rarely been studied, especially when no label information is given for the test data obtained from the target domain. In this paper, it is assumed that common features exist in both domains and that extra (new additional) features are observed in the target domain; hence, the dimensionality of the target domain is higher than that of the source domain. To leverage the homogeneity of the common features, the adaptation between the source and target domains is formulated as an optimal transport (OT) problem. In addition, a learning bound in the target domain for the proposed OT-based method is derived. The experiments with simulated and real-world data show that our proposed algorithm is able to obtain better model for the target domain by considering the extra features given for the target domain. Toshimitsu Aritake, Hideitsu Hino |
IJCNN | 2 |
| 2022 | Unsupervised Domain Adaptation for Extra Features in the Target Domain Using Optimal TransportabstractDomain adaptation aims to transfer knowledge of labeled instances obtained from a source domain to a target domain to fill the gap between the domains. Most domain adaptation methods assume that the source and target domains have the same dimensionality. Methods that are applicable when the number of features is different in each domain have rarely been studied, especially when no label information is given for the test data obtained from the target domain. In this letter, it is assumed that common features exist in both domains and that extra (new additional) features are observed in the target domain; hence, the dimensionality of the target domain is higher than that of the source domain. To leverage the homogeneity of the common features, the adaptation between these source and target domains is formulated as an optimal transport (OT) problem. In addition, a learning bound in the target domain for the proposed OT-based method is derived. The proposed algorithm is validated using both simulated and real-world data. Toshimitsu Aritake, Hideitsu Hino |
Neural Comput. | 2 |
| 2022 | Information Geometrically Generalized Covariate Shift AdaptationabstractMany machine learning methods assume that the training and test data follow the same distribution. However, in the real world, this assumption is often violated. In particular, the marginal distribution of the data changes, called covariate shift, is one of the most important research topics in machine learning. We show that the well-known family of covariate shift adaptation methods is unified in the framework of information geometry. Furthermore, we show that parameter search for a geometrically generalized covariate shift adaptation method can be achieved efficiently. Numerical experiments show that our generalization can achieve better performance than the existing methods it encompasses. Masanari Kimura, Hideitsu Hino |
Neural Comput. | 2 |
| 2022 | Detecting cell assemblies by NMF-based clustering from calcium imaging dataabstractA large number of neurons form cell assemblies that process information in the brain. Recent developments in measurement technology, one of which is calcium imaging, have made it possible to study cell assemblies. In this study, we aim to extract cell assemblies from calcium imaging data. We propose a clustering approach based on non-negative matrix factorization (NMF). The proposed approach first obtains a similarity matrix between neurons by NMF and then performs spectral clustering on it. The application of NMF entails the problem of model selection. The number of bases in NMF affects the result considerably, and a suitable selection method is yet to be established. We attempt to resolve this problem by model averaging with a newly defined estimator based on NMF. Experiments on simulated data suggest that the proposed approach is superior to conventional correlation-based clustering methods over a wide range of sampling rates. We also analyzed calcium imaging data of sleeping/waking mice and the results suggest that the size of the cell assembly depends on the degree and spatial extent of slow wave generation in the cerebral cortex. Mizuo Nagayama, Toshimitsu Aritake, Hideitsu Hino, Takeshi Kanda, Takehiro Miyazaki, Masashi Yanagisawa, Shotaro Akaho, Noboru Murata |
Neural Networks | 3 |
| 2021 | Bayesian Dynamic Mode Decomposition with Variational Matrix FactorizationabstractDynamic mode decomposition (DMD) and its extensions are data-driven methods that have substantially contributed to our understanding of dynamical systems. However, because DMD and most of its extensions are deterministic, it is difficult to treat probabilistic representations of parameters and predictions. In this work, we propose a novel formulation of a Bayesian DMD model. Our Bayesian DMD model is consistent with the procedure of standard DMD, which is to first determine the subspace of observations, and then compute the modes on that subspace. Variational matrix factorization makes it possible to realize a fully-Bayesian scheme of DMD. Moreover, we derive a Bayesian DMD model for incomplete data, which demonstrates the advantage of probabilistic modeling. Finally, both of nonlinear simulated and real-world datasets are used to illustrate the potential of the proposed method. Takahiro Kawashima, Hayaru Shouno, Hideitsu Hino |
AAAI | 3 |
| 2021 | Fast and robust multiplane single-molecule localization microscopy using a deep neural networkabstractSingle-molecule localization microscopy is a widely used technique in biological research for measuring the nanostructures of samples smaller than the diffraction limit.This study uses multifocal plane microscopy and addresses the three-dimensional (3D) single-molecule localization problem, where lateral and axial locations of molecules are estimated.However, when multifocal plane microscopy is used, the estimation accuracy of 3D localization is easily deteriorated by the small lateral drifts of camera positions.A 3D molecule localization problem was presented along with the lateral drift estimation as a compressed sensing problem.A deep neural network (DNN) was applied to solve this problem accurately and efficiently.The results show that the proposed method is robust to lateral drift and achieves an accuracy of 20 nm laterally and 50 nm axially without an explicit drift correction. Toshimitsu Aritake, Hideitsu Hino, Shigeyuki Namiki, Daisuke Asanuma, Kenzo Hirose, Noboru Murata |
Neurocomputing | 2 |
| 2021 | Pre-Training Acquisition Functions by Deep Reinforcement Learning for Fixed Budget Active LearningabstractAbstract There are many situations in supervised learning where the acquisition of data is very expensive and sometimes determined by a user’s budget. One way to address this limitation is active learning. In this study, we focus on a fixed budget regime and propose a novel active learning algorithm for the pool-based active learning problem. The proposed method performs active learning with a pre-trained acquisition function so that the maximum performance can be achieved when the number of data that can be acquired is fixed. To implement this active learning algorithm, the proposed method uses reinforcement learning based on deep neural networks as as a pre-trained acquisition function tailored for the fixed budget situation. By using the pre-trained deep Q-learning-based acquisition function, we can realize the active learner which selects a sample for annotation from the pool of unlabeled samples taking the fixed-budget situation into account. The proposed method is experimentally shown to be comparable with or superior to existing active learning methods, suggesting the effectiveness of the proposed approach for the fixed-budget active learning. Yusuke Taguchi, Hideitsu Hino, Keisuke Kameyama |
Neural Process. Lett. | 2 |
| 2020 | Stopping criterion for active learning based on deterministic generalization boundsabstractActive learning is a framework in which the learning machine can select the samples to be used for training. This technique is promising, particularly when the cost of data acquisition and labeling is high. In active learning, determining the timing at which learning should be stopped is a critical issue. In this study, we propose a criterion for automatically stopping active learning. The proposed stopping criterion is based on the difference in the expected generalization errors and hypothesis testing. We derive a novel upper bound for the difference in expected generalization errors before and after obtaining a new training datum based on PAC-Bayesian theory. Unlike ordinary PAC-Bayesian bounds, though, the proposed bound is deterministic; hence, there is no uncontrollable trade-off between the confidence and tightness of the inequality. We combine the upper bound with a statistical test to derive a stopping criterion for active learning. We demonstrate the effectiveness of the proposed method via experiments with both artificial and real datasets. Hideaki Ishibashi, Hideitsu Hino |
AISTATS | 2 |
| 2020 | Modal Principal Component AnalysisabstractPrincipal component analysis (PCA) is a widely used method for data processing, such as for dimension reduction and visualization. Standard PCA is known to be sensitive to outliers, and various robust PCA methods have been proposed. It has been shown that the robustness of many statistical methods can be improved using mode estimation instead of mean estimation, because mode estimation is not significantly affected by the presence of outliers. Thus, this study proposes a modal principal component analysis (MPCA), which is a robust PCA method based on mode estimation. The proposed method finds the minor component by estimating the mode of the projected data points. As a theoretical contribution, probabilistic convergence property, influence function, finite-sample breakdown point, and its lower bound for the proposed MPCA are derived. The experimental results show that the proposed method has advantages over conventional methods. Keishi Sando, Hideitsu Hino |
Neural Comput. | 2 |
| 2019 | Retrieved Image Refinement by Bootstrap Outlier Test
Hayato Watanabe, Hideitsu Hino, Shotaro Akaho, Noboru Murata |
CAIP (1) | 2 |
| 2019 | Sleep State Analysis Using Calcium Imaging Data by Non-negative Matrix Factorization
Mizuo Nagayama, Toshimitsu Aritake, Hideitsu Hino, Takeshi Kanda, Takehiro Miyazaki, Masashi Yanagisawa, Shotaro Akaho, Noboru Murata |
ICANN (1) | 3 |
| 2019 | On a Convergence Property of a Geometrical Algorithm for Statistical Manifolds
Shotaro Akaho, Hideitsu Hino, Noboru Murata |
ICONIP (5) | 2 |
| 2019 | Active Learning with Interpretable PredictorabstractActive learning is a method of constructing a useful prediction model with the minimum number of annotations or labeling for a response variable. It is widely used as a modern experimental design method, particularly for problems with high annotation cost. An appropriate reason for the selection of the next experimental setting is required for experiments with high annotation cost, for example, situations requiring large-scale experiments or long-term experiments such as agricultural examinations. In conventional active learning, it is only known that the samples that can improve the prediction accuracy of a prediction model are selected. This is not a satisfactory explanation for the selection of the next experimental setting for approving an experiment. In this paper, we propose a novel active learning algorithm with the following two models: a model to predict a response variable and a model to predict the amount of decrease in test loss. A new sample is selected using a model that predicts the amount of decrease in test loss. It is possible to provide a reason for sample selection by employing a model that can evaluate variable importance, e.g., using a random forest as a model of predicting the decrease in test loss. We applied the proposed method to multiple datasets and showed that the prediction performance of the proposed method is comparable to those of existing methods and the computational time is superior to those of existing methods. In addition, we demonstrated that it is possible to provide suitable reasons for selecting a sample in the process of active learning. Yusuke Taguchi, Keisuke Kameyama, Hideitsu Hino |
IJCNN | 3 |
| 2018 | Localizing Current Dipoles from EEG Data Using a Birth-Death Process
Keita Nakamura, Sho Sonoda, Hideitsu Hino, Masahiro Kawasaki, Shotaro Akaho, Noboru Murata |
BIBM | 3 |
| 2018 | Geometrical Formulation of the Nonnegative Matrix Factorization
Shotaro Akaho, Hideitsu Hino, Neneka Nara, Noboru Murata |
ICONIP (3) | 2 |
| 2018 | Information Geometric Perspective of Modal Linear Regression
Keishi Sando, Shotaro Akaho, Noboru Murata, Hideitsu Hino |
ICONIP (3) | 4 |
| 2018 | Estimation of neural connections from partially observed neural spikes
Taishi Iwasaki, Hideitsu Hino, Masami Tatsuno, Shotaro Akaho, Noboru Murata |
Neural Networks | 2 |
| 2018 | EEG dipole source localization with information criteria for multiple particle filters
Sho Sonoda, Keita Nakamura, Yuki Kaneda, Hideitsu Hino, Shotaro Akaho, Noboru Murata, Eri Miyauchi, Masahiro Kawasaki |
Neural Networks | 4 |
| 2018 | Toward Distribution Estimation under Local Differential Privacy with Small SamplesabstractAbstract A number of studies have recently been made on discrete distribution estimation in the local model, in which users obfuscate their personal data (e.g., location, response in a survey) by themselves and a data collector estimates a distribution of the original personal data from the obfuscated data. Unlike the centralized model, in which a trusted database administrator can access all users’ personal data, the local model does not suffer from the risk of data leakage. A representative privacy metric in this model is LDP (Local Differential Privacy), which controls the amount of information leakage by a parameter ∈ called privacy budget. When ∈ is small, a large amount of noise is added to the personal data, and therefore users’ privacy is strongly protected. However, when the number of users ℕ is small (e.g., a small-scale enterprise may not be able to collect large samples) or when most users adopt a small value of ∈, the estimation of the distribution becomes a very challenging task. The goal of this paper is to accurately estimate the distribution in the cases explained above. To achieve this goal, we focus on the EM (Expectation-Maximization) reconstruction method, which is a state-of-the-art statistical inference method, and propose a method to correct its estimation error (i.e., difference between the estimate and the true value) using the theory of Rilstone et al. We prove that the proposed method reduces the MSE (Mean Square Error) under some assumptions.We also evaluate the proposed method using three largescale datasets, two of which contain location data while the other contains census data. The results show that the proposed method significantly outperforms the EM reconstruction method in all of the datasets when ℕ or ∈ is small. Takao Murakami, Hideitsu Hino, Jun Sakuma |
Proc. Priv. Enhancing Technol. | 2 |
| 2017 | Double sparsity for multi-frame super resolution
Toshiyuki Kato, Hideitsu Hino, Noboru Murata |
Neurocomputing | 2 |
| 2017 | Local Intrinsic Dimension Estimation by Generalized Linear ModelingabstractWe propose a method for intrinsic dimension estimation. By fitting the power of distance from an inspection point and the number of samples included inside a ball with a radius equal to the distance, to a regression model, we estimate the goodness of fit. Then, by using the maximum likelihood method, we estimate the local intrinsic dimension around the inspection point. The proposed method is shown to be comparable to conventional methods in global intrinsic dimension estimation experiments. Furthermore, we experimentally show that the proposed method outperforms a conventional local dimension estimation method. Hideitsu Hino, Jun Fujiki, Shotaro Akaho, Noboru Murata |
Neural Comput. | 1 |
| 2017 | Group Sparsity Tensor Factorization for Re-Identification of Open Mobility TracesabstractRe-identification attacks based on a Markov chain model have been widely studied to understand how anonymized traces are linked to users. This approach is known to enable users to be re-identified with high accuracy when an adversary trains a personalized transition matrix for each target user using a large amount of training data, and when all of the anonymized traces are from the target users. In reality, however, the amount of training data for each target user can be very small, since many users disclose only a small amount of their location information to the public. In addition, many of the anonymized traces are from “non-target” users, whose personalized transition matrices cannot be trained in advance. This paper aims to quantify the risk of re-identification in the realistic situation explained earlier. We first utilize the fact that spatial data can form a group structure, and propose group sparsity tensor factorization to effectively train the personalized transition matrices from a small number of training traces. We second formulate a re-identification attack in an “open” scenario, where many of the anonymized traces are from non-target users. Specifically, we regard this type of attack as a biometric verification (or identification) task, and propose a framework and an algorithm for performing this task using a population transition matrix, which is computed from personalized transition matrices. Our experimental results using three real data sets show that a training method using tensor factorization significantly outperforms the maximum likelihood estimation method, and is further improved by incorporating group sparsity regularization. Takao Murakami, Atsunori Kanemura, Hideitsu Hino |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Change-point detection in a sequence of bags-of-dataabstractIn this paper, the limitation that is prominent in most existing works of change-point detection methods is addressed by proposing a nonparametric, computationally efficient method. The limitation is that most works assume that each data point observed at each time step is a single multi-dimensional vector. However, there are many situations where this does not hold. Therefore, a setting where each observation is a collection of random variables, which we call a bag of data, is considered. Kensuke Koshijima, Hideitsu Hino, Noboru Murata |
ICDE | 2 |
| 2016 | An Entropy Estimator Based on Polynomial Regression with Poisson Error Structure
Hideitsu Hino, Shotaro Akaho, Noboru Murata |
ICONIP (2) | 1 |
| 2016 | Non-parametric e-mixture of Density Functions
Hideitsu Hino, Ken Takano, Shotaro Akaho, Noboru Murata |
ICONIP (2) | 1 |
| 2016 | Nonparametric e-Mixture EstimationabstractThis study considers the common situation in data analysis when there are few observations of the distribution of interest or the target distribution, while abundant observations are available from auxiliary distributions. In this situation, it is natural to compensate for the lack of data from the target distribution by using data sets from these auxiliary distributions-in other words, approximating the target distribution in a subspace spanned by a set of auxiliary distributions. Mixture modeling is one of the simplest ways to integrate information from the target and auxiliary distributions in order to express the target distribution as accurately as possible. There are two typical mixtures in the context of information geometry: the [Formula: see text]- and [Formula: see text]-mixtures. The [Formula: see text]-mixture is applied in a variety of research fields because of the presence of the well-known expectation-maximazation algorithm for parameter estimation, whereas the [Formula: see text]-mixture is rarely used because of its difficulty of estimation, particularly for nonparametric models. The [Formula: see text]-mixture, however, is a well-tempered distribution that satisfies the principle of maximum entropy. To model a target distribution with scarce observations accurately, this letter proposes a novel framework for a nonparametric modeling of the [Formula: see text]-mixture and a geometrically inspired estimation algorithm. As numerical examples of the proposed framework, a transfer learning setup is considered. The experimental results show that this framework works well for three types of synthetic data sets, as well as an EEG real-world data set. Ken Takano, Hideitsu Hino, Shotaro Akaho, Noboru Murata |
Neural Comput. | 2 |
| 2015 | Personal Authentication Based on 3D Configuration of Micro-feature Points on Facial Surface
Takao Yoshinuma, Hideitsu Hino, Kazuhiro Fukui |
PSIVT | 2 |
| 2015 | Multi-frame image super resolution based on sparse coding
Toshiyuki Kato, Hideitsu Hino, Noboru Murata |
Neural Networks | 2 |
| 2015 | Change-Point Detection in a Sequence of Bags-of-DataabstractIn this paper, the limitation that is prominent in most existing works of change-point detection methods is addressed by proposing a nonparametric, computationally efficient method. The limitation is that most works assume that each data point observed at each time step is a single multi-dimensional vector. However, there are many situations where this does not hold. Therefore, a setting where each observation is a collection of random variables, which we call a bag of data, is considered. After estimating the underlying distribution behind each bag of data and embedding those distributions in a metric space, the change-point score is derived by evaluating how the sequence of distributions is fluctuating in the metric space using a distance-based information estimator. Also, a procedure that adaptively determines when to raise alerts is incorporated by calculating the confidence interval of the change-point score at each time step. This avoids raising false alarms in highly noisy situations and enables detecting changes of various magnitudes. A number of experimental studies and numerical examples are provided to demonstrate the generality and the effectiveness of our approach with both synthetic and real datasets. Kensuke Koshijima, Hideitsu Hino, Noboru Murata |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | A Non-parametric Maximum Entropy Clustering
Hideitsu Hino, Noboru Murata |
ICANN | 1 |
| 2014 | An Algorithm for Directed Graph Estimation
Hideitsu Hino, Atsushi Noda, Masami Tatsuno, Shotaro Akaho, Noboru Murata |
ICANN | 1 |
| 2014 | A Kernel Method to Extract Common Features Based on Mutual Information
Takamitsu Araki, Hideitsu Hino, Shotaro Akaho |
ICONIP (2) | 2 |
| 2014 | Sensing Visual Attention by Sequential PatternsabstractA method for sensing human visual attention is proposed. The method is based on the analysis of sequential image patterns of faces and irises observed at regular time intervals. The basic concept is to represent the set of image patterns produced by the action of gazing at a certain area as a nonlinear subspace in a high-dimensional pattern vector space. Such a space is called an attention subspace. In this framework, an input subspace from an unknown action is classified into an attention subspace of gazing at a certain area or into attention subspaces of gazing at other areas (named non-attention subspaces) by measuring the canonical angles between the input subspace and pre-computed dictionary subspaces. To maintain performance even in the presence of head movement, two mechanisms are introduced: 1) the kernel orthogonal mutual subspace method, which is suitable for classifying sets of multiple images, and 2) a kernel function for considering the head position in addition to a kernel function for pixel values. The stable performance of the proposed method including situations with head movements is demonstrated through experiments. Yasuyuki Yamazaki, Hideitsu Hino, Kazuhiro Fukui |
ICPR | 2 |
| 2014 | A Nonparametric Clustering Algorithm with a Quantile-Based Likelihood EstimatorabstractClustering is a representative of unsupervised learning and one of the important approaches in exploratory data analysis. By its very nature, clustering without strong assumption on data distribution is desirable. Information-theoretic clustering is a class of clustering methods that optimize information-theoretic quantities such as entropy and mutual information. These quantities can be estimated in a nonparametric manner, and information-theoretic clustering algorithms are capable of capturing various intrinsic data structures. It is also possible to estimate information-theoretic quantities using a data set with sampling weight for each datum. Assuming the data set is sampled from a certain cluster and assigning different sampling weights depending on the clusters, the cluster-conditional information-theoretic quantities are estimated. In this letter, a simple iterative clustering algorithm is proposed based on a nonparametric estimator of the log likelihood for weighted data sets. The clustering algorithm is also derived from the principle of conditional entropy minimization with maximum entropy regularization. The proposed algorithm does not contain a tuning parameter. The algorithm is experimentally shown to be comparable to or outperform conventional nonparametric clustering methods. Hideitsu Hino, Noboru Murata |
Neural Comput. | 1 |
| 2014 | Intrinsic Graph Structure Estimation Using Graph LaplacianabstractA graph is a mathematical representation of a set of variables where some pairs of the variables are connected by edges. Common examples of graphs are railroads, the Internet, and neural networks. It is both theoretically and practically important to estimate the intensity of direct connections between variables. In this study, a problem of estimating the intrinsic graph structure from observed data is considered. The observed data in this study are a matrix with elements representing dependency between nodes in the graph. The dependency represents more than direct connections because it includes influences of various paths. For example, each element of the observed matrix represents a co-occurrence of events at two nodes or a correlation of variables corresponding to two nodes. In this setting, spurious correlations make the estimation of direct connection difficult. To alleviate this difficulty, a digraph Laplacian is used for characterizing a graph. A generative model of this observed matrix is proposed, and a parameter estimation algorithm for the model is also introduced. The notable advantage of the proposed method is its ability to deal with directed graphs, while conventional graph structure estimation methods such as covariance selections are applicable only to undirected graphs. The algorithm is experimentally shown to be able to identify the intrinsic graph structure. Atsushi Noda, Hideitsu Hino, Masami Tatsuno, Shotaro Akaho, Noboru Murata |
Neural Comput. | 2 |
| 2013 | Pairwise Similarity for Line Extraction from Distorted Images
Hideitsu Hino, Jun Fujiki, Shotaro Akaho, Yoshihiko Mochizuki, Noboru Murata |
CAIP (2) | 1 |
| 2013 | Information estimators for weighted observations
Hideitsu Hino, Noboru Murata |
Neural Networks | 1 |
| 2012 | Robust Hypersurface Fitting Based on Random Sampling Approximations
Jun Fujiki, Shotaro Akaho, Hideitsu Hino, Noboru Murata |
ICONIP (3) | 3 |
| 2012 | An improved entropy-based multiple kernel learning
Hideitsu Hino, Tetsuji Ogawa |
ICPR | 1 |
| 2012 | Sliced inverse regression with conditional entropy minimization
Hideitsu Hino, Keigo Wakayama, Noboru Murata |
ICPR | 1 |
| 2012 | Multiple Kernel Learning with Gaussianity MeasuresabstractKernel methods are known to be effective for nonlinear multivariate analysis. One of the main issues in the practical use of kernel methods is the selection of kernel. There have been a lot of studies on kernel selection and kernel learning. Multiple kernel learning (MKL) is one of the promising kernel optimization approaches. Kernel methods are applied to various classifiers including Fisher discriminant analysis (FDA). FDA gives the Bayes optimal classification axis if the data distribution of each class in the feature space is a gaussian with a shared covariance structure. Based on this fact, an MKL framework based on the notion of gaussianity is proposed. As a concrete implementation, an empirical characteristic function is adopted to measure gaussianity in the feature space associated with a convex combination of kernel functions, and two MKL algorithms are derived. From experimental results on some data sets, we show that the proposed kernel learning followed by FDA offers strong classification power. Hideitsu Hino, Nima Reyhani, Noboru Murata |
Neural Comput. | 1 |
| 2011 | Robust Hyperplane Fitting Based on k-th Power Deviation and α-Quantile
Jun Fujiki, Shotaro Akaho, Hideitsu Hino, Noboru Murata |
CAIP (1) | 3 |
| 2011 | A Computationally Efficient Information Estimator for Weighted Data
Hideitsu Hino, Noboru Murata |
ICANN (2) | 1 |
| 2011 | Speaker recognition using multiple kernel learning based on conditional entropy minimizationabstractWe applied a multiple kernel learning (MKL) method based on information-theoretic optimization to speaker recognition. Most of the kernel methods applied to speaker recognition systems require a suitable kernel function and its parameters to be determined for a given data set. In contrast, MKL eliminates the need for strict determination of the kernel function and parameters by using a convex combination of element kernels. In the present paper, we describe an MKL algorithm based on conditional entropy minimization (MCEM). We experimentally verified the effectiveness of MCEM for speaker classification; this method reduced the speaker error rate as compared to conventional methods. Tetsuji Ogawa, Hideitsu Hino, Nima Reyhani, Noboru Murata, Tetsunori Kobayashi |
ICASSP | 2 |
| 2011 | Speaker Verification Robust to Talking Style Variation Using Multiple Kernel Learning Based on Conditional Entropy Minimization
Tetsuji Ogawa, Hideitsu Hino, Noboru Murata, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2011 | New Probabilistic Bounds on Eigenvalues and Eigenvectors of Random Kernel Matrices
Nima Reyhani, Hideitsu Hino, Ricardo Vigário |
UAI | 2 |
| 2011 | An Estimation of Generalized Bradley-Terry Models Based on the em AlgorithmabstractThe Bradley-Terry model is a statistical representation for one's preference or ranking data by using pairwise comparison results of items. For estimation of the model, several methods based on the sum of weighted Kullback-Leibler divergences have been proposed from various contexts. The purpose of this letter is to interpret an estimation mechanism of the Bradley-Terry model from the viewpoint of flatness, a fundamental notion used in information geometry. Based on this point of view, a new estimation method is proposed on a framework of the em algorithm. The proposed method is different in its objective function from that of conventional methods, especially in treating unobserved comparisons, and it is consistently interpreted in a probability simplex. An estimation method with weight adaptation is also proposed from a viewpoint of the sensitivity. Experimental results show that the proposed method works appropriately, and weight adaptation improves accuracy of the estimate. Yu Fujimoto, Hideitsu Hino, Noboru Murata |
Neural Comput. | 2 |
| 2010 | Multiple Kernel Learning by Conditional Entropy MinimizationabstractKernel methods have been successfully used in many practical machine learning problems. Choosing a suitable kernel is left to the practitioner. A common way to an automatic selection of optimal kernels is to learn a linear combination of element kernels. In this paper, a novel framework of multiple kernel learning is proposed based on conditional entropy minimization criterion. For the proposed framework, three multiple kernel learning algorithms are derived. The algorithms are experimentally shown to be comparable to or outperform kernel Fisher discriminant analysis and other multiple kernel learning algorithms on benchmark data sets. Hideitsu Hino, Nima Reyhani, Noboru Murata |
ICMLA | 1 |
| 2010 | Self-Calibration of Radially Symmetric Distortion by Model SelectionabstractFor self-calibration of general radially symmetric distortion (RSD) of omni directional cameras such as fish-eye lenses, calibration parameters are usually estimated so that curved lines, which are supposed to be straight in the real-world, are mapped to straight lines in the calibrated image, which is assumed to be taken by an ideal pin-hole camera. In this paper, a method of calibrating RSD is introduced base on the notion of principal component analysis (PCA). In the proposed method, the distortion function, which maps a distorted image to an ideal pin-hole camera image, is assumed to be a linear combination of a certain class of basis functions, and an algorithm for solving its coefficients by using line patterns is given. Then a method of selecting good basis functions is proposed, which aims to realize appropriate calibration in practice. Experimental results for synthetic data and real images are presented to emonstrate the performance of our calibration method. Jun Fujiki, Hideitsu Hino, Yumi Usami, Shotaro Akaho, Noboru Murata |
ICPR | 2 |
| 2010 | A Grouped Ranking Model for Item Preference ParameterabstractGiven a set of rating data for a set of items, determining preference levels of items is a matter of importance. Various probability models have been proposed to solve this task. One such model is the Plackett-Luce model, which parameterizes the preference level of each item by a real value. In this letter, the Plackett-Luce model is generalized to cope with grouped ranking observations such as movie or restaurant ratings. Since it is difficult to maximize the likelihood of the proposed model directly, a feasible approximation is derived, and the em algorithm is adopted to find the model parameter by maximizing the approximate likelihood which is easily evaluated. The proposed model is extended to a mixture model, and two applications are proposed. To show the effectiveness of the proposed model, numerical experiments with real-world data are carried out. Hideitsu Hino, Yu Fujimoto, Noboru Murata |
Neural Comput. | 1 |
| 2010 | A Conditional Entropy Minimization Criterion for Dimensionality Reduction and Multiple Kernel LearningabstractReducing the dimensionality of high-dimensional data without losing its essential information is an important task in information processing. When class labels of training data are available, Fisher discriminant analysis (FDA) has been widely used. However, the optimality of FDA is guaranteed only in a very restricted ideal circumstance, and it is often observed that FDA does not provide a good classification surface for many real problems. This letter treats the problem of supervised dimensionality reduction from the viewpoint of information theory and proposes a framework of dimensionality reduction based on class-conditional entropy minimization. The proposed linear dimensionality-reduction technique is validated both theoretically and experimentally. Then, through kernel Fisher discriminant analysis (KFDA), the multiple kernel learning problem is treated in the proposed framework, and a novel algorithm, which iteratively optimizes the parameters of the classification function and kernel combination coefficients, is proposed. The algorithm is experimentally shown to be comparable to or outperforms KFDA for large-scale benchmark data sets, and comparable to other multiple kernel learning techniques on the yeast protein function annotation task. Hideitsu Hino, Noboru Murata |
Neural Comput. | 1 |
| 2009 | Calibration of Radially Symmetric Distortion by Fitting Principal Component
Hideitsu Hino, Yumi Usami, Jun Fujiki, Shotaro Akaho, Noboru Murata |
CAIP | 1 |
| 2009 | An Information Theoretic Perspective of the Sparse Coding
Hideitsu Hino, Noboru Murata |
ISNN (1) | 1 |
| 2009 | Item Preference Parameters from Grouped Ranking Observations
Hideitsu Hino, Yu Fujimoto, Noboru Murata |
PAKDD | 1 |