Ryohei Nakano

dblp:43/294 · DBLP profile ↗
← Back
66ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 4 first-authorDatabases, data management, data science and information retrieval · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 51% Optimization for machine learning · 21% Knowledge representation and reasoning · 14%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 88% Mathematical optimization · 12%
Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 46% Web and social media mining · 36% Data models and query languages · 18%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design
influence maximization
0.112007
Extracting Influential Nodes for Information Diffusion on a Social Network · AAAI 2007
Web and social media mining
social network analysis
0.012007
Extracting Influential Nodes for Information Diffusion on a Social Network · AAAI 2007
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.011998
SMEM Algorithm for Mixture Models · NIPS 1998
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation
0.011998
SMEM Algorithm for Mixture Models · NIPS 1998
Machine learning › Optimization for machine learning
second-order optimization
0.011996
Second-order Learning Algorithm with Squared Penalty Term · NIPS 1996
Machine learning › Optimization for machine learning › non-convex optimization › global optimization
deterministic annealing
0.011994
Deterministic Annealing Variant of the EM Algorithm · NIPS 1994
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.011994
Deterministic Annealing Variant of the EM Algorithm · NIPS 1994
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation
0.011994
Deterministic Annealing Variant of the EM Algorithm · NIPS 1994
Query processing and optimization › aggregation
aggregate functions
0.011990
Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions · ACM Trans. Database Syst. 1990
Query processing and optimization › query optimization
logical query optimization
0.011990
Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions · ACM Trans. Database Syst. 1990
Query processing and optimization › query compilation
relational algebra translation
0.011990
Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions · ACM Trans. Database Syst. 1990
Data models and query languages › query language
relational calculus
0.011990
Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions · ACM Trans. Database Syst. 1990
Query processing and optimization
rule rewriting
0.011990
Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions · ACM Trans. Database Syst. 1990
Mathematical optimization
continuous optimization
0.011996
Second-order Learning Algorithm with Squared Penalty Term · NIPS 1996
Mathematical optimization › constrained optimization
penalty methods
0.011996
Second-order Learning Algorithm with Squared Penalty Term · NIPS 1996
Data models and query languages › relational model
null values
0.011990
Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions · ACM Trans. Database Syst. 1990
Data models and query languages
SQL
0.011990
Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions · ACM Trans. Database Syst. 1990

Methods — techniques the papers use, named apart from their topics

influence diffusion models · 0.1influence diffusion model · 0.1squared penalty · 0.0second-order learning · 0.0expectation-maximization · 0.0neural network · 0.0statistical mechanics analogy · 0.0maximum entropy principle · 0.0deterministic annealing · 0.0two-phase rewriting · 0.0heuristic rewriting · 0.0
YearPublicationVenuePosition
2019 Mixture of Multilayer Perceptron Regressions
abstract
This paper investigates mixture of multilayer perceptron (MLP) regressions. Although mixture of MLP regressions (MoMR) can be a strong fitting model for noisy data, the research on it has been rare. We employ soft mixture approach and use the Expectation-Maximization (EM) algorithm as a basic learning method. Our learning method goes in a double-looped manner; the outer loop is controlled by the EM and the inner loop by MLP learning method. Given data, we will have many models; thus, we need a criterion to select the best. Bayesian Information Criterion (BIC) is used here because it works nicely for MLP model selection. Our experiments showed that the proposed MoMR method found the expected MoMR model as the best for artificial data and selected the MoMR model having smaller error than any linear models for real noisy data.
Ryohei Nakano, Seiya Satoh
ICPRAM1
2019 Faster RBF Network Learning Utilizing Singular Regions
abstract
There are two ways to learn radial basis function (RBF) networks: one-stage and two-stage learnings. Recently a very powerful one-stage learning method called RBF-SSF has been proposed, which can stably find a series of excellent solutions, making good use of singular regions, and can monotonically decrease training error along with the increase of hidden units. RBF-SSF was built by applying the SSF (singularity stairs following) paradigm to RBF networks; the SSF paradigm was originally and successfully proposed for multilayer perceptrons. Although RBF-SSF has the strong capability to find excellent solutions, it required a lot of time mainly because it computes the Hessian. This paper proposes a faster version of RBF-SSF called RBF-SSF(pH) by introducing partial calculation of the Hessian. The experiments using two datasets showed RBF-SSF(pH) ran as fast as usual one-stage learning methods while keeping the excellent solution quality.
Seiya Satoh, Ryohei Nakano
ICPRAM2
2017 How New Information Criteria WAIC and WBIC Worked for MLP Model Selection
Seiya Satoh, Ryohei Nakano
ICPRAM2
2017 Performance of Complex-Valued Multilayer Perceptrons Largely Depends on Learning Methods
Seiya Satoh, Ryohei Nakano
IJCCI2
2016 How complex-valued multilayer perceptron can predict the behavior of deterministic chaos
abstract
A complex-valued multilayer perceptron has the capability to represent complicated periodicity. We employ a very powerful learning method called C-SSF for learning a complex-valued multilayer perceptron. C-SSF finds a series of excellent solutions through successive learning. In deterministic chaos, long-term prediction is considered impossible. We apply C-SSF to two kinds of deterministic chaos and evaluate the learning and prediction performance of C-SSF.
Seiya Satoh, Ryohei Nakano
IJCNN2
2015 Complex-valued multilayer perceptron learning using singular regions and search pruning
abstract
In the search space of a complex-valued multilayer perceptron (C-MLP) there exist flat areas called singular regions. Although singular regions cause serious stagnation of learning, there exist descending paths from the regions. Based on this observation, a completely new learning method for C-MLP, called C-SSF1.0, was proposed, making good use of singular regions to stably find excellent solutions of successive C-MLPs. However, the method takes longer time than an existing method. This paper proposes a faster version of C-SSF called C-SSF1.1 by introducing search pruning. Our experiments showed the proposed method ran a few times faster than C-SSF1.0 without losing excellent solutions quality.
Seiya Satoh, Ryohei Nakano
IJCNN2
2014 Complex-Valued Multilayer Perceptron Search Utilizing Singular Regions of Complex-Valued Parameter Space
Seiya Satoh, Ryohei Nakano
ICANN2
2013 Fast and Stable Learning Utilizing Singular Regions of Multilayer Perceptron
Seiya Satoh, Ryohei Nakano
Neural Process. Lett.2
2012 Complex-Valued Multilayer Perceptron Search Utilizing Eigen Vector Descent and Reducibility Mapping
Shinya Suzumura, Ryohei Nakano
ICANN (2)2
2010 Nominally Conditioned Linear Regression
Yusuke Tanahashi, Ryohei Nakano, Kazumi Saito
ICANN (3)2
2010 Extracting influential nodes on a social network for information diffusion
Masahiro Kimura, Kazumi Saito, Ryohei Nakano, Hiroshi Motoda
Data Min. Knowl. Discov.3
2010 Empirical analysis of an on-line adaptive system using a mixture of Bayesian networks
Daisuke Kitakoshi, Hiroyuki Shioya, Ryohei Nakano
Inf. Sci.3
2010 Multi-directional search from the primitive initial point for Gaussian mixture estimation using variational Bayes method
Yuta Ishikawa, Ichiro Takeuchi, Ryohei Nakano
Neural Networks3
2009 Bidirectional Clustering of MLP Weights for Finding Nominally Conditioned Polynomials
Yusuke Tanahashi, Ryohei Nakano
ICANN (2)2
2009 A Bayesian Graph Clustering Approach Using the Prior Based on Degree Distribution
Naoyuki Harada, Yuta Ishikawa, Ichiro Takeuchi, Ryohei Nakano
ICONIP (1)4
2009 Variational Bayes from the Primitive Initial Point for Gaussian Mixture Estimation
Yuta Ishikawa, Ichiro Takeuchi, Ryohei Nakano
ICONIP (1)3
2008 Optimizing Sparse Kernel Ridge Regression hyperparameters based on leave-one-out cross-validation
abstract
Kernel ridge regression (KRR) is a nonlinear extension of the ridge regression. The performance of the KRR depends on its hyperparameters such as a penalty factor C, and RBF kernel parameter sigma. We employ a method called MCV-KRR which optimizes the KRR hyperparameters so that a cross-validation error is minimized. This method becomes equivalent to a predictive approach to Gaussian process. Since the cost of KRR training is O(N3) where N is a data size, to reduce this complexity, some sparse approximation of the KRR is recently studied. In this paper, we apply the minimum cross-validation (MCV) approach to such sparse approximation. Our experiments show the MCV with the sparse approximation of the KRR can achieve almost the same generalization performance as the MCV-KRR with much lower cost.
Masayuki Karasuyama, Ryohei Nakano
IJCNN2
2008 EM Algorithm with PIP Initialization and Temperature-Based Selection
Yuta Ishikawa, Ryohei Nakano
KES (3)2
2008 Reducing SVR Support Vectors by Using Backward Deletion
Masayuki Karasuyama, Ichiro Takeuchi, Ryohei Nakano
KES (3)3
2008 Prediction of Information Diffusion Probabilities for Independent Cascade Model
Kazumi Saito, Ryohei Nakano, Masahiro Kimura
KES (3)2
2007 Extracting Influential Nodes for Information Diffusion on a Social Network
Masahiro Kimura, Kazumi Saito, Ryohei Nakano
AAAI3
2007 Obtaining EM Initial Points by Using the Primitive Initial Point and Subsampling Strategy
abstract
The EM algorithm is an efficient algorithm to obtain the ML estimate for incomplete data, but has the local optimality problem. The deterministic annealing EM (DAEM) algorithm was once proposed to solve this problem, which begins a search from the primitive initial point. Then the mes-EM algorithm was proposed: a variant of the m-EM algorithm which begins the multiple-token EM search from the primitive initial point. The mes-EM could obtain excellent solutions in compensation for rather high computing cost. This paper proposes a lighter version of the mes-EM algorithm using the subsampling strategy and evaluates its performance.
Yuta Ishikawa, Ryohei Nakano
IJCNN2
2007 Optimizing SVR Hyperparameters via Fast Cross-Validation using AOSVR
abstract
The performance of support vector regression (SVR) deeply depends on its hyperparameters such as an insensitive zone thickness, a penalty factor, and kernel parameters. A method called MCV-SVR was once proposed, which optimizes SVR hyperparameters so that cross-validation error is minimized. However, the computational cost of CV is usually high. In this paper we apply accurate online support vector regression (AOSVR) to the MCV-SVR cross-validation procedure. The AOSVR enables an efficient update of a trained SVR function when a sample is removed from training data. We show the AOSVR dramatically accelerates the MCV-SVR. Moreover, our experiments using real-world data showed our faster MCV-SVR has better generalization than other existing methods such as Bayesian SVR or practical setting.
Masayuki Karasuyama, Ryohei Nakano
IJCNN2
2006 Landscape of a Likelihood Surface for a Gaussian Mixture and its use for the EM Algorithm
abstract
The EM algorithm is an efficient algorithm to obtain the ML estimate for incomplete data, but has the local optimality problem. The deterministic annealing EM (DAEM) algorithm was once proposed to solve this problem, which begins a search from the primitive initial point. Then the multi-thread EM (m-EM) algorithm was proposed, which begins the multiple-token EM search from the primitive initial point, resulting in excellent solutions in compensation for a rather heavy computing cost. These previous work indicate the potential of using the primitive initial point as the starting point of the EM algorithm. The paper investigates experimentally the characteristics of a landscape of a likelihood surface around the primitive initial point for a multivariate Gaussian mixture, and based on the observation proposes sensible ways of running the EM algorithm.
Yuta Ishikawa, Ryohei Nakano
IJCNN2
2006 Revised Optimizer of SVR Hyperparameters Minimizing Cross-Validation Error
abstract
The performance of support vector regression (SVR) deeply depends on its hyperparameters such as an insensitive zone thickness epsiv, a penalty factor C, and RBF kernel parameter sigma. A method called MCV-SVR was once proposed, which optimizes SVR hyperparameters so that a cross-validation error is minimized. However, as pointed out in this paper, the MCV-SVR (or its variants) has numerical instability in gradient calculation, which may cause bad influence on performance. Thus, this paper introduces a new method of computing the gradient of parameters with respect to hyperparameters. The revised optimizers incorporating the new method is shown to be free from the instability problem. Our experiments using three data sets showed that the revised optimizers considerably improved generalization performance of the MCV-SVR or its variant, and outperformed other methods such as multi-layer perceptrons or SVR with practical setting of hyperparameters.
Masayuki Karasuyama, Daisuke Kitakoshi, Ryohei Nakano
IJCNN3
2006 Improving Convergence Performance of PageRank Computation Based on Step-Length Calculation Approach
Kazumi Saito, Ryohei Nakano
KES (2)2
2006 Finding Nominally Conditioned Multivariate Polynomials Using a Four-Layer Perceptron Having Shared Weights
Yusuke Tanahashi, Kazumi Saito, Daisuke Kitakoshi, Ryohei Nakano
KES (2)4
2005 Yet faster method to optimize SVR hyperparameters based on minimizing cross-validation error
abstract
The performance of support vector (SV) regression deeply depends on its hyperparameters such as an insensitive zone thickness, a penalty factor, kernel function parameters. A method called MCV-SVR was once proposed, which optimizes SVR hyperparameters /spl lambda/ so that a cross-validation error is minimized. The method iterates two steps until convergence; step 1 optimizes parameters /spl theta/ under given /spl lambda/, while step 2 improves /spl lambda/ under given /spl theta/. Recently a faster version called the MCV-SVR-light was proposed, which accelerates step 2 by pruning. The present paper yet accelerates step 1 of the MCV-SVR-light by pruning without affecting solution quality. Here the pruning means confining the process to support vectors. Our experiments using three data sets show that the proposed method converged faster than the existing methods while the generalization performance remained comparable.
Kenji Kobayashi, Daisuke Kitakoshi, Ryohei Nakano
IJCNN3
2005 Weight sharing on naive Bayes document model
abstract
In this paper, we study weight sharing on the naive Bayes document model. Firstly we consider splitting words into a relatively small number of groups such that words in each group have the same parameter value. This problem can be regarded as a probabilistic parameter sharing task. In this task, we formalize the problem in terms of maximum likelihood estimation, and then propose an algorithm for this purpose. Secondly we focus on an adaptive hyperparameter estimation problem based on prior distributions constructed by using such word groups. This problem can be regarded as a hyperparameter sharing task. In this task, we describe a framework and algorithm, which enables to derive the unique optimal solution in the context of leave-one-out cross validation. In our experiments using a benchmark document set called webkb, we show a series of simulation results using the proposed algorithms.
Kazumi Saito, Ryohei Nakano
IJCNN2
2005 Finding a succinct multi-layer perceptron having shared weights
abstract
We present a method to find a succinct neural network having shared weights. We focus on weight sharing. Weight sharing constrains the freedom of weight values and weights are allowed to have one of common weights. A near-zero common weight can be eliminated, called weight pruning. Recently, a weight sharing method called BCW has been proposed. The BCW employs merge and split operations based on 2nd-order optimal criteria, and can escape local optima through bidirectional clustering. However, the BCW assumes a vital network parameter J, the number of hidden units, is given. This paper modifies the BCW to make the procedure faster so that the selection of J based on cross-validation can be done in reasonable CPU time. Our experiments showed that the proposed method can restore the original model for an artificial data set, and finds a small number of common weights and an interesting tendency for a real data set.
Yusuke Tanahashi, Xiang-Fang Chin, Kazumi Saito, Ryohei Nakano
IJCNN4
2005 Analysis for Adaptability of Policy-Improving System with a Mixture Model of Bayesian Networks to Dynamic Environments
Daisuke Kitakoshi, Hiroyuki Shioya, Ryohei Nakano
KES (4)3
2005 Model Selection and Weight Sharing of Multi-layer Perceptrons
Yusuke Tanahashi, Kazumi Saito, Ryohei Nakano
KES (4)3
2004 Threshold-based multi-thread EM algorithm
abstract
The EM algorithm is an efficient algorithm to obtain the ML estimate for incomplete data, but has the local optimality problem. The deterministic annealing EM (DAEM) algorithm was once proposed to solve the problem, but the global optimality is not guaranteed because of a single-token search. Then, the multi-thread DAEM (m-DAEM) algorithm was proposed by incorporating a multiple-token search with solution quality improvement with a heavy computing cost. Later, another variant of m-DAEM (/spl epsiv/-DAEM) was proposed by introducing threshold-based dynamic annealing with more quality improvement for an adequate threshold /spl epsiv/; however, finding such /spl epsiv/ is not easy. This paper proposes a new variant of EM, called /spl epsiv/-EM, by incorporating a multiple-token search together with threshold-based bifurcation. Our experiments using Gaussian mixture estimation problems showed that the /spl epsiv/-EM finds excellent solutions with relatively small computing cost, and the threshold /spl epsiv/ plays a key role in reducing computing cost.
Tetsuro Kawai, Ryohei Nakano
IJCNN2
2004 Extracting characteristic words of text using neural networks
abstract
In this paper, we consider models for estimating categories of documents and extracting characteristic words of such categories. To this end, we focus on three models, i.e., naive Bayes and two types of neural networks formalized as statistical models. Here, suitable categories of documents are estimated based on posterior probabilities, and characteristic words are extracted based on the magnitude of resulting parameter values. In our experiments using a set of real Web pages, we compare these models in the aspect of categorization performances and extraction capabilities of characteristic words.
Kazumi Saito, Ryohei Nakano
IJCNN2
2004 Piecewise Multivariate Polynomials Using a Four-Layer Perceptron
Yusuke Tanahashi, Kazumi Saito, Ryohei Nakano
KES3
2004 Learning an Evaluation Function for Shogi from Data of Games
Satoshi Tanimoto, Ryohei Nakano
KES2
2003 Optimizing Support Vector regression hyperparameters based on cross-validation
abstract
This paper proposes a method to optimize hyperparameters for Support Vector (SV) regression so that the cross-validation error is minimized. The performance of SV regression depends on its hyperparameters such as /spl epsiv/ (the thickness of a tube), C (a penalty factor), /spl sigma/ (kernel function parameter), and so on. This paper employs the procedure of cross-validation to optimize these hyperparameters together with training the corresponding SV regression models; thus, the learning is performed by using a coordinate descent method. Since an error surface produced by the usual /spl epsiv/-insensitive l/sub 1/ loss is not smooth, not suitable for our approach, we introduce the /spl epsiv/-insensitive l/sub 2/ loss. The experiments show the l/sub 2/ loss produces very smooth error surfaces and our coordinate descent nicely works, reaching the model whose validation performance is globally optimal.
Kentaro Ito, Ryohei Nakano
IJCNN2
2003 Threshold-based dynamic annealing for multi-thread DAEM and its extreme
abstract
The EM algorithm is an efficient algorithm to obtain the ML estimate for incomplete data, but has the local optimality problem. The deterministic annealing EM (DAEM) algorithm was once proposed to solve the problem, but it not guaranteed to obtain the global optimum since it employs a single token search. Then multi-thread DAEM (m-DAEM) algorithm was proposed by incorporating a search framework of multiple tokens, giving further improvement of solution quality with a heavy computing cost. This paper proposes another variant of m-DAEM, called /spl epsiv/-DAEM, by introducing threshold-based dynamic annealing where the Hessian information is made good use of. Given the adequate /spl epsiv/, the /spl epsiv/-DAEM shows excellent performance. Moreover, its extreme case where /spl epsiv/ /spl rarr/ /spl infin/, called the multi-thread EM, also shows rather excellent performance.
Masaharu Takada, Ryohei Nakano
IJCNN2
2002 Structuring Neural Networks through Bidirectional Clustering of Weights
Kazumi Saito, Ryohei Nakano
Discovery Science2
2002 Extracting regression rules from neural networks
Kazumi Saito, Ryohei Nakano
Neural Networks2
2001 Finding Polynomials to Fit Multivariate Data Having Numeric and Nominal Variables
Ryohei Nakano, Kazumi Saito
IDA1
2000 Discovery of Nominally Conditioned Polynomials Using Neural Networks, Vector Quantizers and Decision Trees
Kazumi Saito, Ryohei Nakano
Discovery Science2
2000 Discovery of Relevant Weights by Minimizing Cross-Validation Error
Kazumi Saito, Ryohei Nakano
PAKDD2
2000 Second-Order Learning Algorithm with Squared Penalty Term
abstract
This article compares three penalty terms with respect to the efficiency of supervised learning, by using first- and second-order off-line learning algorithms and a first-order on-line algorithm. Our experiments showed that for a reasonably adequate penalty factor, the combination of the squared penalty term and the second-order learning algorithm drastically improves the convergence performance in comparison to the other combinations, at the same time bringing about excellent generalization performance. Moreover, in order to understand how differently each penalty term works, a function surface evaluation is described. Finally, we show how cross validation can be applied to find an optimal penalty factor.
Kazumi Saito, Ryohei Nakano
Neural Comput.2
2000 SMEM Algorithm for Mixture Models
abstract
We present a split-and-merge expectation-maximization (SMEM) algorithm to overcome the local maxima problem in parameter estimation of finite mixture models. In the case of mixture models, local maxima often involve having too many components of a mixture model in one part of the space and too few in another, widely separated part of the space. To escape from such configurations, we repeatedly perform simultaneous split-and-merge operations using a new criterion for efficiently selecting the split-and-merge candidates. We apply the proposed algorithm to the training of gaussian mixtures and mixtures of factor analyzers using synthetic and real data and show the effectiveness of using the split-and-merge operations to improve the likelihood of both the training data and of held-out test data. We also show the practical usefulness of the proposed algorithm by applying it to image compression and pattern recognition problems.
Naonori Ueda, Ryohei Nakano, Zoubin Ghahramani, Geoffrey E. Hinton
Neural Comput.2
2000 Stable behavior in a recurrent neural network for a finite state machine
Ken-ichi Arai, Ryohei Nakano
Neural Networks2
1999 Discovery of a Set of Nominally Conditioned Polynomials
Ryohei Nakano, Kazumi Saito
Discovery Science1
1998 Computational Characteristics of Law Discovery Using Neural Networks
Ryohei Nakano, Kazumi Saito
Discovery Science1
1998 SMEM Algorithm for Mixture Models
Naonori Ueda, Ryohei Nakano, Zoubin Ghahramani, Geoffrey E. Hinton
NIPS2
1998 Learning dynamical systems by recurrent neural networks from orbits
Masahiro Kimura, Ryohei Nakano
Neural Networks2
1998 Deterministic annealing EM algorithm
Naonori Ueda, Ryohei Nakano
Neural Networks2
1997 Unique Representations of Dynamical Systems Produced by Recurrent Neural Networks
Masahiro Kimura, Ryohei Nakano
ICANN2
1997 Adaptive β Scheduling Learning Method of Finite State Automata by Recurrent Neural Networks
Kenichi Arai, Ryohei Nakano
ICONIP (1)2
1997 Numeric Law Discovery Using Neural Networks
Kazumi Saito, Ryohei Nakano
ICONIP (2)2
1997 Law Discovery using Neural Networks
Kazumi Saito, Ryohei Nakano
IJCAI2
1997 Partial BFGS Update and Efficient Step-Length Calculation for Three-Layer Neural Networks
abstract
Second-order learning algorithms based on quasi-Newton methods have two problems. First, standard quasi-Newton methods are impractical for large-scale problems because they require N2 storage space to maintain an approximation to an inverse Hessian matrix (N is the number of weights). Second, a line search to calculate a reasonably accurate step length is indispensable for these algorithms. In order to provide desirable performance, an efficient and reasonably accurate line search is needed. To overcome these problems, we propose a new second-order learning algorithm. Descent direction is calculated on the basis of a partial Broydon-Fletcher-Goldfarb-Shanno (BFGS) update with 2Ns memory space (s < < N), and a reasonably accurate step length is efficiently calculated as the minimal point of a second-order approximation to the objective function with respect to the step length. Our experiments, which use a parity problem and a speech synthesis problem, have shown that the proposed algorithm outperformed major learning algorithms. Moreover, it turned out that an efficient and accurate step-length calculation plays an important role for the convergence of quasi-Newton algorithms, and a partial BFGS update greatly saves storage space without losing the convergence performance.
Kazumi Saito, Ryohei Nakano
Neural Comput.2
1996 Annealed RNN Learning of Finite State Automata
Ken-ichi Arai, Ryohei Nakano
ICANN2
1996 Learning Dynamical Systems Produced by Recurrent Neural Networks
Masahiro Kimura, Ryohei Nakano
ICANN2
1996 Second-order Learning Algorithm with Squared Penalty Term
Kazumi Saito, Ryohei Nakano
NIPS2
1996 Scheduling by Genetic Local Search with Multi-Step Crossover
Takeshi Yamada, Ryohei Nakano
PPSN2
1994 Deterministic Annealing Variant of the EM Algorithm
abstract
We present a deterministic annealing variant of the EM algorithm for maximum likelihood parameter estimation problems. In our approach, the EM process is reformulated as the problem of min(cid:173) imizing the thermodynamic free energy by using the principle of maximum entropy and statistical mechanics analogy. Unlike simu(cid:173) lated annealing approaches, this minimization is deterministically performed. Moreover, the derived algorithm, unlike the conven(cid:173) tional EM algorithm, can obtain better estimates free of the initial parameter values.
Naonori Ueda, Ryohei Nakano
NIPS2
1994 Optimal Population Size under Constant Computation Cost
Ryohei Nakano, Yuval Davidor, Takeshi Yamada
PPSN1
1994 A new competitive learning approach based on an equidistortion principle for designing optimal vector quantizers
Naonori Ueda, Ryohei Nakano
Neural Networks2
1992 A Genetic Algorithm Applicable to Large-Scale Job-Shop Problems
Takeshi Yamada, Ryohei Nakano
PPSN2
1990 Translation with Optimization from Relational Calculus to Relational Algebra Having Aggregate Functions
abstract
Most of the previous translations of relational calculus to relational algebra aimed at proving that the two languages have the equivalent expressive power, thereby generating very complicated relational algebra expressions, especially when aggregate functions are introduced. This paper presents a rule-based translation method from relational calculus expressions having both aggregate functions and null values to optimized relational algebra expressions. Thus, logical optimization is carried out through translation. The translation method comprises two parts: the translational of the relational calculus kernel and the translation of aggregate functions. The former uses the familiar step-wise rewriting strategy, while the latter adopts a two-phase rewriting strategy via standard aggregate expressions. Each translation proceeds by applying a heuristic rewriting rule in preference to a basic rewriting rule. After introducing SQL-type null values, their impact on the translation is thoroughly investigated, resulting in several extensions of the translation. A translation experiment with many queries shows that the proposed translation method generates optimized relational algebra expressions. It is shown that heuristic rewriting rules play an essential role in the optimization. The correctness of the present translation is also shown. Each translation proceeds by applying a heuristic rewriting rule in preference to a basic rewriting rule. After introducing SQL-type null values, their impact on the translation is thoroughly investigated, resulting in several extensions of the translation. A translation experiment with many queries shows that the proposed translation method generates optimized relational
Ryohei Nakano
ACM Trans. Database Syst.1
1983 Integrity Checking in a Logic-Oriented ER Model
Ryohei Nakano
ER1