Faustino J. Gomez

dblp:73/5325 · DBLP profile ↗
← Back
36ranked-venue papers
10as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 9 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Deep learning architectures and training · 60% Reinforcement learning · 8% Representation and self-supervised learning · 8%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › training dynamics
neural network stability
0.312018
NAIS-Net: Stable Deep Networks from Non-Autonomous Differential Equations · NeurIPS 2018
Machine learning › Deep learning architectures and training
recurrent neural network
0.322014
A Clockwork RNN · ICML 2014
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks · ICML 2006
Machine learning › Deep learning architectures and training
sequence modeling
0.222014
A Clockwork RNN · ICML 2014
Evolino: Hybrid Neuroevolution/Optimal Linear Search for Sequence Learning · IJCAI 2005
Machine learning › Deep learning architectures and training
attention mechanism
0.212014
Deep Networks with Internal Selective Attention through Feedback Connections · NIPS 2014
Machine learning › Deep learning architectures and training
feedback loop
0.212014
Deep Networks with Internal Selective Attention through Feedback Connections · NIPS 2014
Machine learning › Deep learning architectures and training › attention mechanism
selective attention
0.212014
Deep Networks with Internal Selective Attention through Feedback Connections · NIPS 2014
Machine learning › Deep learning architectures and training › neural network training
neuroevolution
0.232008
Accelerated Neural Evolution through Cooperatively Coevolved Synapses · J. Mach. Learn. Res. 2008
Evolino: Hybrid Neuroevolution/Optimal Linear Search for Sequence Learning · IJCAI 2005
Solving Non-Markovian Control Tasks with Neuro-Evolution · IJCAI 1999
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.112012
On the Size of the Online Kernel Sparsification Dictionary · ICML 2012
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.112012
On the Size of the Online Kernel Sparsification Dictionary · ICML 2012
Machine learning › Representation and self-supervised learning › feature transformation
feature construction
0.112011
Incremental Basis Construction from Temporal Difference Error · ICML 2011
Machine learning › Reinforcement learning
temporal difference learning
0.112011
Incremental Basis Construction from Temporal Difference Error · ICML 2011
Machine learning › Reinforcement learning
value function approximation
0.112011
Incremental Basis Construction from Temporal Difference Error · ICML 2011
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.112010
Improving the Asymptotic Performance of Markov Chain Monte-Carlo by Inserting Vortices · NIPS 2010
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
non-reversible markov chain
0.112010
Improving the Asymptotic Performance of Markov Chain Monte-Carlo by Inserting Vortices · NIPS 2010
Machine learning › Optimization for machine learning › evolutionary computation
cooperative coevolution
0.112008
Accelerated Neural Evolution through Cooperatively Coevolved Synapses · J. Mach. Learn. Res. 2008
Machine learning › Optimization for machine learning
evolutionary computation
0.112008
Accelerated Neural Evolution through Cooperatively Coevolved Synapses · J. Mach. Learn. Res. 2008
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › end-to-end speech recognition
connectionist temporal classification
0.112006
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks · ICML 2006
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition
0.112006
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks · ICML 2006
Natural language and speech › Information extraction and text analysis
sequence labeling
0.112006
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks · ICML 2006
Machine learning › Deep learning architectures and training › recurrent neural network
recurrent neural network training
0.112005
Evolino: Hybrid Neuroevolution/Optimal Linear Search for Sequence Learning · IJCAI 2005
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.012013
Compete to Compute · NIPS 2013

Methods — techniques the papers use, named apart from their topics

skip connections · 0.3lipschitz input-output map · 0.3selective attention · 0.2modular hidden layer partitioning · 0.2feedback connections · 0.2clock-rate computation · 0.2gradient-based training · 0.2backpropagation · 0.2online kernel sparsification · 0.1incremental basis construction · 0.1
YearPublicationVenuePosition
2018 NAIS-Net: Stable Deep Networks from Non-Autonomous Differential Equations
abstract
This paper introduces Non-Autonomous Input-Output Stable Network (NAIS-Net), a very deep architecture where each stacked processing block is derived from a time-invariant non-autonomous dynamical system. Non-autonomy is implemented by skip connections from the block input to each of the unrolled processing stages and allows stability to be enforced so that blocks can be unrolled adaptively to a pattern-dependent processing depth. NAIS-Net induces non-trivial, Lipschitz input-output maps, even for an infinite unroll length. We prove that the network is globally asymptotically stable so that for every initial condition there is exactly one input-dependent equilibrium assuming tanh units, and multiple stable equilibria for ReL units. An efficient implementation that enforces the stability under derived conditions for both fully-connected and convolutional layers is also presented. Experimental results show how NAIS-Net exhibits stability in practice, yielding a significant reduction in generalization gap compared to ResNets.
Marco Ciccone, Marco Gallieri, Jonathan Masci, Christian Osendorfer, Faustino J. Gomez
NeurIPS5
2014 Evolving deep unsupervised convolutional networks for vision-based reinforcement learning
abstract
Dealing with high-dimensional input spaces, like visual input, is a challenging task for reinforcement learning (RL). Neuroevolution (NE), used for continuous RL problems, has to either reduce the problem dimensionality by (1) compressing the representation of the neural network controllers or (2) employing a pre-processor (compressor) that transforms the high-dimensional raw inputs into low-dimensional features. In this paper, we are able to evolve extremely small recurrent neural network (RNN) controllers for a task that previously required networks with over a million weights. The high-dimensional visual input, which the controller would normally receive, is first transformed into a compact feature vector through a deep, max-pooling convolutional neural network (MPCNN). Both the MPCNN preprocessor and the RNN controller are evolved successfully to control a car in the TORCS racing simulator using only visual input. This is the first use of deep learning in the context evolutionary RL.
Jan Koutník, Jürgen Schmidhuber, Faustino J. Gomez
GECCO3
2014 A Clockwork RNN
abstract
Sequence prediction and classification are ubiquitous and challenging problems in machine learning that can require identifying complex dependencies between temporally distant inputs. Recurrent Neural Networks (RNNs) have the ability, in theory, to cope with these temporal dependencies by virtue of the short-term memory implemented by their recurrent (feedback) connections. However, in practice they are difficult to train successfully when long-term memory is required. This paper introduces a simple, yet powerful modification to the simple RNN (SRN) architecture, the Clockwork RNN (CW-RNN), in which the hidden layer is partitioned into separate modules, each processing inputs at its own temporal granularity, making computations only at its prescribed clock rate. Rather than making the standard RNN models more complex, CW-RNN reduces the number of SRN parameters, improves the performance significantly in the tasks tested, and speeds up the network evaluation. The network is demonstrated in preliminary experiments involving three tasks: audio signal generation, TIMIT spoken word classification, where it outperforms both SRN and LSTM networks, and online handwriting recognition, where it outperforms SRNs.
Jan Koutník, Klaus Greff, Faustino J. Gomez, Jürgen Schmidhuber
ICML3
2014 Deep Networks with Internal Selective Attention through Feedback Connections
Marijn F. Stollenga, Jonathan Masci, Faustino J. Gomez, Jürgen Schmidhuber
NIPS3
2013 Evolving large-scale neural networks for vision-based TORCS
Jan Koutník, Giuseppe Cuccu, Jürgen Schmidhuber, Faustino J. Gomez
FDG4
2013 Evolving large-scale neural networks for vision-based reinforcement learning
abstract
The idea of using evolutionary computation to train artificial neural networks, or neuroevolution (NE), for reinforcement learning (RL) tasks has now been around for over 20 years. However, as RL tasks become more challenging, the networks required become larger, as do their genomes. But, scaling NE to large nets (i.e. tens of thousands of weights) is infeasible using direct encodings that map genes one-to-one to network components. In this paper, we scale-up our compressed network encoding where network weight matrices are represented indirectly as a set of Fourier-type coefficients, to tasks that require very-large networks due to the high-dimensionality of their input space. The approach is demonstrated successfully on two reinforcement learning tasks in which the control networks receive visual input: (1) a vision-based version of the octopus control task requiring networks with over 3 thousand weights, and (2) a version of the TORCS driving game where networks with over 1 million weights are evolved to drive a car around a track using video images from the driver's perspective.
Jan Koutník, Giuseppe Cuccu, Jürgen Schmidhuber, Faustino J. Gomez
GECCO4
2013 Compete to Compute
abstract
Local competition among neighboring neurons is common in biological neural networks (NNs). We apply the concept to gradient-based, backprop-trained artificial multilayer NNs. NNs with competing linear units tend to outperform those with non-competing nonlinear units, and avoid catastrophic forgetting when training sets change over time.
Rupesh Kumar Srivastava, Jonathan Masci, Sohrob Kazerounian, Faustino J. Gomez, Jürgen Schmidhuber
NIPS4
2012 On the Size of the Online Kernel Sparsification Dictionary
Yi Sun 0003, Faustino J. Gomez, Jürgen Schmidhuber
ICML2
2012 Block Diagonal Natural Evolution Strategies
Giuseppe Cuccu, Faustino J. Gomez
PPSN (2)2
2012 Compressed Network Complexity Search
Faustino J. Gomez, Jan Koutník, Jürgen Schmidhuber
PPSN (1)1
2012 Generalized Compressed Network Search
Rupesh Kumar Srivastava, Jürgen Schmidhuber, Faustino J. Gomez
PPSN (1)3
2011 Novelty-based restarts for evolution strategies
abstract
A major limitation in applying evolution strategies to black box optimization is the possibility of convergence into bad local optima. Many techniques address this problem, mostly through restarting the search. However, deciding the new start location is nontrivial since neither a good location nor a good scale for sampling a random restart position are known. A black box search algorithm can nonetheless obtain some information about this location and scale from past exploration. The method proposed here makes explicit use of such experience, through the construction of an archive of novel solutions during the run. Upon convergence, the most "novel" individual found so far is used to position the new start in the least explored region of the search space, actively looking for a new basin of attraction. We demonstrate the working principle of the method on two multi-modal test problems.
Giuseppe Cuccu, Faustino J. Gomez, Tobias Glasmachers
IEEE Congress on Evolutionary Computation2
2011 Curiosity-driven optimization
abstract
The principle of artificial curiosity directs active exploration towards the most informative or most interesting data. We show its usefulness for global black box optimization when data point evaluations are expensive. Gaussian process regression is used to model the fitness function based on all available observations so far. For each candidate point this model estimates expected Fitness reduction, and yields a novel closed-form expression of expected information gain. A new type of Pareto-front algorithm continually pushes the boundary of candidates not dominated by any other known data according to both criteria, using multi-objective evolutionary search. This makes the exploration-exploitation trade-off explicit, and permits maximally informed data selection. We illustrate the robustness of our approach in a number of experimental scenarios.
Tom Schaul, Yi Sun 0003, Daan Wierstra, Faustino J. Gomez, Jürgen Schmidhuber
IEEE Congress on Evolutionary Computation4
2011 When Novelty Is Not Enough
Giuseppe Cuccu, Faustino J. Gomez
EvoApplications (1)2
2011 Incremental Basis Construction from Temporal Difference Error
Yi Sun 0003, Faustino J. Gomez, Mark B. Ring, Jürgen Schmidhuber
ICML2
2011 Modular deep belief networks that do not forget
abstract
Deep belief networks (DBNs) are popular for learning compact representations of high-dimensional data. However, most approaches so far rely on having a single, complete training set. If the distribution of relevant features changes during subsequent training stages, the features learned in earlier stages are gradually forgotten. Often it is desirable for learning algorithms to retain what they have previously learned, even if the input distribution temporarily changes. This paper introduces the M-DBN, an unsupervised modular DBN that addresses the forgetting problem. M-DBNs are composed of a number of modules that are trained only on samples they best reconstruct. While modularization by itself does not prevent forgetting, the M-DBN additionally uses a learning method that adjusts each module's learning rate proportionally to the fraction of best reconstructed samples. On the MNIST handwritten digit dataset module specialization largely corresponds to the digits discerned by humans. Furthermore, in several learning tasks with changing MNIST digits, M-DBNs retain learned features even after those features are removed from the training data, while monolithic DBNs of comparable size forget feature mappings learned before.
Leo Pape, Faustino J. Gomez, Juergen Ring, Jürgen Schmidhuber
IJCNN2
2010 Evolving neural networks in compressed weight space
abstract
We propose a new indirect encoding scheme for neural networks in which the weight matrices are represented in the frequency domain by sets Fourier coefficients. This scheme exploits spatial regularities in the matrix to reduce the dimensionality of the representation by ignoring high-frequency coefficients, as is done in lossy image compression. We compare the efficiency of searching in this "compressed" network space to searching in the space of directly encoded networks, using the CoSyNE neuroevolution algorithm on three benchmark problems: pole-balancing, ball throwing and octopus arm control. The results show that this encoding can dramatically reduce the search space dimensionality such that solutions can be found in significantly fewer evaluations
Jan Koutník, Faustino J. Gomez, Jürgen Schmidhuber
GECCO2
2010 Improving the Asymptotic Performance of Markov Chain Monte-Carlo by Inserting Vortices
abstract
We present a new way of converting a reversible finite Markov chain into a nonreversible one, with a theoretical guarantee that the asymptotic variance of the MCMC estimator based on the non-reversible chain is reduced. The method is applicable to any reversible chain whose states are not connected through a tree, and can be interpreted graphically as inserting vortices into the state transition graph. Our result confirms that non-reversible chains are fundamentally better than reversible ones in terms of asymptotic performance, and suggests interesting directions for further improving MCMC.
Yi Sun 0003, Faustino J. Gomez, Jürgen Schmidhuber
NIPS2
2009 Sustaining diversity using behavioral information distance
abstract
Conventional similarity metrics used to sustain diversity in evolving populations are not well suited to sequential decision tasks. Genotypes and phenotypic structure are poor predictors of how solutions will actually behave in the environment. In this paper, we propose measuring similarity directly on the behavioral trajectories of evolving candidate policies using a universal similarity measure based on algorithmic information theory: normalized compression distance (NCD). NCD is compared to four other similarity measures in both genotype and phenotype space on the POMDP Tartarus problem, and shown to produce the most fit, general, and complex solutions.
Faustino J. Gomez
GECCO1
2009 Measuring and Optimizing Behavioral Complexity for Evolutionary Reinforcement Learning
Faustino J. Gomez, Julian Togelius, Jürgen Schmidhuber
ICANN (2)1
2009 A reinforcement learning approach for individualizing erythropoietin dosages in hemodialysis patients
José D. Martín-Guerrero, Faustino J. Gomez, Emilio Soria-Olivas, Jürgen Schmidhuber, Mónica Climente-Martí, N. Víctor Jiménez
Expert Syst. Appl.2
2008 Learning what to ignore: Memetic climbing in topology and weight space
abstract
We present the memetic climber, a simple search algorithm that learns topology and weights of neural networks on different time scales. When applied to the problem of learning control for a simulated racing task with carefully selected inputs to the neural network, the memetic climber outperforms a standard hill-climber. When inputs to the network are less carefully selected, the difference is drastic. We also present two variations of the memetic climber and discuss the generalization of the underlying principle to population-based neuroevolution algorithms.
Julian Togelius, Faustino J. Gomez, Jürgen Schmidhuber
IEEE Congress on Evolutionary Computation2
2008 Countering Poisonous Inputs with Memetic Neuroevolution
Julian Togelius, Tom Schaul, Jürgen Schmidhuber, Faustino J. Gomez
PPSN4
2008 Accelerated Neural Evolution through Cooperatively Coevolved Synapses
Faustino J. Gomez, Jürgen Schmidhuber, Risto Miikkulainen
J. Mach. Learn. Res.1
2007 Training Recurrent Networks by Evolino
abstract
In recent years, gradient-based LSTM recurrent neural networks (RNNs) solved many previously RNN-unlearnable tasks. Sometimes, however, gradient information is of little use for training RNNs, due to numerous local minima. For such cases, we present a novel method: EVOlution of systems with LINear Outputs (Evolino). Evolino evolves weights to the nonlinear, hidden nodes of RNNs while computing optimal linear mappings from hidden state to output, using methods such as pseudo-inverse-based linear regression. If we instead use quadratic programming to maximize the margin, we obtain the first evolutionary recurrent support vector machines. We show that Evolino-based LSTM can solve tasks that Echo State nets (Jaeger, 2004a) cannot and achieves higher accuracy in certain continuous function generation tasks than conventional gradient descent RNNs, including gradient-based LSTM.
Jürgen Schmidhuber, Daan Wierstra, Matteo Gagliolo, Faustino J. Gomez
Neural Comput.4
2006 Efficient Non-linear Control Through Neuroevolution
Faustino J. Gomez, Jürgen Schmidhuber, Risto Miikkulainen
ECML1
2006 Evolino for recurrent support vector machines
Jürgen Schmidhuber, Matteo Gagliolo, Daan Wierstra, Faustino J. Gomez
ESANN4
2006 Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
abstract
Many real-world sequence learning tasks require the prediction of sequences of labels from noisy, unsegmented input data. In speech recognition, for example, an acoustic signal is transcribed into words or sub-word units. Recurrent neural networks (RNNs) are powerful sequence learners that would seem well suited to such tasks. However, because they require pre-segmented training data, and post-processing to transform their outputs into label sequences, their applicability has so far been limited. This paper presents a novel method for training RNNs to label unsegmented sequences directly, thereby solving both problems. An experiment on the TIMIT speech corpus demonstrates its advantages over both a baseline HMM and a hybrid HMM-RNN.
Alex Graves, Santiago Fernández, Faustino J. Gomez, Jürgen Schmidhuber
ICML3
2006 A System for Robotic Heart Surgery that Learns to Tie Knots Using Recurrent Neural Networks
abstract
Tying suture knots is a time-consuming task performed frequently during minimally invasive surgery (MIS). Automating this task could greatly reduce total surgery time for patients. Current solutions to this problem replay manually programmed trajectories, but a more general and robust approach is to use supervised machine learning to smooth surgeon-given training trajectories and generalize from them. Since knottying generally requires a controller with internal memory to distinguish between identical inputs that require different actions at different points along a trajectory, it would be impossible to teach the system using traditional feedforward neural nets or support vector machines. Instead we exploit more powerful, recurrent neural networks (RNNs) with adaptive internal states. Results obtained using LSTM RNNs trained by the recent Evolino algorithm show that this approach can significantly increase the efficiency of suture knot tying in MIS over preprogrammed control
Hermann Georg Mayer, Faustino J. Gomez, Daan Wierstra, Istvan Nagy 0002, Alois C. Knoll, Jürgen Schmidhuber
IROS2
2005 Co-evolving recurrent neurons learn deep memory POMDPs
abstract
Recurrent neural networks are theoretically capable of learning complex temporal sequences, but training them through gradient-descent is too slow and unstable for practical use in reinforcement learning environments. Neuroevolution, the evolution of artificial neural networks using genetic algorithms, can potentially solve real-world reinforcement learning tasks that require deep use of memory, i.e. memory spanning hundreds or thousands of inputs, by searching the space of recurrent neural networks directly. In this paper, we introduce a new neuroevolution algorithm called Hierarchical Enforced SubPopulations that simultaneously evolves networks at two levels of granularity: full networks and network components or neurons. We demonstrate the method in two POMDP tasks that involve temporal dependencies of up to thousands of time-steps, and show that it is faster and simpler than the current best conventional reinforcement learning system on these tasks.
Faustino J. Gomez, Jürgen Schmidhuber
GECCO1
2005 Modeling systems with internal state using evolino
abstract
Existing Recurrent Neural Networks (RNNs) are limited in their ability to model dynamical systems with nonlinearities and hidden internal states. Here we use our general framework for sequence learning, EVOlution of recurrent systems with LINear Outputs (Evolino), to discover good RNN hidden node weights through evolution, while using linear regression to compute an optimal linear mapping from hidden state to output. Using the Long Short-Term Memory RNN Architecture, Evolino outperforms previous state-of-the-art methods on several tasks: 1) context-sensitive languages, 2) multiple superimposed sine waves. Categories and Subject Descriptors I.2.6 [Artificial Intelligence]: Learning—Connectionism and neural nets
Daan Wierstra, Faustino J. Gomez, Jürgen Schmidhuber
GECCO2
2005 Evolving Modular Fast-Weight Networks for Control
Faustino J. Gomez, Jürgen Schmidhuber
ICANN (2)1
2005 Evolino: Hybrid Neuroevolution/Optimal Linear Search for Sequence Learning
Jürgen Schmidhuber, Daan Wierstra, Faustino J. Gomez
IJCAI3
2004 Transfer of Neuroevolved Controllers in Unstable Domains
Faustino J. Gomez, Risto Miikkulainen
GECCO (2)1
2003 Active Guidance for a Finless Rocket Using Neuroevolution
Faustino J. Gomez, Risto Miikkulainen
GECCO1
1999 Solving Non-Markovian Control Tasks with Neuro-Evolution
Faustino J. Gomez, Risto Miikkulainen
IJCAI1