VLDB 2026 Research / reviewers in the wild / expert
Jürgen Schmidhuber
dblp:s/JurgenSchmidhuber
· DBLP profile ↗
259ranked-venue papers
34as first author
59since 2021 · last 2025
0000-0002-1468-6758ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 238 · 32 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 13 · 3 first-author · 2 since 2021Systems, architecture and hardware · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3Theory of computation · 3 · 1 since 2021Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language ModelsabstractLong-form writing agents require flexible integration and interaction across information retrieval, reasoning, and composition.Current approaches rely on predefined workflows and rigid thinking patterns to generate outlines before writing, resulting in constrained adaptability during writing.In this paper we propose WriteHERE, a general agent framework that achieves human-like adaptive writing through recursive task decomposition and dynamic integration of three fundamental task types: retrieval, reasoning, and composition.Our methodology features: 1) a planning mechanism that interleaves recursive task decomposition and execution, eliminating artificial restrictions on writing workflow; and 2) integration of task types that facilitates heterogeneous task decomposition.Evaluations on both fiction writing and technical report generation show that our method consistently outperforms state-of-the-art approaches across all automatic evaluation metrics, demonstrating the effectiveness and broad applicability of our proposed framework.We have publicly released our code and prompts to facilitate further research. Ruibin Xiong, Dmitrii Khizbullin, Mingchen Zhuge, Jürgen Schmidhuber |
EMNLP | 5 |
| 2025 | FACTS: A Factored State-Space Framework for World ModellingabstractWorld modelling is essential for understanding and predicting the dynamics of complex systems by learning both spatial and temporal dependencies. However, current frameworks, such as Transformers and selective state-space models like Mambas, exhibit limitations in efficiently encoding spatial and temporal structures, particularly in scenarios requiring long-term high-dimensional sequence modelling. To address these issues, we propose a novel recurrent framework, the FACTored State-space (FACTS) model, for spatial-temporal world modelling. The FACTS framework constructs a graph-structured memory with a routing mechanism that learns permutable memory representations, ensuring invariance to input permutations while adapting through selective state-space propagation. Furthermore, FACTS supports parallel computation of high-dimensional sequences. We empirically evaluate FACTS across diverse tasks, including multivariate time series forecasting, object-centric world modelling, and spatial-temporal graph prediction, demonstrating that it consistently outperforms or matches specialised state-of-the-art models, despite its general-purpose world modelling design. Nanbo Li, Firas Laakom, Jürgen Schmidhuber |
ICLR | 5 |
| 2025 | Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term PlanningabstractThe Value Iteration Network (VIN) is an end-to-end differentiable neural network architecture for planning. It exhibits strong generalization to unseen domains by incorporating a differentiable planning module that operates on a latent Markov Decision Process (MDP). However, VINs struggle to scale to long-term and large-scale planning tasks, such as navigating a $100\times 100$ maze---a task that typically requires thousands of planning steps to solve. We observe that this deficiency is due to two issues: the representation capacity of the latent MDP and the planning module's depth. We address these by augmenting the latent MDP with a dynamic transition kernel, dramatically improving its representational capacity, and, to mitigate the vanishing gradient problem, introduce an "adaptive highway loss" that constructs skip connections to improve gradient flow. We evaluate our method on 2D/3D maze navigation environments, continuous control, and the real-world Lunar rover navigation task. We find that our new method, named Dynamic Transition VIN (DT-VIN), scales to 5000 layers and solves challenging versions of the above tasks. Altogether, we believe that DT-VIN represents a concrete step forward in performing long-term large-scale planning in complex environments. Yuhui Wang 0004, Qingyuan Wu, Dylan R. Ashley, Francesco Faccio, Weida Li, Chao Huang 0015, Jürgen Schmidhuber |
ICML | 7 |
| 2025 | Measuring In-Context Computation Complexity via Hidden State PredictionabstractDetecting when a neural sequence model does "interesting" computation is an open problem. The next token prediction loss is a poor indicator: Low loss can stem from trivially predictable sequences that are uninteresting, while high loss may reflect unpredictable but also irrelevant information that can be ignored by the model. We propose a better metric: measuring the model’s ability to predict its own future hidden states. We show empirically that this metric–in contrast to the next token prediction loss–correlates with the intuitive interestingness of the task. To measure predictability, we introduce the architecture-agnostic "prediction of hidden states" (PHi) layer that serves as an information bottleneck on the main pathway of the network (e.g., the residual stream in Transformers). We propose a novel learned predictive prior that enables us to measure the novel information gained in each computation step, which serves as our metric. We show empirically that our metric predicts the description length of formal languages learned in-context, the complexity of mathematical reasoning problems, and the correctness of self-generated reasoning chains. Vincent Herrmann, Róbert Csordás, Jürgen Schmidhuber |
ICML | 3 |
| 2025 | Fairness Overfitting in Machine Learning: An Information-Theoretic PerspectiveabstractDespite substantial progress in promoting fairness in high-stake applications using machine learning models, existing methods often modify the training process, such as through regularizers or other interventions, but lack formal guarantees that fairness achieved during training will generalize to unseen data. Although overfitting with respect to prediction performance has been extensively studied, overfitting in terms of fairness loss has received far less attention. This paper proposes a theoretical framework for analyzing fairness generalization error through an information-theoretic lens. Our novel bounding technique is based on Efron–Stein inequality, which allows us to derive tight information-theoretic fairness generalization bounds with both Mutual Information (MI) and Conditional Mutual Information (CMI). Our empirical results validate the tightness and practical relevance of these bounds across diverse fairness-aware learning algorithms.
Our framework offers valuable insights to guide the design of algorithms improving fairness generalization. Firas Laakom, Haobo Chen, Jürgen Schmidhuber, Yuheng Bu |
ICML | 3 |
| 2025 | Directly Forecasting Belief for Reinforcement Learning with DelaysabstractReinforcement learning (RL) with delays is challenging as sensory perceptions lag behind the actual events: the RL agent needs to estimate the real state of its environment based on past observations. State-of-the-art (SOTA) methods typically employ recursive, step-by-step forecasting of states. This can cause the accumulation of compounding errors. To tackle this problem, our novel belief estimation method, named Directly Forecasting Belief Transformer (DFBT), directly forecasts states from observations without incrementally estimating intermediate states step-by-step. We theoretically demonstrate that DFBT greatly reduces compounding errors of existing recursively forecasting methods, yielding stronger performance guarantees. In experiments with D4RL offline datasets, DFBT reduces compounding errors with remarkable prediction accuracy. DFBT’s capability to forecast state sequences also facilitates multi-step bootstrapping, thus greatly improving learning efficiency. On the MuJoCo benchmark, our DFBT-based method substantially outperforms SOTA baselines. Code is available at https://github.com/QingyuanWuNothing/DFBT. Qingyuan Wu, Yuhui Wang 0004, Simon Sinong Zhan, Yixuan Wang 0001, Chung-Wei Lin, Chen Lv 0001, Qi Zhu 0002, Jürgen Schmidhuber, Chao Huang 0015 |
ICML | 8 |
| 2025 | Agent-as-a-Judge: Evaluate Agents with AgentsabstractContemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes—ignoring the step-by-step nature of the thinking done by agentic systems—or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is a natural extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire task-solving processes for more precise evaluations. We apply the Agent-as-a-Judge framework to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic AI code generation tasks. DevAI includes rich manual annotations, like a total of 365 hierarchical solution requirements, which make it particularly suitable for an agentic evaluator. We benchmark three of the top code-generating agentic systems using Agent-as-a-Judge and find that our framework dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that this work represents a concrete step towards enabling vastly more sophisticated agentic systems. To help that, our dataset and the full implementation of Agent-as-a-Judge will be publically available at https://github.com/metauto-ai/agent-as-a-judge Mingchen Zhuge, Changsheng Zhao 0002, Dylan R. Ashley, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, Jürgen Schmidhuber |
ICML | 13 |
| 2025 | Towards an Extremely Robust Baby Robot With Rich Interaction Ability for Advanced Machine Learning AlgorithmsabstractAdvanced machine learning algorithms require platforms that are extremely robust and equipped with rich sensory feedback to handle extensive trial-and-error learning without relying on overwhelming inductive biases. Traditional robotic designs, while well-suited for their specific use cases, are often fragile when used with these algorithms as they fail to address the intermediate sub-optimal posterior-based behavior these algorithms exhibit. To address this gap—and inspired by the vision of enabling curiosity-driven baby robots—we present a novel robotic limb designed from scratch. Our design features a semi-soft structure, a high degree of redundancy achieved through rich non-contact sensors (exclusively cameras), and strategically designed, easily replaceable failure points. Proof-of-concept experiments using two contemporary reinforcement learning algorithms on a physical prototype demonstrate that our design is able to succeed in a target-finding task even under simulated sensor failures, all with minimal human oversight during extended learning periods. Additional experiments on the robustness of the design show that it is able to withstand relatively large amounts of mechanical stress. We believe this design represents a concrete step toward more tailored robotic designs capable of supporting general-purpose, generally intelligent robots. Mohannad Alhakami, Dylan R. Ashley, Joel Dunham, Yanning Dai, Francesco Faccio, Eric Feron, Jürgen Schmidhuber |
IROS | 7 |
| 2025 | PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled PriorsabstractEvaluating the scientific discovery capabilities of large language model based agents, particularly how they cope with varying environmental complexity and utilize prior knowledge, requires specialized benchmarks currently lacking in the landscape. To address this gap, we introduce PhysGym, a novel benchmark suite and simulation platform for rigorously assessing LLM-based scientific reasoning in interactive physics environments. PhysGym's primary contribution lies in its sophisticated control over the level of prior knowledge provided to the agent. This allows researchers to dissect agent performance along axes including the complexity of the problem and the prior knowledge levels. The benchmark comprises a suite of interactive simulations, where agents must actively probe environments, gather data sequentially under constraints and formulate hypotheses about underlying physical laws. PhysGym provides standardized evaluation protocols and metrics for assessing hypothesis accuracy and model fidelity. We demonstrate the benchmark's utility by presenting results from baseline LLMs, showcasing its ability to differentiate capabilities based on varying priors and task complexity. Piotr Piekos, Mateusz Ostaszewski, Firas Laakom, Jürgen Schmidhuber |
NeurIPS | 5 |
| 2025 | Curious Causality-Seeking Agents in Open-ended WorldsabstractWhen building a world model, a common assumption is that the environment has a single, unchanging underlying causal rule, like applying Newton's laws to every situation. However, in truly open-ended environments, the apparent causal mechanism may drift over time because the agent continually encounters novel contexts and operates within a limited observational window. This brings about a problem that, when building a world model, even subtle shifts in policy or environment states can alter the very observed causal mechanisms.
In this work, we introduce the Meta-Causal Graph as world models for open-ended environments, a minimal unified representation that efficiently encodes the transformation rules governing how causal structures shift across different latent world states. A single Meta-Causal Graph is composed of multiple causal subgraphs, each triggered by meta state, which is in the latent state space. Building on this representation, we introduce a Causality-Seeking Agent whose objectives are to (1) identify the meta states that trigger each subgraph, (2) discover the corresponding causal relationships by agent curiosity-driven intervention policy, and (3) iteratively refine the Meta-Causal Graph through ongoing curiosity-driven exploration and agent experiences. Experiments on both synthetic tasks and a challenging robot arm manipulation task demonstrate that our method robustly captures shifts in causal dynamics and generalizes effectively to previously unseen contexts. Haoxuan Li 0001, Haifeng Zhang 0002, Jun Wang 0012, Francesco Faccio, Jürgen Schmidhuber, Mengyue Yang |
NeurIPS | 6 |
| 2025 | Falling Walls, WWW, Modern AI, and the Future of the UniverseabstractAround 1990, the Berlin Wall came down, the WWW was born at CERN, mobile phones became popular, self-driving cars appeared in traffic, and modern AI based on very deep artificial neural networks emerged, including the principles behind the G, P, and T in ChatGPT. I place these events in the history of the universe since the Big Bang, and discuss what's next: not just AI behind the screen in the virtual world, but real AI for real robots in the real world, connected through a WWW of machines. Intelligent (but not necessarily super-intelligent) robots that can learn to operate the tools and machines operated by humans can also build (and repair when needed) more of their own kind. This will culminate in life-like, self-replicating and self-improving machine civilisations, which represent the ultimate form of upscaling, and will shape the long-term future of the entire cosmos. The wonderful short-term side effect is that our AI will continue to make people's lives longer, healthier and easier. Jürgen Schmidhuber |
WWW | 1 |
| 2025 | Mindstorms in Natural Language-Based Societies of MindabstractInspired by Minsky's Society of Mind, Schmidhuber's Learning to Think, and other more recent works, this paper proposes and advocates for the concept of natural language-based societies of mind (NLSOMs). We imagine these societies as consisting of a collection of multimodal neural networks, including large language models, which engage in a “mindstorm” to solve problems using a shared natural language interface. Here, we work to identify and discuss key questions about the social structure, governance, and economic principles for NLSOMs, emphasizing their impact on the future of AI. Our demonstrations with NLSOMs-which feature up to 129 agents-show their effectiveness in various tasks, including visual question answering, image captioning, and prompt generation for text-to-image synthesis. Mingchen Zhuge, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li 0024, Guohao Li 0001, Shuming Liu 0001, Jinjie Mai, Piotr Piekos, Aditya A. Ramesh, Imanol Schlag, Aleksandar Stanic, Yuhui Wang 0004, Mengmeng Xu 0006, Deng-Ping Fan, Bernard Ghanem, Jürgen Schmidhuber |
Comput. Vis. Media | 26 |
| 2025 | On the Distillation of Stories for Transferring Narrative Arcs in Collections of Independent MediaabstractThe act of telling stories is a fundamental part of what it means to be human. This work introduces the concept of narrative information, which we define as the overlap in information space between a story and the items that compose the story. Using contrastive learning methods, we show how modern artificial neural networks can be leveraged to distill stories and extract a representation of the narrative information. We then demonstrate how evolutionary algorithms can leverage this to extract a set of narrative template curves and how these-in tandem with a novel curve-fitting algorithm we introduce-can reorder music albums to automatically induce stories in them. In doing so, we give statistically significant evidence that (1) these narrative information template curves are present in existing albums and that (2) people prefer an album ordered through one of these learned template curves over a random one. The premises of our work extend to any form of (largely) independent media, and as evidence, we also show that our method works with image data. Dylan R. Ashley, Vincent Herrmann, Zachary Friggstad, Jürgen Schmidhuber |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Tune-an-Ellipse: CLIP Has Potential to Find what you WantabstractVisual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle, commonly used for highlighting, can guide CLIP's attention to the surrounding region, to identify specific objects within an image. Without precise object proposals, however, it is insufficient for localization. Our novel, simple yet effective approach, i.e., Differentiable Visual Prompting, enables CLIP to zero-shot localize: given an image and a text prompt describing an object, we first pick a rendered ellipse from uniformly distributed anchor ellipses on the image grid via visual prompting, then use three loss functions to tune the ellipse coefficients to encap-sulate the target region gradually. This yields promising ex-perimental results for referring expression comprehension without precisely specified object proposals. In addition, we systematically present the limitations of visual prompting inherent in CLIP and discuss potential solutions. Jinheng Xie, Songhe Deng, Bing Li 0024, Yawen Huang, Yefeng Zheng 0001, Jürgen Schmidhuber, Bernard Ghanem, LinLin Shen, Zheng Shou 0001 |
CVPR | 7 |
| 2024 | Past, Present, Future, and Far Future of AI
Jürgen Schmidhuber |
DATA | 1 |
| 2024 | Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Mingchen Zhuge, Jian Ding 0001, Deyao Zhu, Jürgen Schmidhuber, Mohamed Elhoseiny 0001 |
ECCV (29) | 8 |
| 2024 | Self-organising Neural Discrete Representation Learning à la Kohonen
Kazuki Irie, Róbert Csordás, Jürgen Schmidhuber |
ICANN (1) | 3 |
| 2024 | MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkabstractRecently, remarkable progress has been made on automated problem solving through societies of agents based on large language models (LLMs). Previous LLM-based multi-agent systems can already solve simple dialogue tasks. More complex tasks, however, face challenges through logic inconsistencies due to cascading hallucinations caused by naively chaining LLMs. Here we introduce MetaGPT, an innovative meta-programming framework incorporating efficient human workflows into LLM-based multi-agent collaborations. MetaGPT encodes Standardized Operating Procedures (SOPs) into prompt sequences for more streamlined workflows, thus allowing agents with human-like domain expertise to verify intermediate results and reduce errors. MetaGPT utilizes an assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together. On collaborative software engineering benchmarks, MetaGPT generates more coherent solutions than previous chat-based multi-agent systems. Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu 0001, Jürgen Schmidhuber |
ICLR | 15 |
| 2024 | Exploring the Promise and Limits of Real-Time Recurrent LearningabstractReal-time recurrent learning (RTRL) for sequence-processing recurrent neural networks (RNNs) offers certain conceptual advantages over backpropagation through time (BPTT). RTRL requires neither caching past activations nor truncating context, and enables online learning. However, RTRL's time and space complexity make it impractical. To overcome this problem, most recent work on RTRL focuses on approximation theories, while experiments are often limited to diagnostic settings. Here we explore the practical promise of RTRL in more realistic settings. We study actor-critic methods that combine RTRL and policy gradients, and test them in several subsets of DMLab-30, ProcGen, and Atari-2600 environments. On DMLab memory tasks, our system trained on fewer than 1.2B environmental frames is competitive with or outperforms well-known IMPALA and R2D2 baselines trained on 10B frames. To scale to such challenging tasks, we focus on certain well-known neural architectures with element-wise recurrence, allowing for tractable RTRL without approximation. Importantly, we also discuss rarely addressed limitations of RTRL in real-world applications, such as its complexity in the multi-layer case. Kazuki Irie, Anand Gopalakrishnan, Jürgen Schmidhuber |
ICLR | 3 |
| 2024 | Learning Useful Representations of Recurrent Neural Network Weight MatricesabstractRecurrent Neural Networks (RNNs) are general-purpose parallel-sequential computers. The program of an RNN is its weight matrix. How to learn useful representations of RNN weights that facilitate RNN analysis as well as downstream tasks? While the _mechanistic approach_ directly looks at some RNN's weights to predict its behavior, the _functionalist approach_ analyzes its overall functionality–specifically, its input-output mapping. We consider several mechanistic approaches for RNN weights and adapt the permutation equivariant Deep Weight Space layer for RNNs. Our two novel functionalist approaches extract information from RNN weights by 'interrogating' the RNN through probing inputs. We develop a theoretical framework that demonstrates conditions under which the functionalist approach can generate rich representations that help determine RNN behavior. We create and release the first two 'model zoo' datasets for RNN weight representation learning. One consists of generative models of a class of formal languages, and the other one of classifiers of sequentially processed MNIST digits. With the help of an emulation-based self-supervised learning technique we compare and evaluate the different RNN weight encoding techniques on multiple downstream applications. On the most challenging one, namely predicting which exact task the RNN was trained on, functionalist approaches show clear superiority. Vincent Herrmann, Francesco Faccio, Jürgen Schmidhuber |
ICML | 3 |
| 2024 | Sequence Compression Speeds Up Credit Assignment in Reinforcement LearningabstractTemporal credit assignment in reinforcement learning is challenging due to delayed and stochastic outcomes. Monte Carlo targets can bridge long delays between action and consequence but lead to high-variance targets due to stochasticity. Temporal difference (TD) learning uses bootstrapping to overcome variance but introduces a bias that can only be corrected through many iterations. TD($\lambda$) provides a mechanism to navigate this bias-variance tradeoff smoothly. Appropriately selecting $\lambda$ can significantly improve performance. Here, we propose Chunked-TD, which uses predicted probabilities of transitions from a model for computing $\lambda$-return targets. Unlike other model-based solutions to credit assignment, Chunked-TD is less vulnerable to model inaccuracies. Our approach is motivated by the principle of history compression and ‘chunks’ trajectories for conventional TD learning. Chunking with learned world models compresses near-deterministic regions of the environment-policy interaction to speed up credit assignment while still bootstrapping when necessary. We propose algorithms that can be implemented online and show that they solve some problems much faster than conventional TD($\lambda$). Aditya A. Ramesh, Kenny Young, Louis Kirsch, Jürgen Schmidhuber |
ICML | 4 |
| 2024 | Highway Value Iteration NetworksabstractValue iteration networks (VINs) enable end-to-end learning for planning tasks by employing a differentiable "planning module" that approximates the value iteration algorithm. However, long-term planning remains a challenge because training very deep VINs is difficult. To address this problem, we embed highway value iteration—a recent algorithm designed to facilitate long-term credit assignment—into the structure of VINs. This improvement augments the "planning module" of the VIN with three additional components: 1) an "aggregate gate," which constructs skip connections to improve information flow across many layers; 2) an "exploration module," crafted to increase the diversity of information and gradient flow in spatial dimensions; 3) a "filter gate" designed to ensure safe exploration. The resulting novel highway VIN can be trained effectively with hundreds of layers using standard backpropagation. In long-term planning tasks requiring hundreds of planning steps, deep highway VINs outperform both traditional VINs and several advanced, very deep NNs. Yuhui Wang 0004, Weida Li, Francesco Faccio, Qingyuan Wu, Jürgen Schmidhuber |
ICML | 5 |
| 2024 | Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short DelaysabstractReinforcement learning (RL) is challenging in the common case of delays between events and their sensory perceptions. State-of-the-art (SOTA) state augmentation techniques either suffer from state space explosion or performance degeneration in stochastic environments. To address these challenges, we present a novel *Auxiliary-Delayed Reinforcement Learning (AD-RL)* method that leverages auxiliary tasks involving short delays to accelerate RL with long delays, without compromising performance in stochastic environments. Specifically, AD-RL learns a value function for short delays and uses bootstrapping and policy improvement techniques to adjust it for long delays. We theoretically show that this can greatly reduce the sample complexity. On deterministic and stochastic benchmarks, our method significantly outperforms the SOTAs in both sample efficiency and policy performance. Code is available at https://github.com/QingyuanWuNothing/AD-RL. Qingyuan Wu, Simon Sinong Zhan, Yixuan Wang 0001, Yuhui Wang 0004, Chung-Wei Lin, Chen Lv 0001, Qi Zhu 0002, Jürgen Schmidhuber, Chao Huang 0015 |
ICML | 8 |
| 2024 | GPTSwarm: Language Agents as Optimizable GraphsabstractVarious human-designed prompt engineering techniques have been proposed to improve problem solvers based on Large Language Models (LLMs), yielding many disparate code bases. We unify these approaches by describing LLM-based agents as computational graphs. The nodes implement functions to process multimodal data or query LLMs, and the edges describe the information flow between operations. Graphs can be recursively combined into larger composite graphs representing hierarchies of inter-agent collaboration (where edges connect operations of different agents). Our novel automatic graph optimizers (1) refine node-level LLM prompts (node optimization) and (2) improve agent orchestration by changing graph connectivity (edge optimization). Experiments demonstrate that our framework can be used to efficiently develop, integrate, and automatically improve various LLM agents. Our code is public. Mingchen Zhuge, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, Jürgen Schmidhuber |
ICML | 6 |
| 2024 | Utilizing a Malfunctioning 3D Printer by Modeling Its Dynamics with Machine LearningabstractTo create a self-repairing 3D printer, it must continue operating even after experiencing corruption. This work focuses on developing a method to effectively utilize a malfunctioning printer for reliable printing. This method can be applied by the printer itself for self-repair and enhance the reliability of commercial 3D printers. We achieve this by modeling the dynamics of the corrupted printer using a machine learning model that by observing one trajectory infers the corrupted printer dynamics to improve its accuracy. Our method is evaluated on a digital twin of the 3D printer, demonstrating its capability to enable the printer to operate reliably, even when encountering new corruptions not encountered during training. The scripts are public on https://github.com/piotrpiekos/adaptive-printer. Renzo Caballero, Piotr Piekos, Eric Feron, Jürgen Schmidhuber |
ICRA | 4 |
| 2024 | Past, Present, Future, and Far Future of AI
Jürgen Schmidhuber |
ICSOFT | 1 |
| 2024 | MoEUT: Mixture-of-Experts Universal TransformersabstractPrevious work on Universal Transformers (UTs) has demonstrated the importance of parameter sharing across layers. By allowing recurrence in depth, UTs have advantages over standard Transformers in learning compositional generalizations, but layer-sharing comes with a practical limitation of parameter-compute ratio: it drastically reduces the parameter count compared to the non-shared model with the same dimensionality. Naively scaling up the layer size to compensate for the loss of parameters makes its computational resource requirements prohibitive. In practice, no previous work has succeeded in proposing a shared-layer Transformer design that is competitive in parameter count-dominated tasks such as language modeling. Here we propose MoEUT (pronounced "moot"), an effective mixture-of-experts (MoE)-based shared-layer Transformer architecture, which combines several recent advances in MoEs for both feedforward and attention layers of standard Transformers together with novel layer-normalization and grouping schemes that are specific and crucial to UTs. The resulting UT model, for the first time, slightly outperforms standard Transformers on language modeling tasks such as BLiMP and PIQA, while using significantly less compute and memory. Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, Christopher D. Manning |
NeurIPS | 3 |
| 2024 | SwitchHead: Accelerating Transformers with Mixture-of-Experts AttentionabstractDespite many recent works on Mixture of Experts (MoEs) for resource-efficient Transformer language models, existing methods mostly focus on MoEs for feedforward layers. Previous attempts at extending MoE to the self-attention layer fail to match the performance of the parameter-matched baseline. Our novel SwitchHead is an effective MoE method for the attention layer that successfully reduces both the compute and memory requirements, achieving wall-clock speedup, while matching the language modeling performance of the baseline Transformer. Our novel MoE mechanism allows SwitchHead to compute up to 8 times fewer attention matrices than the standard Transformer. SwitchHead can also be combined with MoE feedforward layers, resulting in fully-MoE "SwitchAll" Transformers. For our 262M parameter model trained on C4, SwitchHead matches the perplexity of standard models with only 44% compute and 27% memory usage. Zero-shot experiments on downstream tasks confirm the performance of SwitchHead, e.g., achieving more than 3.5% absolute improvements on BliMP compared to the baseline with an equal compute resource. Róbert Csordás, Piotr Piekos, Kazuki Irie, Jürgen Schmidhuber |
NeurIPS | 4 |
| 2024 | Recurrent Complex-Weighted Autoencoders for Unsupervised Object DiscoveryabstractCurrent state-of-the-art synchrony-based models encode object bindings with complex-valued activations and compute with real-valued weights in feedforward architectures. We argue for the computational advantages of a recurrent architecture with complex-valued weights. We propose a fully convolutional autoencoder, SynCx, that performs iterative constraint satisfaction: at each iteration, a hidden layer bottleneck encodes statistically regular configurations of features in particular phase relationships; over iterations, local constraints propagate and the model converges to a globally consistent configuration of phase assignments. Binding is achieved simply by the matrix-vector product operation between complex-valued weights and activations, without the need for additional mechanisms that have been incorporated into current synchrony-based models. SynCx outperforms or is strongly competitive with current models for unsupervised object discovery. SynCx also avoids certain systematic grouping errors of current models, such as the inability to separate similarly colored objects without additional supervision. Anand Gopalakrishnan, Aleksandar Stanic, Jürgen Schmidhuber, Michael C. Mozer |
NeurIPS | 3 |
| 2024 | Learning to Generalize With Object-Centric Agents in the Open World Survival Game CrafterabstractReinforcement learning agents must generalize beyond their training experience. Prior work has focused mostly on identical training and evaluation environments. Starting from the recently introducedCrafterbenchmark, a 2-D open world survival game, we introduce a new set of environments suitable for evaluating some agent's ability to generalize on previously unseen (numbers of) objects and to adapt quickly (meta-learning). InCrafter, the agents are evaluated by the number of unlocked achievements (such as collecting resources) when trained for 1 M steps. We show that current agents struggle to generalize, and introduce novel object-centric agents that improve over strong baselines. We also provide critical insights of general interest for future work onCrafterthrough several experiments. We show that careful hyperparameter tuning improves the PPO baseline agent by a large margin and that even feedforward agents can unlock almost all achievements by relying on the inventory display. We achieve a new state-of-the-art performance on the originalCrafterenvironment. In addtion, when trained beyond 1 M steps, our tuned agents can unlock almost all achievements. We show that the recurrent PPO agents improve over feedforward ones, even with the inventory information removed. We introduceCrafterOOD, a set of 15 new environments that evaluate OOD generalization. OnCrafterOOD, we show that the current agents fail to generalize, whereas our novel object-centric agents achieve state-of-the-art OOD generalization while also being interpretable. Our code is public. Aleksandar Stanic, Yujin Tang, David Ha, Jürgen Schmidhuber |
IEEE Trans. Games | 4 |
| 2023 | Goal-Conditioned Generators of Deep PoliciesabstractGoal-conditioned Reinforcement Learning (RL) aims at learning optimal policies, given goals encoded in special command inputs. Here we study goal-conditioned neural nets (NNs) that learn to generate deep NN policies in form of context-specific weight matrices, similar to Fast Weight Programmers and other methods from the 1990s. Using context commands of the form ``generate a policy that achieves a desired expected return,'' our NN generators combine powerful exploration of parameter space with generalization across commands to iteratively find better and better policies. A form of weight-sharing HyperNetworks and policy embeddings scales our method to generate deep NNs. Experiments show how a single learned policy generator can produce policies that achieve any return seen during training. Finally, we evaluate our algorithm on a set of continuous control tasks where it exhibits competitive performance. Our code is public. Francesco Faccio, Vincent Herrmann, Aditya A. Ramesh, Louis Kirsch, Jürgen Schmidhuber |
AAAI | 5 |
| 2023 | Practical Computational Power of Linear Transformers and Their Recurrent and Self-Referential ExtensionsabstractRecent studies of the computational power of recurrent neural networks (RNNs) reveal a hierarchy of RNN architectures, given real-time and finite-precision assumptions.Here we study auto-regressive Transformers with linearised attention, a.k.a.linear Transformers (LTs) or Fast Weight Programmers (FWPs).LTs are special in the sense that they are equivalent to RNN-like sequence processors with a fixed-size state, while they can also be expressed as the now-popular self-attention networks.We show that many well-known results for the standard Transformer directly transfer to LTs/FWPs.Our formal language recognition experiments demonstrate how recently proposed FWP extensions such as recurrent FWPs and self-referential weight matrices successfully overcome certain limitations of the LT, e.g., allowing for generalisation on the parity problem.Our code is public.1 Kazuki Irie, Róbert Csordás, Jürgen Schmidhuber |
EMNLP | 3 |
| 2023 | Learning to Identify Critical States for Reinforcement Learning from VideosabstractRecent work on deep reinforcement learning (DRL) has pointed out that algorithmic information about good policies can be extracted from offline data which lack explicit information about executed actions [45], [46], [30]. For example, videos of humans or robots may convey a lot of implicit information about rewarding action sequences, but a DRL machine that wants to profit from watching such videos must first learn by itself to identify and recognize relevant states/actions/rewards. Without relying on ground-truth annotations, our new method called Deep State Identifier learns to predict returns from episodes encoded as videos. Then it uses a kind of mask-based sensitivity analysis to extract/identify important critical states. Extensive experiments showcase our method’s potential for understanding and improving agent behavior. The source code and the generated datasets are available at https://github.com/AI-Initiative-KAUST/VideoRLCS. Mingchen Zhuge, Bing Li 0024, Yuhui Wang 0004, Francesco Faccio, Bernard Ghanem, Jürgen Schmidhuber |
ICCV | 7 |
| 2023 | Images as Weight Matrices: Sequential Image Generation Through Synaptic Learning Rules
Kazuki Irie, Jürgen Schmidhuber |
ICLR | 2 |
| 2023 | The Benefits of Model-Based Generalization in Reinforcement LearningabstractModel-Based Reinforcement Learning (RL) is widely believed to have the potential to improve sample efficiency by allowing an agent to synthesize large amounts of imagined experience. Experience Replay (ER) can be considered a simple kind of model, which has proved effective at improving the stability and efficiency of deep RL. In principle, a learned parametric model could improve on ER by generalizing from real experience to augment the dataset with additional plausible experience. However, given that learned value functions can also generalize, it is not immediately obvious why model generalization should be better. Here, we provide theoretical and empirical insight into when, and how, we can expect data generated by a learned model to be useful. First, we provide a simple theorem motivating how learning a model as an intermediate step can narrow down the set of possible value functions more than learning a value function directly from data using the Bellman equation. Second, we provide an illustrative example showing empirically how a similar effect occurs in a more concrete setting with neural network function approximation. Finally, we provide extensive experiments showing the benefit of model-based learning for online RL in environments with combinatorial complexity, but factored structure that allows a learned model to generalize. In these experiments, we take care to control for other factors in order to isolate, insofar as possible, the benefit of using experience generated by a learned model relative to ER alone. Kenny Young, Aditya A. Ramesh, Louis Kirsch, Jürgen Schmidhuber |
ICML | 4 |
| 2023 | Contrastive Training of Complex-Valued Autoencoders for Object DiscoveryabstractCurrent state-of-the-art object-centric models use slots and attention-based routing for binding. However, this class of models has several conceptual limitations: the number of slots is hardwired; all slots have equal capacity; training has high computational cost; there are no object-level relational factors within slots. Synchrony-based models in principle can address these limitations by using complex-valued activations which store binding information in their phase components. However, working examples of such synchrony-based models have been developed only very recently, and are still limited to toy grayscale datasets and simultaneous storage of less than three objects in practice. Here we introduce architectural modifications and a novel contrastive learning method that greatly improve the state-of-the-art synchrony-based model. For the first time, we obtain a class of synchrony-based models capable of discovering objects in an unsupervised manner in multi-object color datasets and simultaneously representing more than three objects. Aleksandar Stanic, Anand Gopalakrishnan, Kazuki Irie, Jürgen Schmidhuber |
NeurIPS | 4 |
| 2023 | Unsupervised Learning of Temporal Abstractions With Slot-Based TransformersabstractThe discovery of reusable subroutines simplifies decision making and planning in complex reinforcement learning problems. Previous approaches propose to learn such temporal abstractions in an unsupervised fashion through observing state-action trajectories gathered from executing a policy. However, a current limitation is that they process each trajectory in an entirely sequential manner, which prevents them from revising earlier decisions about subroutine boundary points in light of new incoming information. In this work, we propose slot-based transformer for temporal abstraction (SloTTAr), a fully parallel approach that integrates sequence processing transformers with a slot attention module to discover subroutines in an unsupervised fashion while leveraging adaptive computation for learning about the number of such subroutines solely based on their empirical distribution. We demonstrate how SloTTAr is capable of outperforming strong baselines in terms of boundary point discovery, even for sequences containing variable amounts of subroutines, while being up to seven times faster to train on existing benchmarks. Anand Gopalakrishnan, Kazuki Irie, Jürgen Schmidhuber, Sjoerd van Steenkiste |
Neural Comput. | 3 |
| 2022 | Reward-Weighted Regression Converges to a Global OptimumabstractReward-Weighted Regression (RWR) belongs to a family of widely known iterative Reinforcement Learning algorithms based on the Expectation-Maximization framework. In this family, learning at each iteration consists of sampling a batch of trajectories using the current policy and fitting a new policy to maximize a return-weighted log-likelihood of actions. Although RWR is known to yield monotonic improvement of the policy under certain circumstances, whether and under which conditions RWR converges to the optimal policy have remained open questions. In this paper, we provide for the first time a proof that RWR converges to a global optimum when no function approximation is used, in a general compact setting. Furthermore, for the simpler case with finite state and action spaces we prove R-linear convergence of the state-value function to the optimum. Miroslav Strupl, Francesco Faccio, Dylan R. Ashley, Rupesh Kumar Srivastava, Jürgen Schmidhuber |
AAAI | 5 |
| 2022 | CTL++: Evaluating Generalization on Never-Seen Compositional Patterns of Known Functions, and Compatibility of Neural RepresentationsabstractWell-designed diagnostic tasks have played a key role in studying the failure of neural nets (NNs) to generalize systematically.Famous examples include SCAN and Compositional Table Lookup (CTL).Here we introduce CTL++, a new diagnostic dataset based on compositions of unary symbolic functions.While the original CTL is used to test length generalization or productivity, CTL++ is designed to test systematicity of NNs, that is, their capability to generalize to unseen compositions of known functions.CTL++ splits functions into groups and tests performance on group elements composed in a way not seen during training.We show that recent CTL-solving Transformer variants fail on CTL++.The simplicity of the task design allows for fine-grained control of task difficulty, as well as many insightful analyses.For example, we measure how much overlap between groups is needed by tested NNs for learning to compose.We also visualize how learned symbol representations in outputs of functions from different groups are compatible in case of success but not in case of failure.These results provide insights into failure cases reported on more complex compositions in the natural language domain.Our code is public.1 Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber |
EMNLP | 3 |
| 2022 | The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber |
ICLR | 3 |
| 2022 | The Dual Form of Neural Networks Revisited: Connecting Test Time Predictions to Training Patterns via Spotlights of AttentionabstractLinear layers in neural networks (NNs) trained by gradient descent can be expressed as a key-value memory system which stores all training datapoints and the initial weights, and produces outputs using unnormalised dot attention over the entire training experience. While this has been technically known since the 1960s, no prior work has effectively studied the operations of NNs in such a form, presumably due to prohibitive time and space complexities and impractical model sizes, all of them growing linearly with the number of training patterns which may get very large. However, this dual formulation offers a possibility of directly visualising how an NN makes use of training patterns at test time, by examining the corresponding attention weights. We conduct experiments on small scale supervised image classification tasks in single-task, multi-task, and continual learning settings, as well as language modelling, and discuss potentials and limits of this view for better understanding and interpreting how NNs exploit training patterns. Our code is public. Kazuki Irie, Róbert Csordás, Jürgen Schmidhuber |
ICML | 3 |
| 2022 | A Modern Self-Referential Weight Matrix That Learns to Modify ItselfabstractThe weight matrix (WM) of a neural network (NN) is its program. The programs of many traditional NNs are learned through gradient descent in some error function, then remain fixed. The WM of a self-referential NN, however, can keep rapidly modifying all of itself during runtime. In principle, such NNs can meta-learn to learn, and meta-meta-learn to meta-learn to learn, and so on, in the sense of recursive self-improvement. While NN architectures potentially capable of implementing such behaviour have been proposed since the ’90s, there have been few if any practical studies. Here we revisit such NNs, building upon recent successes of fast weight programmers and closely related linear Transformers. We propose a scalable self-referential WM (SRWM) that learns to use outer products and the delta update rule to modify itself. We evaluate our SRWM in supervised few-shot learning and in multi-task reinforcement learning with procedurally generated game environments. Our experiments demonstrate both practical applicability and competitive performance of the proposed SRWM. Our code is public. Kazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen Schmidhuber |
ICML | 4 |
| 2022 | Neural Differential Equations for Learning to Program Neural Nets Through Continuous Learning RulesabstractNeural ordinary differential equations (ODEs) have attracted much attention as continuous-time counterparts of deep residual neural networks (NNs), and numerous extensions for recurrent NNs have been proposed. Since the 1980s, ODEs have also been used to derive theoretical results for NN learning rules, e.g., the famous connection between Oja's rule and principal component analysis. Such rules are typically expressed as additive iterative update processes which have straightforward ODE counterparts. Here we introduce a novel combination of learning rules and Neural ODEs to build continuous-time sequence processing nets that learn to manipulate short-term memory in rapidly changing synaptic connections of other nets. This yields continuous-time counterparts of Fast Weight Programmers and linear Transformers. Our novel models outperform the best existing Neural Controlled Differential Equation based models on various time series classification tasks, while also addressing their fundamental scalability limitations. Our code is public. Kazuki Irie, Francesco Faccio, Jürgen Schmidhuber |
NeurIPS | 3 |
| 2022 | Exploring through Random Curiosity with General Value FunctionsabstractEfficient exploration in reinforcement learning is a challenging problem commonly addressed through intrinsic rewards. Recent prominent approaches are based on state novelty or variants of artificial curiosity. However, directly applying them to partially observable environments can be ineffective and lead to premature dissipation of intrinsic rewards. Here we propose random curiosity with general value functions (RC-GVF), a novel intrinsic reward function that draws upon connections between these distinct approaches. Instead of using only the current observation’s novelty or a curiosity bonus for failing to predict precise environment dynamics, RC-GVF derives intrinsic rewards through predicting temporally extended general value functions. We demonstrate that this improves exploration in a hard-exploration diabolical lock problem. Furthermore, RC-GVF significantly outperforms previous methods in the absence of ground-truth episodic counts in the partially observable MiniGrid environments. Panoramic observations on MiniGrid further boost RC-GVF's performance such that it is competitive to baselines exploiting privileged information in form of episodic counts. Aditya A. Ramesh, Louis Kirsch, Sjoerd van Steenkiste, Jürgen Schmidhuber |
NeurIPS | 4 |
| 2022 | Recurrent Neural-Linear Posterior Sampling for Nonstationary Contextual BanditsabstractAn agent in a nonstationary contextual bandit problem should balance between exploration and the exploitation of (periodic or structured) patterns present in its previous experiences. Handcrafting an appropriate historical context is an attractive alternative to transform a nonstationary problem into a stationary problem that can be solved efficiently. However, even a carefully designed historical context may introduce spurious relationships or lack a convenient representation of crucial information. In order to address these issues, we propose an approach that learns to represent the relevant context for a decision based solely on the raw history of interactions between the agent and the environment. This approach relies on a combination of features extracted by recurrent neural networks with a contextual linear bandit algorithm based on posterior sampling. Our experiments on a diverse selection of contextual and noncontextual nonstationary problems show that our recurrent approach consistently outperforms its feedforward counterpart, which requires handcrafted historical contexts, while being more widely applicable than conventional nonstationary bandit algorithms. Although it is very difficult to provide theoretical performance guarantees for our new approach, we also prove a novel regret bound for linear posterior sampling with measurement error that may serve as a foundation for future theoretical work. Aditya A. Ramesh, Paulo E. Rauber, Michelangelo Conserva, Jürgen Schmidhuber |
Neural Comput. | 4 |
| 2022 | Bayesian Brains and the Rényi DivergenceabstractUnder the Bayesian brain hypothesis, behavioral variations can be attributed to different priors over generative model parameters. This provides a formal explanation for why individuals exhibit inconsistent behavioral preferences when confronted with similar choices. For example, greedy preferences are a consequence of confident (or precise) beliefs over certain outcomes. Here, we offer an alternative account of behavioral variability using Rényi divergences and their associated variational bounds. Rényi bounds are analogous to the variational free energy (or evidence lower bound) and can be derived under the same assumptions. Importantly, these bounds provide a formal way to establish behavioral differences through an α parameter, given fixed priors. This rests on changes in α that alter the bound (on a continuous scale), inducing different posterior estimates and consequent variations in behavior. Thus, it looks as if individuals have different priors and have reached different conclusions. More specifically, α→0+ optimization constrains the variational posterior to be positive whenever the true posterior is positive. This leads to mass-covering variational estimates and increased variability in choice behavior. Furthermore, α→+∞ optimization constrains the variational posterior to be zero whenever the true posterior is zero. This leads to mass-seeking variational posteriors and greedy preferences. We exemplify this formulation through simulations of the multiarmed bandit task. We note that these α parameterizations may be especially relevant (i.e., shape preferences) when the true posterior is not in the same family of distributions as the assumed (simpler) approximate density, which may be the case in many real-world scenarios. The ensuing departure from vanilla variational inference provides a potentially useful explanation for differences in behavioral preferences of biological (or artificial) agents under the assumption that the brain performs variational Bayesian inference. Noor Sajid, Francesco Faccio, Lancelot Da Costa, Thomas Parr, Jürgen Schmidhuber, Karl J. Friston |
Neural Comput. | 5 |
| 2021 | Hierarchical Relational InferenceabstractCommon-sense physical reasoning in the real world requires learning about the interactions of objects and their dynamics. The notion of an abstract object, however, encompasses a wide variety of physical objects that differ greatly in terms of the complex behaviors they support. To address this, we propose a novel approach to physical reasoning that models objects as hierarchies of parts that may locally behave separately, but also act more globally as a single whole. Unlike prior approaches, our method learns in an unsupervised fashion directly from raw visual images to discover objects, parts, and their relations. It explicitly distinguishes multiple levels of abstraction and improves over a strong baseline at modeling synthetic and real-world videos. Aleksandar Stanic, Sjoerd van Steenkiste, Jürgen Schmidhuber |
AAAI | 3 |
| 2021 | The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of TransformersabstractRecently, many datasets have been proposed to test the systematic generalization ability of neural networks.The companion baseline Transformers, typically trained with default hyper-parameters from standard tasks, are shown to fail dramatically.Here we demonstrate that by revisiting model configurations as basic as scaling of embeddings, early stopping, relative positional embedding, and Universal Transformer variants, we can drastically improve the performance of Transformers on systematic generalization.We report improvements on five popular datasets: SCAN, CFQ, PCFG, COGS, and Mathematics dataset.Our models improve accuracy from 50% to 85% on the PCFG productivity split, and from 35% to 81% on COGS.On SCAN, relative positional embedding largely mitigates the EOS decision problem (Newman et al., 2020), yielding 100% accuracy on the length split with a cutoff at 26. Importantly, performance differences between these models are typically invisible on the IID data split.This calls for proper generalization validation sets for developing neural networks that generalize systematically.We publicly release the code to reproduce our results 1 . Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber |
EMNLP (1) | 3 |
| 2021 | Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
Róbert Csordás, Sjoerd van Steenkiste, Jürgen Schmidhuber |
ICLR | 3 |
| 2021 | Parameter-Based Value Functions
Francesco Faccio, Louis Kirsch, Jürgen Schmidhuber |
ICLR | 3 |
| 2021 | Unsupervised Object Keypoint Learning using Local Spatial Predictability
Anand Gopalakrishnan, Sjoerd van Steenkiste, Jürgen Schmidhuber |
ICLR | 3 |
| 2021 | Spatial Dependency Networks: Neural Layers for Improved Generative Image Modeling
Ðorðe Miladinovic, Aleksandar Stanic, Stefan Bauer, Jürgen Schmidhuber, Joachim M. Buhmann |
ICLR | 4 |
| 2021 | Learning Associative Inference Using Fast Weight Memory
Imanol Schlag, Tsendsuren Munkhdalai, Jürgen Schmidhuber |
ICLR | 3 |
| 2021 | Linear Transformers Are Secretly Fast Weight ProgrammersabstractWe show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early ’90s, where a slow neural net learns by gradient descent to program the fast weights of another net through sequences of elementary programming instructions which are additive outer products of self-invented activation patterns (today called keys and values). Such Fast Weight Programmers (FWPs) learn to manipulate the contents of a finite memory and dynamically interact with it. We infer a memory capacity limitation of recent linearised softmax attention variants, and replace the purely additive outer products by a delta rule-like programming instruction, such that the FWP can more easily learn to correct the current mapping from keys to values. The FWP also learns to compute dynamically changing learning rates. We also propose a new kernel function to linearise attention which balances simplicity and effectiveness. We conduct experiments on synthetic retrieval problems as well as standard machine translation and language modelling tasks which demonstrate the benefits of our methods. Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber |
ICML | 3 |
| 2021 | Improving Stateful Premise Selection with Transformers
Krsto Prorokovic, Michael Wand 0002, Jürgen Schmidhuber |
CICM | 3 |
| 2021 | Going Beyond Linear Transformers with Recurrent Fast Weight ProgrammersabstractTransformers with linearised attention (''linear Transformers'') have demonstrated the practical scalability and effectiveness of outer product-based Fast Weight Programmers (FWPs) from the '90s. However, the original FWP formulation is more general than the one of linear Transformers: a slow neural network (NN) continually reprograms the weights of a fast NN with arbitrary architecture. In existing linear Transformers, both NNs are feedforward and consist of a single layer. Here we explore new variations by adding recurrence to the slow and fast nets. We evaluate our novel recurrent FWPs (RFWPs) on two synthetic algorithmic tasks (code execution and sequential ListOps), Wikitext-103 language models, and on the Atari 2600 2D game environment. Our models exhibit properties of Transformers and RNNs. In the reinforcement learning setting, we report large improvements over LSTM in several Atari games. Our code is public. Kazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen Schmidhuber |
NeurIPS | 4 |
| 2021 | Meta Learning Backpropagation And Improving ItabstractMany concepts have been proposed for meta learning with neural networks (NNs), e.g., NNs that learn to reprogram fast weights, Hebbian plasticity, learned learning rules, and meta recurrent NNs. Our Variable Shared Meta Learning (VSML) unifies the above and demonstrates that simple weight-sharing and sparsity in an NN is sufficient to express powerful learning algorithms (LAs) in a reusable fashion. A simple implementation of VSML where the weights of a neural network are replaced by tiny LSTMs allows for implementing the backpropagation LA solely by running in forward-mode. It can even meta learn new LAs that differ from online backpropagation and generalize to datasets outside of the meta training distribution without explicit gradient calculation. Introspection reveals that our meta learned LAs learn through fast association in a way that is qualitatively different from gradient descent. Louis Kirsch, Jürgen Schmidhuber |
NeurIPS | 2 |
| 2021 | Reinforcement Learning in Sparse-Reward Environments With Hindsight Policy GradientsabstractA reinforcement learning agent that needs to pursue different goals across episodes requires a goal-conditional policy. In addition to their potential to generalize desirable behavior to unseen goals, such policies may also enable higher-level planning based on subgoals. In sparse-reward environments, the capacity to exploit information about the degree to which an arbitrary goal has been achieved while another goal was intended appears crucial to enabling sample efficient learning. However, reinforcement learning agents have only recently been endowed with such capacity for hindsight. In this letter, we demonstrate how hindsight can be introduced to policy gradient methods, generalizing this idea to a broad class of successful algorithms. Our experiments on a diverse selection of sparse-reward environments show that hindsight leads to a remarkable increase in sample efficiency. Paulo E. Rauber, Avinash Ummadisingu, Filipe Wall Mutz, Jürgen Schmidhuber |
Neural Comput. | 4 |
| 2021 | Deep neural network representation and Generative Adversarial Learning
Ariel Ruiz-Garcia, Jürgen Schmidhuber, Vasile Palade, Clive Cheong Took, Danilo P. Mandic |
Neural Networks | 2 |
| 2020 | Motion Dynamics Improve Speaker-Independent LipreadingabstractWe present a novel lipreading system that improves on the task of speaker-independent word recognition by decoupling motion and content dynamics. We achieve this by implementing a deep learning architecture that uses two distinct pipelines to process motion and content and subsequently merges them, implementing an end-to-end trainable system that performs fusion of independently learned representations. We obtain a average relative word accuracy improvement of ≈6.8% on unseen speakers and of ≈3.3% on known speakers, with respect to a baseline which uses a standard architecture. Matteo Riva, Michael Wand 0002, Jürgen Schmidhuber |
ICASSP | 3 |
| 2020 | Improving Generalization in Meta Reinforcement Learning using Learned Objectives
Louis Kirsch, Sjoerd van Steenkiste, Jürgen Schmidhuber |
ICLR | 3 |
| 2020 | The DeepScoresV2 Dataset and Benchmark for Music Object DetectionabstractIn this paper, we present DeepScoresV2, an extended version of the DeepScores dataset for optical music recognition (OMR). We improve upon the original DeepScores dataset by providing much more detailed annotations, namely (a) annotations for 135 classes including fundamental symbols of non-fixed size and shape, increasing the number of annotated symbols by 23%; (b) oriented bounding boxes; (c) higher-level rhythm and pitch information (onset beat for all symbols and line position for noteheads); and (d) a compatibility mode for easy use in conjunction with the MUSCIMA++ dataset for OMR on handwritten documents. These additions open up the potential for future advancement in OMR research. Additionally, we release two state-of-the-art baselines for DeepScoresV2 based on Faster R-CNN and the Deep Watershed Detector. An analysis of the baselines shows that regular orthogonal bounding boxes are unsuitable for objects which are long, small, and potentially rotated, such as ties and beams, which demonstrates the need for detection algorithms that naturally incorporate object angles. The dataset, code and pre-trained models, as well as user instructions, are publicly available at https://zenodo.org/record/4012193. Lukas Tuggener, Yvan Putra Satyawan, Alexander Pacha, Jürgen Schmidhuber, Thilo Stadelmann |
ICPR | 4 |
| 2020 | Fusion Architectures for Word-Based Audiovisual Speech Recognition
Michael Wand 0002, Jürgen Schmidhuber |
INTERSPEECH | 2 |
| 2020 | Generative Adversarial Networks are special cases of Artificial Curiosity (1990) and also closely related to Predictability Minimization (1991)
Jürgen Schmidhuber |
Neural Networks | 1 |
| 2020 | Investigating object compositionality in Generative Adversarial Networks
Sjoerd van Steenkiste, Karol Kurach, Jürgen Schmidhuber, Sylvain Gelly |
Neural Networks | 3 |
| 2019 | Improving Differentiable Neural Computers Through Memory Masking, De-allocation, and Link Distribution Sharpness Control
Róbert Csordás, Jürgen Schmidhuber |
ICLR (Poster) | 2 |
| 2019 | Hindsight policy gradients
Paulo E. Rauber, Avinash Ummadisingu, Filipe Wall Mutz, Jürgen Schmidhuber |
ICLR (Poster) | 4 |
| 2019 | Are Disentangled Representations Helpful for Abstract Visual Reasoning?abstractA disentangled representation encodes information about the salient factors of variation in the data independently. Although it is often argued that this representational format is useful in learning to solve many real-world down-stream tasks, there is little empirical evidence that supports this claim. In this paper, we conduct a large-scale study that investigates whether disentangled representations are more suitable for abstract reasoning tasks. Using two new tasks similar to Raven's Progressive Matrices, we evaluate the usefulness of the representations learned by 360 state-of-the-art unsupervised disentanglement models. Based on these representations, we train 3600 abstract reasoning models and observe that disentangled representations do in fact lead to better down-stream performance. In particular, they enable quicker learning using fewer samples. Sjoerd van Steenkiste, Francesco Locatello, Jürgen Schmidhuber, Olivier Bachem |
NeurIPS | 3 |
| 2018 | Investigations on End- to-End Audiovisual FusionabstractAudiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural network architecture which is trained end-to-end, without the need to separately model the process of decision fusion as in conventional (e.g. HMM-based) systems. The fusion system outperforms single-modality recognition under all noise conditions. Investigation of the saliency of the input features shows that the neural network automatically adapts to different noise levels in the acoustic signal. Michael Wand 0002, Jürgen Schmidhuber, Ngoc Thang Vu |
ICASSP | 2 |
| 2018 | Relational Neural Expectation Maximization: Unsupervised Discovery of Objects and their Interactions
Sjoerd van Steenkiste, Klaus Greff, Jürgen Schmidhuber |
ICLR (Poster) | 4 |
| 2018 | DeepScores-A Dataset for Segmentation, Detection and Classification of Tiny ObjectsabstractWe present the DeepScores dataset with the goal of advancing the state-of-the-art in small object recognition by placing the question of object recognition in the context of scene understanding. DeepScores contains high quality images of musical scores, partitioned into 300, 000 sheets of written music that contain symbols of different shapes and sizes. With close to a hundred million small objects, this makes our dataset not only unique, but also the largest public dataset. DeepScores comes with ground truth for object classification, detection and semantic segmentation. DeepScores thus poses a relevant challenge for computer vision in general, and optical music recognition (OMR) research in particular. We present a detailed statistical analysis of the dataset, comparing it with other computer vision datasets like PASCAL VOC, SUN, SVHN, ImageNet, MS-COCO, as well as with other OMR datasets. Finally, we provide baseline performances for object classification, intuition for the inherent difficulty that DeepScores poses to state-of-the-art object detectors like YOLO or R-CNN, and give pointers to future research based on this dataset. Lukas Tuggener, Ismail Elezi, Jürgen Schmidhuber, Marcello Pelillo, Thilo Stadelmann |
ICPR | 3 |
| 2018 | Domain-Adversarial Training for Session Independent EMG-based Speech Recognition
Michael Wand 0002, Tanja Schultz, Jürgen Schmidhuber |
INTERSPEECH | 3 |
| 2018 | Recurrent World Models Facilitate Policy EvolutionabstractA generative recurrent neural network is quickly trained in an unsupervised manner to model popular reinforcement learning environments through compressed spatio-temporal representations. The world model's extracted features are fed into compact and simple policies trained by evolution, achieving state of the art results in various environments. We also train our agent entirely inside of an environment generated by its own internal world model, and transfer this policy back into the actual environment. Interactive version of this paper is available at https://worldmodels.github.io David Ha, Jürgen Schmidhuber |
NeurIPS | 2 |
| 2018 | Learning to Reason with Third Order Tensor ProductsabstractWe combine Recurrent Neural Networks with Tensor Product Representations to learn combinatorial representations of sequential data. This improves symbolic interpretation and systematic generalisation. Our architecture is trained end-to-end through gradient descent on a variety of simple natural language reasoning tasks, significantly outperforming the latest state-of-the-art models in single-task and all-tasks settings. We also augment a subset of the data such that training and test data exhibit large systematic differences and show that our approach generalises better than the previous state-of-the-art. Imanol Schlag, Jürgen Schmidhuber |
NeurIPS | 2 |
| 2017 | Highway and Residual Networks learn Unrolled Iterative Estimation
Klaus Greff, Rupesh Kumar Srivastava, Jürgen Schmidhuber |
ICLR (Poster) | 3 |
| 2017 | Recurrent Highway NetworksabstractMany sequential processing tasks require complex nonlinear transition functions from one step to the next. However, recurrent neural networks with “deep” transition functions remain difficult to train, even when using Long Short-Term Memory (LSTM) networks. We introduce a novel theoretical analysis of recurrent networks based on Gersgorin’s circle theorem that illuminates several modeling and optimization issues and improves our understanding of the LSTM cell. Based on this analysis we propose Recurrent Highway Networks, which extend the LSTM architecture to allow step-to-step transition depths larger than one. Several language modeling experiments demonstrate that the proposed architecture results in powerful and efficient models. On the Penn Treebank corpus, solely increasing the transition depth from 1 to 10 improves word-level perplexity from 90.6 to 65.4 using the same number of parameters. On the larger Wikipedia datasets for character prediction (text8 and enwik8), RHNs outperform all previous results and achieve an entropy of 1.27 bits per character. Julian G. Zilly, Rupesh Kumar Srivastava, Jan Koutník, Jürgen Schmidhuber |
ICML | 4 |
| 2017 | Improving Speaker-Independent Lipreading with Domain-Adversarial TrainingabstractWe present a Lipreading system, i.e. a speech recognition system using only visual features, which uses domain-adversarial training for speaker independence. Domain-adversarial training is integrated into the optimization of a lipreader based on a stack of feedforward and LSTM (Long Short-Term Memory) recurrent neural networks, yielding an end-to-end trainable system which only requires a very small number of frames of untranscribed target data to substantially improve the recognition accuracy on the target speaker. On pairs of different source and target speakers, we achieve a relative accuracy improvement of around 40% with only 15 to 20 seconds of untranscribed target speech data. On multi-speaker training setups, the accuracy improvements are smaller but still substantial. Michael Wand 0002, Jürgen Schmidhuber |
INTERSPEECH | 2 |
| 2017 | Neural Expectation MaximizationabstractMany real world tasks such as reasoning and physical interaction require identification and manipulation of conceptual entities. A first step towards solving these tasks is the automated discovery of distributed symbol-like representations. In this paper, we explicitly formalize this problem as inference in a spatial mixture model where each component is parametrized by a neural network. Based on the Expectation Maximization framework we then derive a differentiable clustering method that simultaneously learns how to group and represent individual entities. We evaluate our method on the (sequential) perceptual grouping task and find that it is able to accurately recover the constituent objects. We demonstrate that the learned representations are useful for next-step prediction. Klaus Greff, Sjoerd van Steenkiste, Jürgen Schmidhuber |
NIPS | 3 |
| 2017 | Continual curiosity-driven skill acquisition from high-dimensional video inputs for humanoid robotsabstractIn the absence of external guidance, how can a robot learn to map the many raw pixels of high-dimensional visual inputs to useful action sequences? We propose here Continual Curiosity driven Skill Acquisition (CCSA). CCSA makes robots intrinsically motivated to acquire, store and reuse skills. Previous curiosity-based agents acquired skills by associating intrinsic rewards with world model improvements, and used reinforcement learning to learn how to get these intrinsic rewards. CCSA also does this, but unlike previous implementations, the world model is a set of compact low-dimensional representations of the streams of high-dimensional visual information, which are learned through incremental slow feature analysis. These representations augment the robot's state space with new information about the environment. We show how this information can have a higher-level (compared to pixels) and useful interpretation, for example, if the robot has grasped a cup in its field of view or not. After learning a representation, large intrinsic rewards are given to the robot for performing actions that greatly change the feature output, which has the tendency otherwise to change slowly in time. We show empirically what these actions are (e.g., grasping the cup) and how they can be useful as skills. An acquired skill includes both the learned actions and the learned slow feature representation. Skills are stored and reused to generate new observations, enabling continual acquisition of complex skills. We present results of experiments with an iCub humanoid robot that uses CCSA to incrementally acquire skills to topple, grasp and pick-place a cup, driven by its intrinsic motivation from raw pixel vision. Varun Raj Kompella, Marijn F. Stollenga, Matthew D. Luciw, Jürgen Schmidhuber |
Artif. Intell. | 4 |
| 2017 | LSTM: A Search Space OdysseyabstractSeveral variants of the long short-term memory (LSTM) architecture for recurrent neural networks have been proposed since its inception in 1995. In recent years, these networks have become the state-of-the-art models for a variety of machine learning problems. This has led to a renewed interest in understanding the role and utility of various computational components of typical LSTM variants. In this paper, we present the first large-scale analysis of eight LSTM variants on three representative tasks: speech recognition, handwriting recognition, and polyphonic music modeling. The hyperparameters of all LSTM variants for each task were optimized separately using random search, and their importance was assessed using the powerful functional ANalysis Of VAriance framework. In total, we summarize the results of 5400 experimental runs ( ≈ 15 years of CPU time), which makes our study the largest of its kind on LSTM networks. Our results show that none of the variants can improve upon the standard LSTM architecture significantly, and demonstrate the forget gate and the output activation function to be its most critical components. We further observe that the studied hyperparameters are virtually independent and derive guidelines for their efficient adjustment. Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R. Steunebrink, Jürgen Schmidhuber |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2016 | A Wavelet-based Encoding for NeuroevolutionabstractA new indirect scheme for encoding neural network connection weights as sets of wavelet-domain coefficients is proposed in this paper. It exploits spatial regularities in the weight-space to reduce the gene-space dimension by considering the low-frequency wavelet coefficients only. The wavelet-based encoding builds on top of a frequency-domain encoding, but unlike when using a Fourier-type transform, it offers gene locality while preserving continuity of the genotype-phenotype mapping. We argue that this added property allows for more efficient evolutionary search and demonstrate this on the octopus-arm control task, where superior solutions were found in fewer generations. The scalability of the wavelet-based encoding is shown by evolving networks with many parameters to control game-playing agents in the Arcade Learning Environment. Sjoerd van Steenkiste, Jan Koutník, Kurt Driessens, Jürgen Schmidhuber |
GECCO | 4 |
| 2016 | Lipreading with long short-term memoryabstractLipreading, i.e. speech recognition from visual-only recordings of a speaker's face, can be achieved with a processing pipeline based solely on neural networks, yielding significantly better accuracy than conventional methods. Feedforward and recurrent neural network layers (namely Long Short-Term Memory; LSTM) are stacked to form a single structure which is trained by back-propagating error gradients through all the layers. The performance of such a stacked network was experimentally evaluated and compared to a standard Support Vector Machine classifier using conventional computer vision features (Eigenlips and Histograms of Oriented Gradients). The evaluation was performed on data from 19 speakers of the publicly available GRID corpus. With 51 different words to classify, we report a best word accuracy on held-out evaluation speakers of 79.6% using the end-to-end neural network-based solution (11.6% improvement over the best feature-based solution evaluated). Michael Wand 0002, Jan Koutník, Jürgen Schmidhuber |
ICASSP | 3 |
| 2016 | Deep Neural Network Frontend for Continuous EMG-Based Speech Recognition
Michael Wand 0002, Jürgen Schmidhuber |
INTERSPEECH | 2 |
| 2016 | Tagger: Deep Unsupervised Perceptual GroupingabstractWe present a framework for efficient perceptual inference that explicitly reasons about the segmentation of its inputs and features. Rather than being trained for any specific segmentation, our framework learns the grouping process in an unsupervised manner or alongside any supervised task. We enable a neural network to group the representations of different objects in an iterative manner through a differentiable mechanism. We achieve very fast convergence by allowing the system to amortize the joint iterative inference of the groupings and their representations. In contrast to many other recently proposed methods for addressing multi-object scenes, our system does not assume the inputs to be images and can therefore directly handle other modalities. We evaluate our method on multi-digit classification of very cluttered images that require texture segmentation. Remarkably our method achieves improved classification performance over convolutional networks despite being fully connected, by making use of the grouping mechanism. Furthermore, we observe that our system greatly improves upon the semi-supervised result of a baseline Ladder network on our dataset. These results are evidence that grouping is a powerful tool that can help to improve sample efficiency. Klaus Greff, Antti Rasmus, Mathias Berglund, Tele Hao, Harri Valpola, Jürgen Schmidhuber |
NIPS | 6 |
| 2016 | Optimal Curiosity-Driven Modular Incremental Slow Feature AnalysisabstractConsider a self-motivated artificial agent who is exploring a complex environment. Part of the complexity is due to the raw high-dimensional sensory input streams, which the agent needs to make sense of. Such inputs can be compactly encoded through a variety of means; one of these is slow feature analysis (SFA). Slow features encode spatiotemporal regularities, which are information-rich explanatory factors (latent variables) underlying the high-dimensional input streams. In our previous work, we have shown how slow features can be learned incrementally, while the agent explores its world, and modularly, such that different sets of features are learned for different parts of the environment (since a single set of regularities does not explain everything). In what order should the agent explore the different parts of the environment? Following Schmidhuber's theory of artificial curiosity, the agent should always concentrate on the area where it can learn the easiest-to-learn set of features that it has not already learned. We formalize this learning problem and theoretically show that, using our model, called curiosity-driven modular incremental slow feature analysis, the agent on average will learn slow feature representations in order of increasing learning difficulty, under certain mild conditions. We provide experimental results to support the theoretical analysis. Varun Raj Kompella, Matthew D. Luciw, Marijn F. Stollenga, Jürgen Schmidhuber |
Neural Comput. | 4 |
| 2015 | Training Very Deep NetworksabstractTheoretical and empirical evidence indicates that the depth of neural networks is crucial for their success. However, training becomes more difficult as depth increases, and training of very deep networks remains an open problem. Here we introduce a new architecture designed to overcome this. Our so-called highway networks allow unimpeded information flow across many layers on information highways. They are inspired by Long Short-Term Memory recurrent networks and use adaptive gating units to regulate the information flow. Even with hundreds of layers, highway networks can be trained directly through simple gradient descent. This enables the study of extremely deep and efficient architectures. Rupesh Kumar Srivastava, Klaus Greff, Jürgen Schmidhuber |
NIPS | 3 |
| 2015 | Parallel Multi-Dimensional LSTM, With Application to Fast Biomedical Volumetric Image SegmentationabstractConvolutional Neural Networks (CNNs) can be shifted across 2D images or 3D videos to segment them. They have a fixed input size and typically perceive only small local contexts of the pixels to be classified as foreground or background. In contrast, Multi-Dimensional Recurrent NNs (MD-RNNs) can perceive the entire spatio-temporal context of each pixel in a few sweeps through all pixels, especially when the RNN is a Long Short-Term Memory (LSTM). Despite these theoretical advantages, however, unlike CNNs, previous MD-LSTM variants were hard to parallelise on GPUs. Here we re-arrange the traditional cuboid order of computations in MD-LSTM in pyramidal fashion. The resulting PyraMiD-LSTM is easy to parallelise, especially for 3D data such as stacks of brain slice images. PyraMiD-LSTM achieved best known pixel-wise brain image segmentation results on MRBrainS13 (and competitive results on EM-ISBI12). Marijn F. Stollenga, Wonmin Byeon, Marcus Liwicki, Jürgen Schmidhuber |
NIPS | 4 |
| 2015 | Assessment of algorithms for mitosis detection in breast cancer histopathology images
Mitko Veta, Paul J. van Diest, Stefan M. Willems, Anant Madabhushi, Angel Cruz-Roa, Fabio A. González 0001, Anders Boesen Lindbo Larsen, Jacob S. Vestergaard, Anders Bjorholm Dahl, Dan C. Ciresan, Jürgen Schmidhuber, Alessandro Giusti, Luca Maria Gambardella, Faik Boray Tek, Thomas Walter 0003, Ching-Wei Wang, Satoshi Kondo, Bogdan J. Matuszewski, Frédéric Precioso, Violet Snell, Josef Kittler, Teófilo Emídio de Campos, Adnan Mujahid Khan, Nasir M. Rajpoot, Evdokia Arkoumani, Miangela M. Lacle, Max A. Viergever, Josien P. W. Pluim |
Medical Image Anal. | 12 |
| 2015 | Deep learning in neural networks: An overview
Jürgen Schmidhuber |
Neural Networks | 1 |
| 2014 | Evolving deep unsupervised convolutional networks for vision-based reinforcement learningabstractDealing with high-dimensional input spaces, like visual input, is a challenging task for reinforcement learning (RL). Neuroevolution (NE), used for continuous RL problems, has to either reduce the problem dimensionality by (1) compressing the representation of the neural network controllers or (2) employing a pre-processor (compressor) that transforms the high-dimensional raw inputs into low-dimensional features. In this paper, we are able to evolve extremely small recurrent neural network (RNN) controllers for a task that previously required networks with over a million weights. The high-dimensional visual input, which the controller would normally receive, is first transformed into a compact feature vector through a deep, max-pooling convolutional neural network (MPCNN). Both the MPCNN preprocessor and the RNN controller are evolved successfully to control a car in the TORCS racing simulator using only visual input. This is the first use of deep learning in the context evolutionary RL. Jan Koutník, Jürgen Schmidhuber, Faustino J. Gomez |
GECCO | 2 |
| 2014 | Human-robot cooperation: fast, interactive learning from binary feedbackabstractNo abstract available. Jawad Nagi, Hung Quoc Ngo 0001, Jürgen Schmidhuber, Luca Maria Gambardella, Gianni A. Di Caro |
HRI | 3 |
| 2014 | Reactive Reaching and Grasping on a Humanoid - Towards Closing the Action-Perception Loop on the iCubabstractWe propose a system incorporating a tight integration between computer vision and robot control modules on a complex, high-DOF humanoid robot. Its functionality is showcased by having our iCub humanoid robot pick-up objects from a table in front of it. An important feature is that the system can avoid obstacles - other objects detected in the visual stream - while reaching for the intended target object. Our integration also allows for non-static environments, i.e. the reaching is adapted on-the-fly from the visual feedback received, e.g. when an obstacle is moved into the trajectory. Furthermore we show that this system can be used both in autonomous and tele-operation scenarios. Jürgen Leitner, Mikhail Frank, Alexander Förster, Jürgen Schmidhuber |
ICINCO (1) | 4 |
| 2014 | A Clockwork RNNabstractSequence prediction and classification are ubiquitous and challenging problems in machine learning that can require identifying complex dependencies between temporally distant inputs. Recurrent Neural Networks (RNNs) have the ability, in theory, to cope with these temporal dependencies by virtue of the short-term memory implemented by their recurrent (feedback) connections. However, in practice they are difficult to train successfully when long-term memory is required. This paper introduces a simple, yet powerful modification to the simple RNN (SRN) architecture, the Clockwork RNN (CW-RNN), in which the hidden layer is partitioned into separate modules, each processing inputs at its own temporal granularity, making computations only at its prescribed clock rate. Rather than making the standard RNN models more complex, CW-RNN reduces the number of SRN parameters, improves the performance significantly in the tasks tested, and speeds up the network evaluation. The network is demonstrated in preliminary experiments involving three tasks: audio signal generation, TIMIT spoken word classification, where it outperforms both SRN and LSTM networks, and online handwriting recognition, where it outperforms SRNs. Jan Koutník, Klaus Greff, Faustino J. Gomez, Jürgen Schmidhuber |
ICML | 4 |
| 2014 | Explore to see, learn to perceive, get the actions for free: SKILLABILITYabstractHow can a humanoid robot autonomously learn and refine multiple sensorimotor skills as a byproduct of curiosity driven exploration, upon its high-dimensional unprocessed visual input? We present SKILLABILITY, which makes this possible. It combines the recently introduced Curiosity Driven Modular Incremental Slow Feature Analysis (Curious Dr. MISFA) with the well-known options framework. Curious Dr. MISFA's objective is to acquire abstractions as quickly as possible. These abstractions map high-dimensional pixel-level vision to a low-dimensional manifold. We find that each learnable abstraction augments the robot's state space (a set of poses) with new information about the environment, for example, when the robot is grasping a cup. The abstraction is a function on an image, called a slow feature, which can effectively discretize a high-dimensional visual sequence. For example, it maps the sequence of the robot watching its arm as it moves around, grasping randomly, then grasping a cup, and moving around some more while holding the cup, into a step function having two outputs: when the cup is or is not currently grasped. The new state space includes this grasped/not grasped information. Each abstraction is coupled with an option. The reward function for the option's policy (learned through Least Squares Policy Iteration) is high for transitions that produce a large change in the step-functionlike slow features. This corresponds to finding bottleneck states, which are known good subgoals for hierarchical reinforcement learning - in the example, the subgoal corresponds to grasping the cup. The final skill includes both the learned policy and the learned abstraction. SKILLABILITY makes our iCub the first humanoid robot to learn complex skills such as to topple or grasp an object, from raw high-dimensional video input, driven purely by its intrinsic motivations. Varun Raj Kompella, Marijn F. Stollenga, Matthew D. Luciw, Jürgen Schmidhuber |
IJCNN | 4 |
| 2014 | Improving robot vision models for object detection through interactionabstractWe propose a method for learning specific object representations that can be applied (and reused) in visual detection and identification tasks. A machine learning technique called Cartesian Genetic Programming (CGP) is used to create these models based on a series of images. Our research investigates how manipulation actions might allow for the development of better visual models and therefore better robot vision. This paper describes how visual object representations can be learned and improved by performing object manipulation actions, such as, poke, push and pick-up with a humanoid robot. The improvement can be measured and allows for the robot to select and perform the `right' action, i.e. the action with the best possible improvement of the detector. Jürgen Leitner, Alexander Förster, Jürgen Schmidhuber |
IJCNN | 3 |
| 2014 | Candidate Sampling for Neuron Reconstruction from Anisotropic Electron Microscopy Volumes
Jan Funke, Julien N. P. Martel, Stephan Gerhard, Bjoern Andres, Dan C. Ciresan, Alessandro Giusti, Luca Maria Gambardella, Jürgen Schmidhuber, Hanspeter Pfister, Albert Cardona, Matthew Cook 0001 |
MICCAI (1) | 8 |
| 2014 | Deep Networks with Internal Selective Attention through Feedback Connections
Marijn F. Stollenga, Jonathan Masci, Faustino J. Gomez, Jürgen Schmidhuber |
NIPS | 4 |
| 2014 | Natural evolution strategies
Daan Wierstra, Tom Schaul, Tobias Glasmachers, Yi Sun 0003, Jan Peters 0001, Jürgen Schmidhuber |
J. Mach. Learn. Res. | 6 |
| 2014 | Multimodal Similarity-Preserving HashingabstractWe introduce an efficient computational framework for hashing data belonging to multiple modalities into a single representation space where they become mutually comparable. The proposed approach is based on a novel coupled siamese neural network architecture and allows unified treatment of intra- and inter-modality similarity learning. Unlike existing cross-modality similarity learning approaches, our hashing functions are not limited to binarized linear projections and can assume arbitrarily complex forms. We show experimentally that our method significantly outperforms state-of-the-art hashing approaches on multimedia retrieval tasks. Jonathan Masci, Michael M. Bronstein, Alexander M. Bronstein, Jürgen Schmidhuber |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Efficient Interactive Multiclass Learning from Binary FeedbackabstractWe introduce a novel algorithm called upper confidence - weighted learning (UCWL) for online multiclass learning from binary feedback (e.g., feedback that indicates whether the prediction was right or wrong). UCWL combines the upper confidence bound (UCB) framework with the soft confidence-weighted (SCW) online learning scheme. In UCB, each instance is classified using both score and uncertainty. For a given instance in the sequence, the algorithm might guess its class label primarily to reduce the class uncertainty. This is a form of informed exploration, which enables the performance to improve with lower sample complexity compared to the case without exploration. Combining UCB with SCW leads to the ability to deal well with noisy and nonseparable data, and state-of-the-art performance is achieved without increasing the computational cost. A potential application setting is human-robot interaction (HRI), where the robot is learning to classify some set of inputs while the human teaches it by providing only binary feedback—or sometimes even the wrong answer entirely. Experimental results in the HRI setting and with two benchmark datasets from other settings show that UCWL outperforms other state-of-the-art algorithms in the online binary feedback setting—and surprisingly even sometimes outperforms state-of-the-art algorithms that get full feedback (e.g., the true class label), whereas UCWL gets only binary feedback on the same data sequence. Hung Quoc Ngo 0001, Matthew D. Luciw, Jawad Nagi, Alexander Förster, Jürgen Schmidhuber, Ngo Anh Vien |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2013 | Humanoid learns to detect its own handsabstractRobust object manipulation is still a hard problem in robotics, even more so in high degree-of-freedom (DOF) humanoid robots. To improve performance a closer integration of visual and motor systems is needed. We herein present a novel method for a robot to learn robust detection of its own hands and fingers enabling sensorimotor coordination. It does so solely using its own camera images and does not require any external systems or markers. Our system based on Cartesian Genetic Programming (CGP) allows to evolve programs to perform this image segmentation task in real-time on the real hardware. We show results for a Nao and an iCub humanoid each detecting its own hands and fingers. Jürgen Leitner, Simon Harding, Mikhail Frank, Alexander Förster, Jürgen Schmidhuber |
IEEE Congress on Evolutionary Computation | 5 |
| 2013 | Evolving large-scale neural networks for vision-based TORCS
Jan Koutník, Giuseppe Cuccu, Jürgen Schmidhuber, Faustino J. Gomez |
FDG | 3 |
| 2013 | Evolving large-scale neural networks for vision-based reinforcement learningabstractThe idea of using evolutionary computation to train artificial neural networks, or neuroevolution (NE), for reinforcement learning (RL) tasks has now been around for over 20 years. However, as RL tasks become more challenging, the networks required become larger, as do their genomes. But, scaling NE to large nets (i.e. tens of thousands of weights) is infeasible using direct encodings that map genes one-to-one to network components. In this paper, we scale-up our compressed network encoding where network weight matrices are represented indirectly as a set of Fourier-type coefficients, to tasks that require very-large networks due to the high-dimensionality of their input space. The approach is demonstrated successfully on two reinforcement learning tasks in which the control networks receive visual input: (1) a vision-based version of the octopus control task requiring networks with over 3 thousand weights, and (2) a version of the TORCS driving game where networks with over 1 million weights are evolved to drive a car around a track using video images from the driver's perspective. Jan Koutník, Giuseppe Cuccu, Jürgen Schmidhuber, Faustino J. Gomez |
GECCO | 3 |
| 2013 | Fast image scanning with deep max-pooling convolutional neural networksabstractDeep Neural Networks now excel at image classification, detection and segmentation. When used to scan images by means of a sliding window, however, their high computational complexity can bring even the most powerful hardware to its knees. We show how dynamic programming can speedup the process by orders of magnitude, even when max-pooling layers are present. Alessandro Giusti, Dan C. Ciresan, Jonathan Masci, Luca Maria Gambardella, Jürgen Schmidhuber |
ICIP | 5 |
| 2013 | A fast learning algorithm for image segmentation with max-pooling convolutional networksabstractWe present a fast algorithm for training MaxPooling Convolutional Networks to segment images. This type of network yields record-breaking performance in a variety of tasks, but is normally trained on a computationally expensive patch-by-patch basis. Our new method processes each training image in a single pass, which is vastly more efficient. We validate the approach in different scenarios and report a 1500-fold speed-up. In an application to automated steel defect detection and segmentation, we obtain excellent performance with short training times. Jonathan Masci, Alessandro Giusti, Dan C. Ciresan, Gabriel Fricout, Jürgen Schmidhuber |
ICIP | 5 |
| 2013 | Upper Confidence Weighted Learning for Efficient Exploration in Multiclass Prediction with Binary Feedback
Hung Quoc Ngo 0001, Matthew D. Luciw, Ngo Anh Vien, Jürgen Schmidhuber |
IJCAI | 4 |
| 2013 | Artificial neural networks for spatial perception: Towards visual object localisation in humanoid robotsabstractIn this paper, we present our on-going research to allow humanoid robots to learn spatial perception. We are using artificial neural networks (ANN) to estimate the location of objects in the robot's environment. The method is using only the visual inputs and the joint encoder readings, no camera calibration and information is necessary, nor is a kinematic model. We find that these ANNs can be trained to allow spatial perception in Cartesian (3D) coordinates. These lightweight networks are providing estimates that are comparable to current state of the art approaches and can easily be used together with existing operational space controllers. Jürgen Leitner, Simon Harding, Mikhail Frank, Alexander Förster, Jürgen Schmidhuber |
IJCNN | 5 |
| 2013 | Multi-scale pyramidal pooling network for generic steel defect classificationabstractWe introduce a Multi-Scale Pyramidal Pooling Network tailored to generic steel defect classification, featuring a novel pyramidal pooling layer at multiple scales and a novel encoding layer. Thanks to the former, the network does not require all images of a given classification task to be of equal size. The latter narrows the gap to bag-of-features approaches. On various benchmark datasets, we evaluate and compare our system to convolutional neural networks and state-of-the-art computer vision methods. We also present results on a real industrial steel defect classification problem, where existing architectures are not applicable as they require equally sized input images. Our method substantially outperforms previous methods based on engineered features. It can be seen as a fully supervised hierarchical bag-of-features extension that is trained online and can be fine-tuned for any given task. Jonathan Masci, Ueli Meier, Gabriel Fricout, Jürgen Schmidhuber |
IJCNN | 4 |
| 2013 | Task-relevant roadmaps: A framework for humanoid motion planningabstractTo plan complex motions of robots with many degrees of freedom, our novel, very flexible framework builds task-relevant roadmaps (TRMs), using a new sampling-based optimizer called Natural Gradient Inverse Kinematics (NGIK) based on natural evolution strategies (NES). To build TRMs, NGIK iteratively optimizes postures covering task-spaces expressed by arbitrary task-functions, subject to constraints expressed by arbitrary cost-functions, transparently dealing with both hard and soft constraints. TRMs are grown to maximally cover the task-space while minimizing costs. Unlike Jacobian-based methods, our algorithm does not rely on calculation of gradients, making application of the algorithm much simpler. We show how NGIK outperforms recent related sampling algorithms. A video demo (http://youtu.be/N6x2e1Zf_yg) successfully applies TRMs to an iCub humanoid robot with 41 DOF in its upper body, arms, hands, head, and eyes. To our knowledge, no similar methods exhibit such a degree of flexibility in defining movements. Marijn F. Stollenga, Leo Pape, Mikhail Frank, Jürgen Leitner, Alexander Förster, Jürgen Schmidhuber |
IROS | 6 |
| 2013 | Mitosis Detection in Breast Cancer Histology Images with Deep Neural Networks
Dan C. Ciresan, Alessandro Giusti, Luca Maria Gambardella, Jürgen Schmidhuber |
MICCAI (2) | 4 |
| 2013 | Compete to ComputeabstractLocal competition among neighboring neurons is common in biological neural networks (NNs). We apply the concept to gradient-based, backprop-trained artificial multilayer NNs. NNs with competing linear units tend to outperform those with non-competing nonlinear units, and avoid catastrophic forgetting when training sets change over time. Rupesh Kumar Srivastava, Jonathan Masci, Sohrob Kazerounian, Faustino J. Gomez, Jürgen Schmidhuber |
NIPS | 5 |
| 2013 | PUF Modeling Attacks on Simulated and Silicon DataabstractWe discuss numerical modeling attacks on several proposed strong physical unclonable functions (PUFs). Given a set of challenge-response pairs (CRPs) of a Strong PUF, the goal of our attacks is to construct a computer algorithm which behaves indistinguishably from the original PUF on almost all CRPs. If successful, this algorithm can subsequently impersonate the Strong PUF, and can be cloned and distributed arbitrarily. It breaks the security of any applications that rest on the Strong PUF's unpredictability and physical unclonability. Our method is less relevant for other PUF types such as Weak PUFs. The Strong PUFs that we could attack successfully include standard Arbiter PUFs of essentially arbitrary sizes, and XOR Arbiter PUFs, Lightweight Secure PUFs, and Feed-Forward Arbiter PUFs up to certain sizes and complexities. We also investigate the hardness of certain Ring Oscillator PUF architectures in typical Strong PUF applications. Our attacks are based upon various machine learning techniques, including a specially tailored variant of logistic regression and evolution strategies. Our results are mostly obtained on CRPs from numerical simulations that use established digital models of the respective PUFs. For a subset of the considered PUFs-namely standard Arbiter PUFs and XOR Arbiter PUFs-we also lead proofs of concept on silicon data from both FPGAs and ASICs. Over four million silicon CRPs are used in this process. The performance on silicon CRPs is very close to simulated CRPs, confirming a conjecture from earlier versions of this work. Our findings lead to new design requirements for secure electrical Strong PUFs, and will be useful to PUF designers and attackers alike. Ulrich Rührmair, Jan Sölter, Frank Sehnke, Xiaolin Xu 0001, Ahmed Mahmoud, Vera Stoyanova, Gideon Dror, Jürgen Schmidhuber, Wayne P. Burleson, Srini Devadas |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2012 | Multi-column deep neural networks for image classificationabstractTraditional methods of computer vision and machine learning cannot match human performance on tasks such as the recognition of handwritten digits or traffic signs. Our biologically plausible, wide and deep artificial neural network architectures can. Small (often minimal) receptive fields of convolutional winner-take-all neurons yield large network depth, resulting in roughly as many sparsely connected neural layers as found in mammals between retina and visual cortex. Only winner neurons are trained. Several deep neural columns become experts on inputs preprocessed in different ways; their predictions are averaged. Graphics cards allow for fast training. On the very competitive MNIST handwriting benchmark, our method is the first to achieve near-human performance. On a traffic sign recognition benchmark it outperforms humans by a factor of two. We also improve the state-of-the-art on a plethora of common image classification benchmarks. Dan C. Ciresan, Ueli Meier, Jürgen Schmidhuber |
CVPR | 3 |
| 2012 | MT-CGP: mixed type cartesian genetic programmingabstractThe majority of genetic programming implementations build expressions that only use a single data type. This is in contrast to human engineered programs that typically make use of multiple data types, as this provides the ability to express solutions in a more natural fashion. In this paper, we present a version of Cartesian Genetic Programming that handles multiple data types. We demonstrate that this allows evolution to quickly find competitive, compact, and human readable solutions on multiple classification tasks. Simon Harding, Vincent Graziano, Jürgen Leitner, Jürgen Schmidhuber |
GECCO | 4 |
| 2012 | Reflexive Collision Response with Virtual Skin - Roadmap Planning Meets Reinforcement Learning
Mikhail Frank, Alexander Förster, Jürgen Schmidhuber |
ICAART (1) | 3 |
| 2012 | Low Complexity Proto-Value Function Learning from Sensory Observations with Incremental Slow Feature Analysis
Matthew D. Luciw, Jürgen Schmidhuber |
ICANN (2) | 2 |
| 2012 | The Modular Behavioral Environment for Humanoids and other Robots (MoBeE)
Mikhail Frank, Jürgen Leitner, Marijn F. Stollenga, Simon Harding, Alexander Förster, Jürgen Schmidhuber |
ICINCO (2) | 6 |
| 2012 | On the Size of the Online Kernel Sparsification Dictionary
Yi Sun 0003, Faustino J. Gomez, Jürgen Schmidhuber |
ICML | 3 |
| 2012 | Transfer learning for Latin and Chinese characters with Deep Neural NetworksabstractWe analyze transfer learning with Deep Neural Networks (DNN) on various character recognition tasks. DNN trained on digits are perfectly capable of recognizing uppercase letters with minimal retraining. They are on par with DNN fully trained on uppercase letters, but train much faster. DNN trained on Chinese characters easily recognize uppercase Latin letters. Learning Chinese characters is accelerated by first pretraining a DNN on a small subset of all classes and then continuing to train on all classes. Furthermore, pretrained nets consistently outperform randomly initialized nets on new tasks with few labeled data. Dan C. Ciresan, Ueli Meier, Jürgen Schmidhuber |
IJCNN | 3 |
| 2012 | Steel defect classification with Max-Pooling Convolutional Neural NetworksabstractWe present a Max-Pooling Convolutional Neural Network approach for supervised steel defect classification. On a classification task with 7 defects, collected from a real production line, an error rate of 7% is obtained. Compared to SVM classifiers trained on commonly used feature descriptors our best net performs at least two times better. Not only we do obtain much better results, but the proposed method also works directly on raw pixel intensities of detected and segmented steel defects, avoiding further time consuming and hard to optimize ad-hoc preprocessing. Jonathan Masci, Ueli Meier, Dan C. Ciresan, Jürgen Schmidhuber, Gabriel Fricout |
IJCNN | 4 |
| 2012 | Learning skills from play: Artificial curiosity on a Katana robot armabstractArtificial curiosity tries to maximize learning progress. We apply this concept to a physical system. Our Katana robot arm curiously plays with wooden blocks, using vision, reaching, and grasping. It is intrinsically motivated to explore its world. As a by-product, it learns how to place blocks stably, and how to stack blocks. Hung Quoc Ngo 0001, Matthew D. Luciw, Alexander Förster, Jürgen Schmidhuber |
IJCNN | 4 |
| 2012 | Transferring spatial perception between robots operating in a shared workspaceabstractWe use a Katana robotic arm to teach an iCub humanoid robot how to perceive the location of the objects it sees. To do this, the Katana positions an object within the shared workspace, and tells the iCub where it has placed it. While the iCub moves it observes the object, and a neural network then learns how to relate its pose and visual inputs to the object location. We show that satisfactory results can be obtained for localisation even in scenarios where the kinematic model is imprecise or not available. Furthermore, we demonstrate that this task can be accomplished safely. For this task we extend our collision avoidance software for the iCub to prevent collisions between multiple, independently controlled, heterogeneous robots in the same workspace. Jürgen Leitner, Simon Harding, Mikhail Frank, Alexander Förster, Jürgen Schmidhuber |
IROS | 5 |
| 2012 | Deep Neural Networks Segment Neuronal Membranes in Electron Microscopy ImagesabstractWe address a central problem of neuroanatomy, namely, the automatic segmentation of neuronal structures depicted in stacks of electron microscopy (EM) images. This is necessary to efficiently map 3D brain structure and connectivity. To segment {\em biological} neuron membranes, we use a special type of deep {\em artificial} neural network as a pixel classifier. The label of each pixel (membrane or non-membrane) is predicted from raw pixel values in a square window centered on it. The input layer maps each window pixel to a neuron. It is followed by a succession of convolutional and max-pooling layers which preserve 2D information and extract features with increasing levels of abstraction. The output layer produces a calibrated probability for each class. The classifier is trained by plain gradient descent on a $512 \times 512 \times 30$ stack with known ground truth, and tested on a stack of the same size (ground truth unknown to the authors) by the organizers of the ISBI 2012 EM Segmentation Challenge. Even without problem-specific post-processing, our approach outperforms competing techniques by a large margin in all three considered metrics, i.e. \emph{rand error}, \emph{warping error} and \emph{pixel error}. For pixel error, our approach is the only one outperforming a second human observer. Dan C. Ciresan, Alessandro Giusti, Luca Maria Gambardella, Jürgen Schmidhuber |
NIPS | 4 |
| 2012 | Compressed Network Complexity Search
Faustino J. Gomez, Jan Koutník, Jürgen Schmidhuber |
PPSN (1) | 3 |
| 2012 | Generalized Compressed Network Search
Rupesh Kumar Srivastava, Jürgen Schmidhuber, Faustino J. Gomez |
PPSN (1) | 2 |
| 2012 | Incremental learning using partial feedback for gesture-based human-swarm interactionabstractIn this paper we consider a human-swarm interaction scenario based on hand gestures. We study how the swarm can incrementally learn hand gestures through the interaction with a human instructor providing training gestures and correction feedback. The main contribution of the paper is a novel incremental machine learning approach that makes the robot swarm learn and recognize the gestures in a distributed and decentralized fashion using binary (i.e., yes/no) feedback. It exploits cooperative information exchange and swarm's intrinsic parallelism and redundancy. We perform extensive tests using real gesture images, showing that good classification accuracies are obtained even with rather few training samples and relatively small swarms. We also show the good scalability of the approach and its relatively low requirements in terms of communication overhead. Jawad Nagi, Hung Quoc Ngo 0001, Alessandro Giusti, Luca Maria Gambardella, Jürgen Schmidhuber, Gianni A. Di Caro |
RO-MAN | 5 |
| 2012 | Incremental Slow Feature Analysis: Adaptive Low-Complexity Slow Feature Updating from High-Dimensional Input StreamsabstractWe introduce here an incremental version of slow feature analysis (IncSFA), combining candid covariance-free incremental principal components analysis (CCIPCA) and covariance-free incremental minor components analysis (CIMCA). IncSFA's feature updating complexity is linear with respect to the input dimensionality, while batch SFA's (BSFA) updating complexity is cubic. IncSFA does not need to store, or even compute, any covariance matrices. The drawback to IncSFA is data efficiency: it does not use each data point as effectively as BSFA. But IncSFA allows SFA to be tractably applied, with just a few parameters, directly on high-dimensional input streams (e.g., visual input of an autonomous agent), while BSFA has to resort to hierarchical receptive-field-based architectures when the input dimension is too high. Further, IncSFA's updates have simple Hebbian and anti-Hebbian forms, extending the biological plausibility of SFA. Experimental results show IncSFA learns the same set of features as BSFA and can handle a few cases where BSFA fails. Varun Raj Kompella, Matthew D. Luciw, Jürgen Schmidhuber |
Neural Comput. | 3 |
| 2012 | Multi-column deep neural network for traffic sign classification
Dan C. Ciresan, Ueli Meier, Jonathan Masci, Jürgen Schmidhuber |
Neural Networks | 4 |
| 2011 | Curiosity-driven optimizationabstractThe principle of artificial curiosity directs active exploration towards the most informative or most interesting data. We show its usefulness for global black box optimization when data point evaluations are expensive. Gaussian process regression is used to model the fitness function based on all available observations so far. For each candidate point this model estimates expected Fitness reduction, and yields a novel closed-form expression of expected information gain. A new type of Pareto-front algorithm continually pushes the boundary of candidates not dominated by any other known data according to both criteria, using multi-objective evolutionary search. This makes the exploration-exploitation trade-off explicit, and permits maximally informed data selection. We illustrate the robustness of our approach in a number of experimental scenarios. Tom Schaul, Yi Sun 0003, Daan Wierstra, Faustino J. Gomez, Jürgen Schmidhuber |
IEEE Congress on Evolutionary Computation | 5 |
| 2011 | High dimensions and heavy tails for natural evolution strategiesabstractThe family of natural evolution strategies (NES) offers a principled approach to real-valued evolutionary optimization. NES follows the natural gradient of the expected fitness on the parameters of its search distribution. While general in its formulation, previous research has focused on multivariate Gaussian search distributions. Here we exhibit problem classes for which other search distributions are more appropriate, and then derive corresponding NES-variants. Tom Schaul, Tobias Glasmachers, Jürgen Schmidhuber |
GECCO | 3 |
| 2011 | Stacked Convolutional Auto-Encoders for Hierarchical Feature Extraction
Jonathan Masci, Ueli Meier, Dan C. Ciresan, Jürgen Schmidhuber |
ICANN (1) | 4 |
| 2011 | Convolutional Neural Network Committees for Handwritten Character ClassificationabstractIn 2010, after many years of stagnation, the MNIST handwriting recognition benchmark record dropped from 0.40% error rate to 0.35%. Here we report 0.27% for a committee of seven deep CNNs trained on graphics cards, narrowing the gap to human performance. We also apply the same architecture to NIST SD 19, a more challenging dataset including lower and upper case letters. A committee of seven CNNs obtains the best results published so far for both NIST digits and NIST letters. The robustness of our method is verified by analyzing 78125 different 7-net committees. Dan C. Ciresan, Ueli Meier, Luca Maria Gambardella, Jürgen Schmidhuber |
ICDAR | 4 |
| 2011 | Better Digit Recognition with a Committee of Simple Neural NetsabstractWe present a new method to train the members of a committee of one-hidden-layer neural nets. Instead of training various nets on subsets of the training data we preprocess the training data for each individual model such that the corresponding errors are decor related. On the MNIST digit recognition benchmark set we obtain a recognition error rate of 0.39%, using a committee of 25 one-hidden-layer neural nets, which is on par with state-of-the-art recognition rates of more complicated systems. Ueli Meier, Dan C. Ciresan, Luca Maria Gambardella, Jürgen Schmidhuber |
ICDAR | 4 |
| 2011 | Incremental Basis Construction from Temporal Difference Error
Yi Sun 0003, Faustino J. Gomez, Mark B. Ring, Jürgen Schmidhuber |
ICML | 4 |
| 2011 | Flexible, High Performance Convolutional Neural Networks for Image ClassificationabstractWe present a fast, fully parameterizable GPU implementation of Convolutional Neural Network variants. Our feature extractors are neither carefully designed nor pre-wired, but rather learned in a supervised way. Our deep hierarchical architectures achieve the best published results on benchmarks for object classification (NORB, CIFAR10) and handwritten digit recognition (MNIST), with error rates of 2.53%, 19.51%, 0.35%, respectively. Deep nets trained by simple back-propagation perform better than more shallow ones. Learning is surprisingly rapid. NORB is completely trained within five epochs. Test error rates on MNIST drop to 2.42%, 0.97 % and 0.48 % after 1, 3 and 17 epochs, respectively. Dan C. Ciresan, Ueli Meier, Jonathan Masci, Luca Maria Gambardella, Jürgen Schmidhuber |
IJCAI | 5 |
| 2011 | Incremental Slow Feature Analysis
Varun Raj Kompella, Matthew D. Luciw, Jürgen Schmidhuber |
IJCAI | 3 |
| 2011 | A committee of neural networks for traffic sign classificationabstractWe describe the approach that won the preliminary phase of the German traffic sign recognition benchmark with a better-than-human recognition rate of 98.98%.We obtain an even better recognition rate of 99.15% by further training the nets. Our fast, fully parameterizable GPU implementation of a Convolutional Neural Network does not require careful design of pre-wired feature extractors, which are rather learned in a supervised way. A CNN/MLP committee further boosts recognition performance. Dan C. Ciresan, Ueli Meier, Jonathan Masci, Jürgen Schmidhuber |
IJCNN | 4 |
| 2011 | Modular deep belief networks that do not forgetabstractDeep belief networks (DBNs) are popular for learning compact representations of high-dimensional data. However, most approaches so far rely on having a single, complete training set. If the distribution of relevant features changes during subsequent training stages, the features learned in earlier stages are gradually forgotten. Often it is desirable for learning algorithms to retain what they have previously learned, even if the input distribution temporarily changes. This paper introduces the M-DBN, an unsupervised modular DBN that addresses the forgetting problem. M-DBNs are composed of a number of modules that are trained only on samples they best reconstruct. While modularization by itself does not prevent forgetting, the M-DBN additionally uses a learning method that adjusts each module's learning rate proportionally to the fraction of best reconstructed samples. On the MNIST handwritten digit dataset module specialization largely corresponds to the digits discerned by humans. Furthermore, in several learning tasks with changing MNIST digits, M-DBNs retain learned features even after those features are removed from the training data, while monolithic DBNs of comparable size forget feature mappings learned before. Leo Pape, Faustino J. Gomez, Juergen Ring, Jürgen Schmidhuber |
IJCNN | 4 |
| 2011 | Unsupervised Modeling of Partially Observable Environments
Vincent Graziano, Jan Koutník, Jürgen Schmidhuber |
ECML/PKDD (1) | 3 |
| 2010 | Robust Texture Recognition Using Credal ClassifiersabstractTexture classification is used for many vision systems; in this paper we focus on improving the reliability of the classification through the so-called imprecise (or credal) classifiers, which suspend the judgment on the doubtful instances by returning a set of classes instead of a single class. Our view is that on critical instances it is more sensible to return a reliable set of classes rather than an unreliable single class. We compare the traditional naive Bayes classifier (NBC) against its imprecise counterpart, the naive credal classifier (NCC); we consider a standard classification dataset, when the problem is made progressively harder by introducing different image degradations or by providing smaller training sets. Experiments show that on the instances for which NCC returns more classes, NBC issues in fact unreliable classifications; the indeterminate classifications of NCC preserve reliability but at the same time also convey significant information, reducing the set of possible classes (on most critical instances) from 24 to some 2-3. Giorgio Corani, Alessandro Giusti, Davide Migliore, Jürgen Schmidhuber |
BMVC | 4 |
| 2010 | Modeling attacks on physical unclonable functionsabstractWe show in this paper how several proposed Physical Unclonable Functions (PUFs) can be broken by numerical modeling attacks. Given a set of challenge-response pairs (CRPs) of a PUF, our attacks construct a computer algorithm which behaves indistinguishably from the original PUF on almost all CRPs. This algorithm can subsequently impersonate the PUF, and can be cloned and distributed arbitrarily. This breaks the security of essentially all applications and protocols that are based on the respective PUF. The PUFs we attacked successfully include standard Arbited PUFs and Ring Oscillator PUFs of arbitrary sizes, and XO Arbiter PUFs, Lightweight Secure PUFs, and Feed-Forward Arbiter PUFs of up to a given size and complexity. Our attacks are based upon various machine learning techniques including Logistic Regression and Evolution Strategies. Our work leads to new design requirements for secure electrical PUFs, and will be useful to PUF designers and attackers alike. Ulrich Rührmair, Frank Sehnke, Jan Sölter, Gideon Dror, Srini Devadas, Jürgen Schmidhuber |
CCS | 6 |
| 2010 | Exponential natural evolution strategiesabstractThe family of natural evolution strategies (NES) offers a principled approach to real-valued evolutionary optimization by following the natural gradient of the expected fitness. Like the well-known CMA-ES, the most competitive algorithm in the field, NES comes with important invariance properties. In this paper, we introduce a number of elegant and efficient improvements of the basic NES algorithm. First, we propose to parameterize the positive definite covariance matrix using the exponential map, which allows the covariance matrix to be updated in a vector space. This new technique makes the algorithm completely invariant under linear transformations of the underlying search space, which was previously achieved only in the limit of small step sizes. Second, we compute all updates in the natural coordinate system, such that the natural gradient coincides with the vanilla gradient. This way we avoid the computation of the inverse Fisher information matrix, which is the main computational bottleneck of the original NES algorithm. Our new algorithm, exponential NES (xNES), is significantly simpler than its predecessors. We show that the various update rules in CMA-ES are closely related to the natural gradient updates of xNES. However, xNES is more principled than CMA-ES, as all the update rules needed for covariance matrix adaptation are derived from a single principle. We empirically assess the performance of the new algorithm on standard benchmark functions Tobias Glasmachers, Tom Schaul, Yi Sun 0003, Daan Wierstra, Jürgen Schmidhuber |
GECCO | 5 |
| 2010 | Evolving neural networks in compressed weight spaceabstractWe propose a new indirect encoding scheme for neural networks in which the weight matrices are represented in the frequency domain by sets Fourier coefficients. This scheme exploits spatial regularities in the matrix to reduce the dimensionality of the representation by ignoring high-frequency coefficients, as is done in lossy image compression. We compare the efficiency of searching in this "compressed" network space to searching in the space of directly encoded networks, using the CoSyNE neuroevolution algorithm on three benchmark problems: pole-balancing, ball throwing and octopus arm control. The results show that this encoding can dramatically reduce the search space dimensionality such that solutions can be found in significantly fewer evaluations Jan Koutník, Faustino J. Gomez, Jürgen Schmidhuber |
GECCO | 3 |
| 2010 | Multi-Dimensional Deep Memory Atari-Go Players for Parameter Exploring Policy Gradients
Mandy Grüttner, Frank Sehnke, Tom Schaul, Jürgen Schmidhuber |
ICANN (2) | 4 |
| 2010 | Policy Gradients for Cryptanalysis
Frank Sehnke, Christian Osendorfer, Jan Sölter, Jürgen Schmidhuber, Ulrich Rührmair |
ICANN (3) | 4 |
| 2010 | Multimodal Parameter-exploring Policy GradientsabstractPolicy Gradients with Parameter-based Exploration (PGPE) is a novel model-free reinforcement learning method that alleviates the problem of high-variance gradient estimates encountered in normal policy gradient methods. It has been shown to drastically speed up convergence for several large-scale reinforcement learning tasks. However the independent normal distributions used by PGPE to search through parameter space are inadequate for some problems with multimodal reward surfaces. This paper extends the basic PGPE algorithm to use multimodal mixture distributions for each parameter, while remaining efficient. Experimental results on the Rastrigin function and the inverted pendulum benchmark demonstrate the advantages of this modification, with faster convergence to better optima. Frank Sehnke, Alex Graves, Christian Osendorfer, Jürgen Schmidhuber |
ICMLA | 4 |
| 2010 | Improving the Asymptotic Performance of Markov Chain Monte-Carlo by Inserting VorticesabstractWe present a new way of converting a reversible finite Markov chain into a nonreversible one, with a theoretical guarantee that the asymptotic variance of the MCMC estimator based on the non-reversible chain is reduced. The method is applicable to any reversible chain whose states are not connected through a tree, and can be interpreted graphically as inserting vortices into the state transition graph. Our result confirms that non-reversible chains are fundamentally better than reversible ones in terms of asymptotic performance, and suggests interesting directions for further improving MCMC. Yi Sun 0003, Faustino J. Gomez, Jürgen Schmidhuber |
NIPS | 3 |
| 2010 | Characteristic Kernels on Structured Domains Excel in Robotics and Human Action Recognition
Somayeh Danafar, Arthur Gretton, Jürgen Schmidhuber |
ECML/PKDD (1) | 3 |
| 2010 | Formal Theory of Fun and Creativity
Jürgen Schmidhuber |
ECML/PKDD (1) | 1 |
| 2010 | A Natural Evolution Strategy for Multi-objective Optimization
Tobias Glasmachers, Tom Schaul, Jürgen Schmidhuber |
PPSN (1) | 3 |
| 2010 | PyBrain
Tom Schaul, Justin Bayer, Daan Wierstra, Yi Sun 0003, Martin Felder, Frank Sehnke, Thomas Rückstieß, Jürgen Schmidhuber |
J. Mach. Learn. Res. | 8 |
| 2010 | Deep, Big, Simple Neural Nets for Handwritten Digit RecognitionabstractGood old online backpropagation for plain multilayer perceptrons yields a very low 0.35% error rate on the MNIST handwritten digits benchmark. All we need to achieve this best result so far are many hidden layers, many neurons per layer, numerous deformed training images to avoid overfitting, and graphics cards to greatly speed up learning. Dan C. Ciresan, Ueli Meier, Luca Maria Gambardella, Jürgen Schmidhuber |
Neural Comput. | 4 |
| 2010 | Parameter-exploring policy gradients
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters 0001, Jürgen Schmidhuber |
Neural Networks | 6 |
| 2009 | Robust player imitation using multiobjective evolutionabstractThe problem of how to create NPC AI for videogames that believably imitates particular human players is addressed. Previous approaches to learning player behaviour is found to either not generalize well to new environments and noisy perceptions, or to not reproduce human behaviour in sufficient detail. It is proposed that better solutions to this problem can be built on multiobjective evolutionary algorithms, with objectives relating both to traditional progress-based fitness (playing the game well) and similarity to recorded human behaviour (behaving like the recorded player). This idea is explored in the context of a modern racing game. Niels van Hoorn, Julian Togelius, Daan Wierstra, Jürgen Schmidhuber |
IEEE Congress on Evolutionary Computation | 4 |
| 2009 | Efficient natural evolution strategiesabstractEfficient Natural Evolution Strategies (eNES) is a novel alternative to conventional evolutionary algorithms, using the natural gradient to adapt the mutation distribution. Unlike previous methods based on natural gradients, eNES uses a fast algorithm to calculate the inverse of the exact Fisher information matrix, thus increasing both robustness and performance of its evolution gradient estimation, even in higher dimensions. Additional novel aspects of eNES include optimal fitness baselines and importance mixing (a procedure for updating the population with very few fitness evaluations). The algorithm yields competitive results on both unimodal and multimodal benchmarks. Yi Sun 0003, Daan Wierstra, Tom Schaul, Jürgen Schmidhuber |
GECCO | 4 |
| 2009 | Evolving Memory Cell Structures for Sequence Learning
Justin Bayer, Daan Wierstra, Julian Togelius, Jürgen Schmidhuber |
ICANN (2) | 4 |
| 2009 | Measuring and Optimizing Behavioral Complexity for Evolutionary Reinforcement Learning
Faustino J. Gomez, Julian Togelius, Jürgen Schmidhuber |
ICANN (2) | 3 |
| 2009 | Scalable Neural Networks for Board Games
Tom Schaul, Jürgen Schmidhuber |
ICANN (1) | 2 |
| 2009 | An EM Based Training Algorithm for Recurrent Neural Networks
Jan Unkelbach, Yi Sun 0003, Jürgen Schmidhuber |
ICANN (1) | 3 |
| 2009 | Stochastic search using the natural gradientabstractTo optimize unknown 'fitness' functions, we present Natural Evolution Strategies, a novel algorithm that constitutes a principled alternative to standard stochastic search methods. It maintains a multinormal distribution on the set of solution candidates. The Natural Gradient is used to update the distribution's parameters in the direction of higher expected fitness, by efficiently calculating the inverse of the exact Fisher information matrix whereas previous methods had to use approximations. Other novel aspects of our method include optimal fitness baselines and importance mixing, a procedure adjusting batches with minimal numbers of fitness evaluations. The algorithm yields competitive results on a number of benchmarks. Yi Sun 0003, Daan Wierstra, Tom Schaul, Jürgen Schmidhuber |
ICML | 4 |
| 2009 | A reinforcement learning approach for individualizing erythropoietin dosages in hemodialysis patients
José D. Martín-Guerrero, Faustino J. Gomez, Emilio Soria-Olivas, Jürgen Schmidhuber, Mónica Climente-Martí, N. Víctor Jiménez |
Expert Syst. Appl. | 4 |
| 2009 | A Novel Connectionist System for Unconstrained Handwriting RecognitionabstractRecognizing lines of unconstrained handwritten text is a challenging task. The difficulty of segmenting cursive or overlapping characters, combined with the need to exploit surrounding context, has led to low recognition rates for even the best current recognizers. Most recent progress in the field has been made either through improved preprocessing or through advances in language modeling. Relatively little work has been done on the basic recognition algorithms. Indeed, most systems rely on the same hidden Markov models that have been used for decades in speech and handwriting recognition, despite their well-known shortcomings. This paper proposes an alternative approach based on a novel type of recurrent neural network, specifically designed for sequence labeling tasks where the data is hard to segment and contains long-range bidirectional interdependencies. In experiments on two large unconstrained handwriting databases, our approach achieves word recognition accuracies of 79.7 percent on online data and 74.1 percent on offline data, significantly outperforming a state-of-the-art HMM-based system. In addition, we demonstrate the network's robustness to lexicon size, measure the individual influence of its hidden layers, and analyze its use of context. Last, we provide an in-depth discussion of the differences between the network and HMMs, suggesting reasons for the network's superior performance. Alex Graves, Marcus Liwicki, Santiago Fernández, Roman Bertolami, Horst Bunke, Jürgen Schmidhuber |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2008 | Learning what to ignore: Memetic climbing in topology and weight spaceabstractWe present the memetic climber, a simple search algorithm that learns topology and weights of neural networks on different time scales. When applied to the problem of learning control for a simulated racing task with carefully selected inputs to the neural network, the memetic climber outperforms a standard hill-climber. When inputs to the network are less carefully selected, the difference is drastic. We also present two variations of the memetic climber and discuss the generalization of the underlying principle to population-based neuroevolution algorithms. Julian Togelius, Faustino J. Gomez, Jürgen Schmidhuber |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | Natural Evolution StrategiesabstractThis paper presents natural evolution strategies (NES), a novel algorithm for performing real-valued dasiablack boxpsila function optimization: optimizing an unknown objective function where algorithm-selected function measurements constitute the only information accessible to the method. Natural evolution strategies search the fitness landscape using a multivariate normal distribution with a self-adapting mutation matrix to generate correlated mutations in promising regions. NES shares this property with covariance matrix adaption (CMA), an evolution strategy (ES) which has been shown to perform well on a variety of high-precision optimization tasks. The natural evolution strategies algorithm, however, is simpler, less ad-hoc and more principled. Self-adaptation of the mutation matrix is derived using a Monte Carlo estimate of the natural gradient towards better expected fitness. By following the natural gradient instead of the dasiavanillapsila gradient, we can ensure efficient update steps while preventing early convergence due to overly greedy updates, resulting in reduced sensitivity to local suboptima. We show NES has competitive performance with CMA on unimodal tasks, while outperforming it on several multimodal tasks that are rich in deceptive local optima. Daan Wierstra, Tom Schaul, Jan Peters 0001, Jürgen Schmidhuber |
IEEE Congress on Evolutionary Computation | 4 |
| 2008 | Policy Gradients with Parameter-Based Exploration for Control
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters 0001, Jürgen Schmidhuber |
ICANN (1) | 6 |
| 2008 | Episodic Reinforcement Learning by Logistic Reward-Weighted Regression
Daan Wierstra, Tom Schaul, Jan Peters 0001, Jürgen Schmidhuber |
ICANN (1) | 4 |
| 2008 | Driven by Compression Progress
Jürgen Schmidhuber |
KES (1) | 1 |
| 2008 | Offline Handwriting Recognition with Multidimensional Recurrent Neural NetworksabstractOffline handwriting recognition---the transcription of images of handwritten text---is an interesting task, in that it combines computer vision with sequence learning. In most systems the two elements are handled separately, with sophisticated preprocessing techniques used to extract the image features and sequential models such as HMMs used to provide the transcriptions. By combining two recent innovations in neural networks---multidimensional recurrent neural networks and connectionist temporal classification---this paper introduces a globally trained offline handwriting recogniser that takes raw pixel data as input. Unlike competing systems, it does not require any alphabet specific preprocessing, and can therefore be used unchanged for any language. Evidence of its generality and power is provided by data from a recent international Arabic recognition competition, where it outperformed all entries (91.4% accuracy compared to 87.2% for the competition winner) despite the fact that neither author understands a word of Arabic. Alex Graves, Jürgen Schmidhuber |
NIPS | 2 |
| 2008 | State-Dependent Exploration for Policy Gradient Methods
Thomas Rückstieß, Martin Felder, Jürgen Schmidhuber |
ECML/PKDD (2) | 3 |
| 2008 | Countering Poisonous Inputs with Memetic Neuroevolution
Julian Togelius, Tom Schaul, Jürgen Schmidhuber, Faustino J. Gomez |
PPSN | 3 |
| 2008 | Fitness Expectation Maximization
Daan Wierstra, Tom Schaul, Jan Peters 0001, Jürgen Schmidhuber |
PPSN | 4 |
| 2008 | Accelerated Neural Evolution through Cooperatively Coevolved Synapses
Faustino J. Gomez, Jürgen Schmidhuber, Risto Miikkulainen |
J. Mach. Learn. Res. | 2 |
| 2007 | Simple Algorithmic Principles of Discovery, Subjective Beauty, Selective Attention, Curiosity and Creativity
Jürgen Schmidhuber |
ALT | 1 |
| 2007 | Simple Algorithmic Principles of Discovery, Subjective Beauty, Selective Attention, Curiosity & Creativity
Jürgen Schmidhuber |
Discovery Science | 1 |
| 2007 | Policy Gradient Critics
Daan Wierstra, Jürgen Schmidhuber |
ECML | 2 |
| 2007 | RNN-based Learning of Compact Maps for Efficient Robot Localization
Alexander Förster, Alex Graves, Jürgen Schmidhuber |
ESANN | 3 |
| 2007 | An Application of Recurrent Neural Networks to Discriminative Keyword Spotting
Santiago Fernández, Alex Graves, Jürgen Schmidhuber |
ICANN (2) | 3 |
| 2007 | Multi-dimensional Recurrent Neural Networks
Alex Graves, Santiago Fernández, Jürgen Schmidhuber |
ICANN (1) | 3 |
| 2007 | Solving Deep Memory POMDPs with Recurrent Policy Gradients
Daan Wierstra, Alexander Förster, Jan Peters 0001, Jürgen Schmidhuber |
ICANN (1) | 4 |
| 2007 | Sequence Labelling in Structured Domains with Hierarchical Recurrent Neural Networks
Santiago Fernández, Alex Graves, Jürgen Schmidhuber |
IJCAI | 3 |
| 2007 | Learning Restart Strategies
Matteo Gagliolo, Jürgen Schmidhuber |
IJCAI | 2 |
| 2007 | Unconstrained On-line Handwriting Recognition with Recurrent Neural NetworksabstractOn-line handwriting recognition is unusual among sequence labelling tasks in that the underlying generator of the observed data, i.e. the movement of the pen, is recorded directly. However, the raw data can be difficult to interpret because each letter is spread over many pen locations. As a consequence, sophisticated pre-processing is required to obtain inputs suitable for conventional sequence labelling algorithms, such as HMMs. In this paper we describe a system capable of directly transcribing raw on-line handwriting data. The system consists of a recurrent neural network trained for sequence labelling, combined with a probabilistic language model. In experiments on an unconstrained on-line database, we record excellent results using either raw or pre-processed data, well outperforming a benchmark HMM in both cases. Alex Graves, Santiago Fernández, Marcus Liwicki, Horst Bunke, Jürgen Schmidhuber |
NIPS | 5 |
| 2007 | Algorithmic complexity bounds on future prediction errors
Alexey V. Chernov, Marcus Hutter, Jürgen Schmidhuber |
Inf. Comput. | 3 |
| 2007 | Training Recurrent Networks by EvolinoabstractIn recent years, gradient-based LSTM recurrent neural networks (RNNs) solved many previously RNN-unlearnable tasks. Sometimes, however, gradient information is of little use for training RNNs, due to numerous local minima. For such cases, we present a novel method: EVOlution of systems with LINear Outputs (Evolino). Evolino evolves weights to the nonlinear, hidden nodes of RNNs while computing optimal linear mappings from hidden state to output, using methods such as pseudo-inverse-based linear regression. If we instead use quadratic programming to maximize the margin, we obtain the first evolutionary recurrent support vector machines. We show that Evolino-based LSTM can solve tasks that Echo State nets (Jaeger, 2004a) cannot and achieves higher accuracy in certain continuous function generation tasks than conventional gradient descent RNNs, including gradient-based LSTM. Jürgen Schmidhuber, Daan Wierstra, Matteo Gagliolo, Faustino J. Gomez |
Neural Comput. | 1 |
| 2006 | Prefix-Like Complexities and Computability in the Limit
Alexey V. Chernov, Jürgen Schmidhuber |
CiE | 2 |
| 2006 | Impact of Censored Sampling on the Performance of Restart Strategies
Matteo Gagliolo, Jürgen Schmidhuber |
CP | 2 |
| 2006 | Efficient Non-linear Control Through Neuroevolution
Faustino J. Gomez, Jürgen Schmidhuber, Risto Miikkulainen |
ECML | 2 |
| 2006 | Evolino for recurrent support vector machines
Jürgen Schmidhuber, Matteo Gagliolo, Daan Wierstra, Faustino J. Gomez |
ESANN | 1 |
| 2006 | Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networksabstractMany real-world sequence learning tasks require the prediction of sequences of labels from noisy, unsegmented input data. In speech recognition, for example, an acoustic signal is transcribed into words or sub-word units. Recurrent neural networks (RNNs) are powerful sequence learners that would seem well suited to such tasks. However, because they require pre-segmented training data, and post-processing to transform their outputs into label sequences, their applicability has so far been limited. This paper presents a novel method for training RNNs to label unsegmented sequences directly, thereby solving both problems. An experiment on the TIMIT speech corpus demonstrates its advantages over both a baseline HMM and a hybrid HMM-RNN. Alex Graves, Santiago Fernández, Faustino J. Gomez, Jürgen Schmidhuber |
ICML | 4 |
| 2006 | Quasi-online Reinforcement Learning for RobotsabstractThis paper describes quasi-online reinforcement learning: while a robot is exploring its environment, in the background a probabilistic model of the environment is built on the fly as new experiences arrive; the policy is trained concurrently based on this model using an anytime algorithm. Prioritized sweeping, directed exploration, and transformed reward functions provide additional speed-ups. The robot quickly learns goal-directed policies from scratch, requiring few interactions with the environment and making efficient use of available computation time. From an outside perspective it learns the behavior online and in real time. We describe comparisons with standard methods and show the individual utility of each of the proposed techniques Bram Bakker, Viktor Zhumatiy, Gabriel Gruener, Jürgen Schmidhuber |
ICRA | 4 |
| 2006 | A System for Robotic Heart Surgery that Learns to Tie Knots Using Recurrent Neural NetworksabstractTying suture knots is a time-consuming task performed frequently during minimally invasive surgery (MIS). Automating this task could greatly reduce total surgery time for patients. Current solutions to this problem replay manually programmed trajectories, but a more general and robust approach is to use supervised machine learning to smooth surgeon-given training trajectories and generalize from them. Since knottying generally requires a controller with internal memory to distinguish between identical inputs that require different actions at different points along a trajectory, it would be impossible to teach the system using traditional feedforward neural nets or support vector machines. Instead we exploit more powerful, recurrent neural networks (RNNs) with adaptive internal states. Results obtained using LSTM RNNs trained by the recent Evolino algorithm show that this approach can significantly increase the efficiency of suture knot tying in MIS over preprogrammed control Hermann Georg Mayer, Faustino J. Gomez, Daan Wierstra, Istvan Nagy 0002, Alois C. Knoll, Jürgen Schmidhuber |
IROS | 6 |
| 2006 | Developmental robotics, optimal artificial curiosity, creativity, music, and the fine artsabstractEven in the absence of external reward, babies and scientists and others explore their world. Using some sort of adaptive predictive world model, they improve their ability to answer questions such as what happens if I do this or that? They lose interest in both the predictable things and those predicted to remain unpredictable despite some effort. One can design curious robots that do the same. The author’s basic idea (1990, 1991) for doing so is a reinforcement learning (RL) controller is rewarded for action sequences that improve the predictor. Here, this idea is revisited in the context of recent results on optimal predictors and optimal RL machines. Several new variants of the basic principle are proposed. Finally, it is pointed out how the fine arts can be formally understood as a consequence of the principle: given some subjective observer, great works of art and music yield observation histories exhibiting more novel, previously unknown compressibility/regularity/predictability (with respect to the observer’s particular learning algorithm) than lesser works, thus deepening the observer’s understanding of the world and what is possible in it. Jürgen Schmidhuber |
Connect. Sci. | 1 |
| 2005 | Co-evolving recurrent neurons learn deep memory POMDPsabstractRecurrent neural networks are theoretically capable of learning complex temporal sequences, but training them through gradient-descent is too slow and unstable for practical use in reinforcement learning environments. Neuroevolution, the evolution of artificial neural networks using genetic algorithms, can potentially solve real-world reinforcement learning tasks that require deep use of memory, i.e. memory spanning hundreds or thousands of inputs, by searching the space of recurrent neural networks directly. In this paper, we introduce a new neuroevolution algorithm called Hierarchical Enforced SubPopulations that simultaneously evolves networks at two levels of granularity: full networks and network components or neurons. We demonstrate the method in two POMDP tasks that involve temporal dependencies of up to thousands of time-steps, and show that it is faster and simpler than the current best conventional reinforcement learning system on these tasks. Faustino J. Gomez, Jürgen Schmidhuber |
GECCO | 2 |
| 2005 | Modeling systems with internal state using evolinoabstractExisting Recurrent Neural Networks (RNNs) are limited in their ability to model dynamical systems with nonlinearities and hidden internal states. Here we use our general framework for sequence learning, EVOlution of recurrent systems with LINear Outputs (Evolino), to discover good RNN hidden node weights through evolution, while using linear regression to compute an optimal linear mapping from hidden state to output. Using the Long Short-Term Memory RNN Architecture, Evolino outperforms previous state-of-the-art methods on several tasks: 1) context-sensitive languages, 2) multiple superimposed sine waves. Categories and Subject Descriptors I.2.6 [Artificial Intelligence]: Learning—Connectionism and neural nets Daan Wierstra, Faustino J. Gomez, Jürgen Schmidhuber |
GECCO | 3 |
| 2005 | Classifying Unprompted Speech by Retraining LSTM Nets
Nicole Beringer, Alex Graves, Florian Schiel, Jürgen Schmidhuber |
ICANN (1) | 4 |
| 2005 | A Neural Network Model for Inter-problem Adaptive Online Time Allocation
Matteo Gagliolo, Jürgen Schmidhuber |
ICANN (2) | 2 |
| 2005 | Fast Color-Based Object Recognition Independent of Position and Orientation
Martijn van de Giessen, Jürgen Schmidhuber |
ICANN (1) | 2 |
| 2005 | Evolving Modular Fast-Weight Networks for Control
Faustino J. Gomez, Jürgen Schmidhuber |
ICANN (2) | 2 |
| 2005 | Bidirectional LSTM Networks for Improved Phoneme Classification and Recognition
Alex Graves, Santiago Fernández, Jürgen Schmidhuber |
ICANN (2) | 3 |
| 2005 | Completely Self-referential Optimal Reinforcement Learners
Jürgen Schmidhuber |
ICANN (2) | 1 |
| 2005 | Evolino: Hybrid Neuroevolution/Optimal Linear Search for Sequence Learning
Jürgen Schmidhuber, Daan Wierstra, Faustino J. Gomez |
IJCAI | 1 |
| 2005 | Framewise phoneme classification with bidirectional LSTM networksabstractIn this paper, we apply bidirectional training to a long short term memory (LSTM) network for the first time. We also present a modified, full gradient version of the LSTM learning algorithm. We discuss the significance of framewise phoneme classification to continuous speech recognition, and the validity of using bidirectional networks for online causal tasks. On the TIMIT speech database, we measure the framewise phoneme classification scores of bidirectional and unidirectional variants of both LSTM and conventional recurrent neural networks (RNNs). We find that bidirectional LSTM outperforms both RNNs and unidirectional LSTM. Alex Graves, Jürgen Schmidhuber |
IJCNN | 2 |
| 2005 | Framewise phoneme classification with bidirectional LSTM and other neural network architectures
Alex Graves, Jürgen Schmidhuber |
Neural Networks | 2 |
| 2004 | Adaptive Online Time Allocation to Search Algorithms
Matteo Gagliolo, Viktor Zhumatiy, Jürgen Schmidhuber |
ECML | 3 |
| 2004 | Optimal Ordered Problem Solver
Jürgen Schmidhuber |
Mach. Learn. | 1 |
| 2004 | Self-organizing nets for optimizationabstractGiven some optimization problem and a series of typically expensive trials of solution candidates sampled from a search space, how can we efficiently select the next candidate? We address this fundamental problem by embedding simple optimization strategies in learning algorithms inspired by Kohonen's self-organizing maps and neural gas networks. Our adaptive nets or grids are used to identify and exploit search space regions that maximize the probability of generating points closer to the optima. Net nodes are attracted by candidates that lead to improved evaluations, thus, quickly biasing the active data selection process toward promising regions, without loss of ability to escape from local optima. On standard benchmark functions, our techniques perform more reliably than the widely used covariance matrix adaptation evolution strategy. The proposed algorithm is also applied to the problem of drag reduction in a flow past an actively controlled circular cylinder, leading to unprecedented drag reduction. Michele Milano, Petros Koumoutsakos, Jürgen Schmidhuber |
IEEE Trans. Neural Networks | 3 |
| 2003 | A robot that reinforcement-learns to identify and memorize important previous observationsabstractIt is difficult to apply traditional reinforcement learning algorithms to robots, due to problems with large and continuous domains, partial observability, and limited numbers of learning experiences. This paper deals with these problems by combining: (1) reinforcement learning with memory, implemented using an LSTM recurrent neural network whose inputs are discrete events extracted from raw inputs; (2) online exploration and offline policy learning. An experiment with a real robot demonstrates the methodology's feasibility. Bram Bakker, Viktor Zhumatiy, Gabriel Gruener, Jürgen Schmidhuber |
IROS | 4 |
| 2003 | Kalman filters improve LSTM network performance in problems unsolvable by traditional recurrent nets
Juan Antonio Pérez-Ortiz, Felix A. Gers, Douglas Eck, Jürgen Schmidhuber |
Neural Networks | 4 |
| 2002 | The Speed Prior: A New Simplicity Measure Yielding Near-Optimal Computable Predictions
Jürgen Schmidhuber |
COLT | 1 |
| 2002 | DEKF-LSTM
Felix A. Gers, Juan Antonio Pérez-Ortiz, Douglas Eck, Jürgen Schmidhuber |
ESANN | 4 |
| 2002 | Learning the Long-Term Structure of the Blues
Douglas Eck, Jürgen Schmidhuber |
ICANN | 2 |
| 2002 | Learning Context Sensitive Languages with LSTM Trained with Kalman Filters
Felix A. Gers, Juan Antonio Pérez-Ortiz, Douglas Eck, Jürgen Schmidhuber |
ICANN | 4 |
| 2002 | Improving Long-Term Online Prediction with Decoupled Extended Kalman Filters
Juan Antonio Pérez-Ortiz, Jürgen Schmidhuber, Felix A. Gers, Douglas Eck |
ICANN | 2 |
| 2002 | Reinforcement learning in partially observable mobile robot domains using unsupervised event extractionabstractThis paper describes how learning tasks in partially observable mobile robot domains can be solved by combining reinforcement learning with an unsupervised learning "event extraction" mechanism, called ARAVQ. ARAVQ transforms the robot's continuous, noisy, high-dimensional sensory input stream into a compact sequence of high-level events. The resulting hierarchical control system uses an LSTM recurrent neural network as the reinforcement learning component, which learns high-level actions in response to the history of high-level events. The high-level actions select low-level behaviors which take care of the real-time motor control. Illustrative experiments based on the Khepera mobile robot simulator are presented. Bram Bakker, Fredrik Linåker, Jürgen Schmidhuber |
IROS | 3 |
| 2002 | Bias-Optimal Incremental Problem SolvingabstractGiven is a problem sequence and a probability distribution (the bias) on programs computing solution candidates. We present an optimally fast way of incrementally solving each task in the sequence. Bias shifts are computed by program prefixes that modify the distribution on their suf- fixes by reusing successful code for previous tasks (stored in non-modifi- able memory). No tested program gets more runtime than its probability times the total search time. In illustrative experiments, ours becomes the first general system to learn a universal solver for arbitrary disk Tow- ers of Hanoi tasks (minimal solution size ). It demonstrates the advantages of incremental learning by profiting from previously solved, simpler tasks involving samples of a simple context free language. 1 Brief Introduction to Optimal Universal Search Consider an asymptotically optimal method for tasks with quickly verifiable solutions: Jürgen Schmidhuber |
NIPS | 1 |
| 2002 | Learning Precise Timing with LSTM Recurrent Networks
Felix A. Gers, Nicol N. Schraudolph, Jürgen Schmidhuber |
J. Mach. Learn. Res. | 3 |
| 2002 | Learning Nonregular Languages: A Comparison of Simple Recurrent Networks and LSTMabstractIn response to Rodriguez's recent article (2001), we compare the performance of simple recurrent nets and long short-term memory recurrent nets on context-free and context-sensitive languages. Jürgen Schmidhuber, Felix A. Gers, Douglas Eck |
Neural Comput. | 1 |
| 2001 | Applying LSTM to Time Series Predictable through Time-Window Approaches
Felix A. Gers, Douglas Eck, Jürgen Schmidhuber |
ICANN | 3 |
| 2001 | Unsupervised Learning in LSTM Recurrent Neural Networks
Magdalena Klapper-Rybicka, Nicol N. Schraudolph, Jürgen Schmidhuber |
ICANN | 3 |
| 2001 | Market-Based Reinforcement Learning in Partially Observable Worlds
Ivo Kwee, Marcus Hutter, Jürgen Schmidhuber |
ICANN | 3 |
| 2001 | Active Learning with Adaptive Grids
Michele Milano, Jürgen Schmidhuber, Petros Koumoutsakos |
ICANN | 2 |
| 2001 | LSTM recurrent networks learn simple context-free and context-sensitive languagesabstractPrevious work on learning regular languages from exemplary training sequences showed that long short-term memory (LSTM) outperforms traditional recurrent neural networks (RNNs). We demonstrate LSTMs superior performance on context-free language benchmarks for RNNs, and show that it works even better than previous hardwired or highly specialized architectures. To the best of our knowledge, LSTM variants are also the first RNNs to learn a simple context-sensitive language, namely a(n)b(n)c(n). Felix A. Gers, Jürgen Schmidhuber |
IEEE Trans. Neural Networks | 2 |
| 2000 | Evolving strategies for active flow controlabstractRechenberg and Schwefel (Rechenberg, 1994) came up with the idea of evolution strategies for flow optimization. Since then advances in computer architectures and numerical algorithms have greatly decreased computational costs of realistic flow simulations, and today computational fluid dynamics (CFD) is complementing flow experiments as a key guiding tool for aerodynamic design. Of particular interest are designs with active devices controlling the inherently unsteady flow fields, promising potentially drastic performance leaps. We demonstrate that CFD-based design of active control strategies can benefit from evolutionary computation. We optimize the flow past an actively controlled circular cylinder, a fundamental prototypical configuration. The flow is controlled using surface-mounted vortex generators; evolutionary algorithms are used to optimize actuator placement and operating parameters. We achieve drag reduction of up to 60 percent, outperforming the best methods previously reported in the fluid dynamics literature on this benchmark problem. Michele Milano, Petros Koumoutsakos, Xavier Giannakopoulos, Jürgen Schmidhuber |
CEC | 4 |
| 2000 | Recurrent Nets that Time and CountabstractThe size of the time intervals between events conveys information essential for numerous sequential tasks such as motor control and rhythm detection. While hidden Markov models tend to ignore this information, recurrent neural networks (RNN) can in principle learn to make use of it. We focus on long short-term memory (LSTM) because it usually outperforms other RNN. Surprisingly, LSTM augmented by "peephole connections" from its internal cells to its multiplicative gates can learn the fine distinction between sequences of spikes separated by either 50 or 49 discrete time steps, without the help of any short training exemplars. Without external resets or teacher forcing or loss of performance on tasks reported earlier, our LSTM variant also learns to generate very stable sequences of highly nonlinear, precisely timed spikes. This makes LSTM a promising approach for real-world tasks that require to time and count. Felix A. Gers, Jürgen Schmidhuber |
IJCNN (3) | 2 |
| 2000 | Neural Processing of Complex Continual Input StreamsabstractLong short-term memory (LSTM) can learn algorithms for temporal pattern processing not learnable by alternative recurrent neural networks or other methods such as hidden Markov models and symbolic grammar learning. Here, we present tasks involving arithmetic operations on continual input streams that even LSTM cannot solve. However, an LSTM variant based on "forget gates," has superior arithmetic capabilities and does solve the tasks. Felix A. Gers, Jürgen Schmidhuber |
IJCNN (4) | 2 |
| 2000 | Learning to Forget: Continual Prediction with LSTMabstractLong short-term memory (LSTM; Hochreiter & Schmidhuber, 1997) can solve numerous tasks not solvable by previous learning algorithms for recurrent neural networks (RNNs). We identify a weakness of LSTM networks processing continual input streams that are not a priori segmented into subsequences with explicitly marked ends at which the network's internal state could be reset. Without resets, the state may grow indefinitely and eventually cause the network to break down. Our remedy is a novel, adaptive "forget gate" that enables an LSTM cell to learn to reset itself at appropriate times, thus releasing internal resources. We review illustrative benchmark problems on which standard LSTM outperforms other RNN algorithms. All algorithms (including LSTM) fail to solve continual versions of these problems. LSTM with forget gates, however, easily solves them, and in an elegant way. Felix A. Gers, Jürgen Schmidhuber, Fred A. Cummins |
Neural Comput. | 2 |
| 1999 | Artificial curiosity based on discovering novel algorithmic predictability through coevolutionabstractOne explores a spatio-temporal domain by predicting and learning from success/failure what's predictable and what's not. The author studies a "curious" embedded agent that differs from previous explorers in the sense that it can limit its predictions to fairly arbitrary, computable aspects of event sequences and thus can explicitly ignore almost arbitrary unpredictable, random aspects. It constructs initially random algorithms mapping event sequences to abstract internal representations (IRs). It also constructs algorithms predicting IRs from IRs computed earlier. It wants to learn novel algorithms creating IRs useful for correct IR predictions, without wasting time on those learned before. This is achieved by a co-evolutionary scheme involving two competing modules co-evolutionary designing single algorithms to be executed. The modules can bet on the outcome of IR predictions computed by the algorithms they have agreed upon. If their opinions differ then the system checks who's right, punishes the loser (the surprised one), and rewards the winner. A reinforcement learning algorithm forces each module to maximise reward. This motivates both modules to lure the other into agreeing upon algorithms involving predictions that surprise it. Since each module essentially can put in its veto against algorithms it does not consider profitable, the system is motivated to focus on those computable aspects of the environment where both modules still have confident but different opinions. Once both share the same opinion on a particular issue, the winner loses a source of reward-an incentive to shift the focus of interest onto novel, yet unknown algorithms. Jürgen Schmidhuber |
CEC | 1 |
| 1999 | Language identification from prosody without explicit featuresabstractMost current language identification (LID) systems make little or no use of prosodic information, despite the importance of prosody in LID by humans. The greatest obstacle has been that of finding an appropriate feature set which captures linguistically relevant prosodic information. The only system to attempt LID entirely on the basis of prosodic variables uses a set of over 200 features which are selected and combined in a task-specific manner [12]. We apply a novel recurrent neural network model to the task of pairwise discrimination among languages. Network inputs are limited to delta-F 0 and the first difference of the band limited amplitude envelope. Initial results are based on all pairwise combinations of English, German, Japanese, Mandarin and Spanish, with 90 speakers per language. Keywords: Language identification, Recurrent neural networks, prosody 1. PROSODY AND LANGUAGE IDENTIFICATION Most current approaches to automatic language identification use some form of segment re... Fred A. Cummins, Felix A. Gers, Jürgen Schmidhuber |
EUROSPEECH | 3 |
| 1999 | Feature Extraction Through LOCOCODEabstractLow-complexity coding and decoding (LOCOCODE) is a novel approach to sensory coding and unsupervised learning. Unlike previous methods, it explicitly takes into account the information-theoretic complexity of the code generator. It computes lococodes that convey information about the input data and can be computed and decoded by low-complexity mappings. We implement LOCOCODE by training autoassociators with flat minimum search, a recent, general method for discovering low-complexity neural nets. It turns out that this approach can unmix an unknown number of independent data sources by extracting a minimal number of low-complexity features necessary for representing the data. Experiments show that unlike codes obtained with standard autoencoders, lococodes are based on feature detectors, never unstructured, usually sparse, and sometimes factorial or local (depending on statistical properties of the data). Although LOCOCODE is not explicitly designed to enforce sparse or factorial codes, it extracts optimal codes for difficult versions of the "bars" benchmark problem, whereas independent component analysis (ICA) and principal component analysis (PCA) do not. It produces familiar, biologically plausible feature detectors when applied to real-world images and codes with fewer bits per pixel than ICA and PCA. Unlike ICA, it does not need to know the number of independent sources. As a preprocessor for a vowel recognition benchmark problem, it sets the stage for excellent classification performance. Our results reveal an interesting, previously ignored connection between two important fields: regularizer research and ICA-related research. They may represent a first step toward unification of regularization and unsupervised learning. Sepp Hochreiter, Jürgen Schmidhuber |
Neural Comput. | 2 |
| 1998 | Speeding up Q(lambda)-Learning
Marco A. Wiering, Jürgen Schmidhuber |
ECML | 2 |
| 1998 | Evolving Structured Programs with Hierarchical Instructions and Skip Nodes
Rafal Salustowicz, Jürgen Schmidhuber |
ICML | 2 |
| 1998 | Source Separation as a By-Product of Regularization
Sepp Hochreiter, Jürgen Schmidhuber |
NIPS | 2 |
| 1998 | Learning Team Strategies: Soccer Case Studies
Rafal Salustowicz, Marco A. Wiering, Jürgen Schmidhuber |
Mach. Learn. | 3 |
| 1998 | Fast Online Q(lambda)
Marco A. Wiering, Jürgen Schmidhuber |
Mach. Learn. | 2 |
| 1997 | Probabilistic Incremental Program Evolution: Stochastic Search Through Program Space
Rafal Salustowicz, Jürgen Schmidhuber |
ECML | 2 |
| 1997 | Unsupervised Coding with LOCOCODE
Sepp Hochreiter, Jürgen Schmidhuber |
ICANN | 2 |
| 1997 | On Learning Soccer Strategies
Rafal Salustowicz, Marco A. Wiering, Jürgen Schmidhuber |
ICANN | 3 |
| 1997 | Evolving Soccer Strategies
Rafal Salustowicz, Marco A. Wiering, Jürgen Schmidhuber |
ICONIP (1) | 3 |
| 1997 | Probabilistic Incremental Program EvolutionabstractProbabilistic incremental program evolution (PIPE) is a novel technique for automatic program synthesis. We combine probability vector coding of program instructions, population-based incremental learning, and tree-coded programs like those used in some variants of genetic programming (GP). PIPE iteratively generates successive populations of functional programs according to an adaptive probability distribution over all possible programs. Each iteration, it uses the best program to refine the distribution. Thus, it stochastically generates better and better programs. Since distribution refinements depend only on the best program of the current population, PIPE can evaluate program populations efficiently when the goal is to discover a program with minimal runtime. We compare PIPE to GP on a function regression problem and the 6-bit parity problem. We also use PIPE to solve tasks in partially observable mazes, where the best programs have minimal runtime. Rafal Salustowicz, Jürgen Schmidhuber |
Evol. Comput. | 2 |
| 1997 | Shifting Inductive Bias with Success-Story Algorithm, Adaptive Levin Search, and Incremental Self-Improvement
Jürgen Schmidhuber, Marco A. Wiering |
Mach. Learn. | 1 |
| 1997 | Long Short-Term MemoryabstractLearning to store information over extended time intervals by recurrent backpropagation takes a very long time, mostly because of insufficient, decaying error backflow. We briefly review Hochreiter's (1991) analysis of this problem, then address it by introducing a novel, efficient, gradient-based method called long short-term memory (LSTM). Truncating the gradient where this does not do harm, LSTM can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units. Multiplicative gate units learn to open and close access to the constant error flow. LSTM is local in space and time; its computational complexity per time step and weight is O(1). Our experiments with artificial data involve local, distributed, real-valued, and noisy pattern representations. In comparisons with real-time recurrent learning, back propagation through time, recurrent cascade correlation, Elman nets, and neural sequence chunking, LSTM leads to many more successful runs, and learns much faster. LSTM also solves complex, artificial long-time-lag tasks that have never been solved by previous recurrent network algorithms. Sepp Hochreiter, Jürgen Schmidhuber |
Neural Comput. | 2 |
| 1997 | Flat MinimaabstractWe present a new algorithm for finding low-complexity neural networks with high generalization capability. The algorithm searches for a "flat" minimum of the error function. A flat minimum is a large connected region in weight space where the error remains approximately constant. An MDL-based, Bayesian argument suggests that flat minima correspond to "simple" networks and low expected overfitting. The argument is based on a Gibbs algorithm variant and a novel way of splitting generalization error into underfitting and overfitting error. Unlike many previous approaches, ours does not require gaussian assumptions and does not depend on a "good" weight prior. Instead we have a prior over input-output functions, thus taking into account net architecture and training set. Although our algorithm requires the computation of second-order derivatives, it has backpropagation's order of complexity. Automatically, it effectively prunes units, weights, and input lines. Various experiments with feedforward and recurrent nets are described. In an application to stock market prediction, flat minimum search outperforms conventional backprop, weight decay, and "optimal brain surgeon/optimal brain damage". Sepp Hochreiter, Jürgen Schmidhuber |
Neural Comput. | 2 |
| 1997 | Discovering Neural Nets with Low Kolmogorov Complexity and High Generalization Capability
Jürgen Schmidhuber |
Neural Networks | 1 |
| 1996 | Solving POMDPs with Levin Search and EIRA
Marco A. Wiering, Jürgen Schmidhuber |
ICML | 2 |
| 1996 | LSTM can Solve Hard Long Time Lag Problems
Sepp Hochreiter, Jürgen Schmidhuber |
NIPS | 2 |
| 1996 | Semilinear Predictability Minimization Produces Well-Known Feature DetectorsabstractPredictability minimization (PM—Schmidhuber 1992) exhibits various intuitive and theoretical advantages over many other methods for unsupervised redundancy reduction. So far, however, there have not been any serious practical applications of PM. In this paper, we apply semilinear PM to static real world images and find that without a teacher and without any significant preprocessing, the system automatically learns to generate distributed representations based on well-known feature detectors, such as orientation-sensitive edge detectors and off-center–on-surround detectors, thus extracting simple features related to those considered useful for image preprocessing and compression. Jürgen Schmidhuber, Martin Eldracher, Bernhard Foltin |
Neural Comput. | 1 |
| 1996 | Sequential neural text compressionabstractThe purpose of this paper is to show that neural networks may be promising tools for data compression without loss of information. We combine predictive neural nets and statistical coding techniques to compress text files. We apply our methods to certain short newspaper articles and obtain compression ratios exceeding those of the widely used Lempel-Ziv algorithms (which build the basis of the UNIX functions "compress" and "gzip"). The main disadvantage of our methods is that they are about three orders of magnitude slower than standard methods. Jürgen Schmidhuber, Stefan Heil |
IEEE Trans. Neural Networks | 1 |
| 1995 | Discovering Solutions with Low Kolmogorov Complexity and High Generalization Capability
Jürgen Schmidhuber |
ICML | 1 |
| 1994 | Simplifying Neural Nets by Discovering Flat MinimaabstractWe present a new algorithm for finding low complexity networks with high generalization capability. The algorithm searches for large connected regions of so-called ''fiat'' minima of the error func(cid:173) tion. In the weight-space environment of a "flat" minimum, the error remains approximately constant. Using an MDL-based ar(cid:173) gument, flat minima can be shown to correspond to low expected overfitting. Although our algorithm requires the computation of second order derivatives, it has backprop's order of complexity. Experiments with feedforward and recurrent nets are described. In an application to stock market prediction, the method outperforms conventional backprop, weight decay, and "optimal brain surgeon" . Sepp Hochreiter, Jürgen Schmidhuber |
NIPS | 2 |
| 1994 | Predictive Coding with Neural Nets: Application to Text CompressionabstractTo compress text files, a neural predictor network P is used to ap(cid:173) proximate the conditional probability distribution of possible "next characters", given n previous characters. P's outputs are fed into standard coding algorithms that generate short codes for characters with high predicted probability and long codes for highly unpre(cid:173) dictable characters. Tested on short German newspaper articles, our method outperforms widely used Lempel-Ziv algorithms (used in UNIX functions such as "compress" and "gzip"). Jürgen Schmidhuber, Stefan Heil |
NIPS | 1 |
| 1993 | Discovering Predictable ClassificationsabstractPrediction problems are among the most common learning problems for neural networks (e.g., in the context of time series prediction, control, etc.). With many such problems, however, perfect prediction is inherently impossible. For such cases we present novel unsupervised systems that learn to classify patterns such that the classifications are predictable while still being as specific as possible. The approach can be related to the IMAX method of Becker and Hinton (1989) and Zemel and Hinton (1991). Experiments include a binary stereo task proposed by Becker and Hinton, which can be solved more readily by our system. Jürgen Schmidhuber, Daniel Prelinger |
Neural Comput. | 1 |
| 1992 | Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent NetworksabstractPrevious algorithms for supervised sequence learning are based on dynamic recurrent networks. This paper describes an alternative class of gradient-based systems consisting of two feedforward nets that learn to deal with temporal sequences using fast weights: The first net learns to produce context-dependent weight changes for the second net whose weights may vary very quickly. The method offers the potential for STM storage efficiency: A single weight (instead of a full-fledged unit) may be sufficient for storing temporal information. Various learning methods are derived. Two experiments with unknown time delays illustrate the approach. One experiment shows how the system can be used for adaptive temporary variable binding. Jürgen Schmidhuber |
Neural Comput. | 1 |
| 1992 | Learning Complex, Extended Sequences Using the Principle of History CompressionabstractPrevious neural network learning algorithms for sequence processing are computationally expensive and perform poorly when it comes to long time lags. This paper first introduces a simple principle for reducing the descriptions of event sequences without loss of information. A consequence of this principle is that only unexpected inputs can be relevant. This insight leads to the construction of neural architectures that learn to “divide and conquer” by recursively decomposing sequences. I describe two architectures. The first functions as a self-organizing multilevel hierarchy of recurrent networks. The second, involving only two recurrent networks, tries to collapse a multilevel predictor hierarchy into a single recurrent net. Experiments show that the system can require less computation per time step and many fewer training sequences than conventional training algorithms for recurrent nets. Jürgen Schmidhuber |
Neural Comput. | 1 |
| 1992 | A Fixed Size Storage O(n3) Time Complexity Learning Algorithm for Fully Recurrent Continually Running NetworksabstractThe real-time recurrent learning (RTRL) algorithm (Robinson and Fallside 1987; Williams and Zipser 1989) requires O(n4) computations per time step, where n is the number of noninput units. I describe a method suited for on-line learning that computes exactly the same gradient and requires fixed-size storage of the same order but has an average time complexity per time step of O(n3). Jürgen Schmidhuber |
Neural Comput. | 1 |
| 1992 | Learning Factorial Codes by Predictability MinimizationabstractI propose a novel general principle for unsupervised learning of distributed nonredundant internal representations of input patterns. The principle is based on two opposing forces. For each representational unit there is an adaptive predictor, which tries to predict the unit from the remaining units. In turn, each unit tries to react to the environment such that it minimizes its predictability. This encourages each unit to filter "abstract concepts" out of the environmental input such that these concepts are statistically independent of those on which the other units focus. I discuss various simple yet potentially powerful implementations of the principle that aim at finding binary factorial codes (Barlow et al. 1989), i.e., codes where the probability of the occurrence of a particular input is simply the product of the probabilities of the corresponding code symbols. Such codes are potentially relevant for (1) segmentation tasks, (2) speeding up supervised learning, and (3) novelty detection. Methods for finding factorial codes automatically implement Occam's razor for finding codes using a minimal number of units. Unlike previous methods the novel principle has a potential for removing not only linear but also nonlinear output redundancy. Illustrative experiments show that algorithms based on the principle of predictability minimization are practically feasible. The final part of this paper describes an entirely local algorithm that has a potential for learning unique representations of extended input sequences. Jürgen Schmidhuber |
Neural Comput. | 1 |
| 1991 | Learning Unambiguous Reduced Sequence Descriptions
Jürgen Schmidhuber |
NIPS | 1 |
| 1991 | Learning to Generate Artificial Fovea Trajectories for Target DetectionabstractThis paper shows how ‘static’ neural approaches to adaptive target detection can be replaced by a more efficient and more sequential alternative. The latter is inspired by the observation that biological systems employ sequential eye movements for pattern recognition. A system is described, which builds an adaptive model of the time-varying inputs of an artificial fovea controlled by an adaptive neural controller. The controller uses the adaptive model for learning the sequential generation of fovea trajectories causing the fovea to move to a target in a visual scene. The system also learns to track moving targets. No teacher provides the desired activations of ‘eye muscles’ at various times. The only goal information is the shape of the target. Since the task is a ‘reward-only-at-goal’ task, it involves a complex temporal credit assignment problem. Some implications for adaptive attentive systems in general are discussed. Jürgen Schmidhuber, Rudolf Huber |
Int. J. Neural Syst. | 1 |
| 1990 | An on-line algorithm for dynamic reinforcement learning and planning in reactive environmentsabstractAn online learning algorithm for reinforcement learning with continually running recurrent networks in nonstationary reactive environments is described. Various kinds of reinforcement are considered as special types of input to an agent living in the environment. The agent's only goal is to maximize the amount of reinforcement received over time. Supervised learning techniques for recurrent networks serve to construct a differentiable model of the environmental dynamics which includes a model of future reinforcement. This model is used for learning goal-directed behavior in an online fashion. The possibility of using the system for planning future action sequences is investigated and this approach is compared to approaches based on temporal difference methods. A connection to metalearning (learning how to learn) is noted Jürgen Schmidhuber |
IJCNN | 1 |
| 1990 | Reinforcement Learning in Markovian and Non-Markovian Environments
Jürgen Schmidhuber |
NIPS | 1 |