EDBT 2026 Demo / reviewers in the wild / expert
Simon Schmitt
dblp:61/3969
· DBLP profile ↗
18ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 55% Trustworthy machine learning · 27% Deep learning architectures and training · 11% | |
| Network and information security
1 paper |
Systems and software security · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › uncertainty estimation
epistemic uncertainty |
1.5 | 2 | 2025 | General Uncertainty Estimation with Delta Variances · AAAI 2025 Exploration via Epistemic Value Estimation · AAAI 2023 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | General Uncertainty Estimation with Delta Variances · AAAI 2025 |
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | Exploration via Epistemic Value Estimation · AAAI 2023 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.6 | 1 | 2022 | Chaining Value Functions for Off-Policy Learning · AAAI 2022 |
Machine learning › Deep learning architectures and training › neural network training
backpropagation-free training |
0.5 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Machine learning › Deep learning architectures and training › feedforward neural network › piecewise linear network
gated linear networks |
0.5 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.5 | 1 | 2021 | Learning and Planning in Complex Action Spaces · ICML 2021 |
Machine learning › Learning theory
online learning |
0.5 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Machine learning › Reinforcement learning
planning and learning |
0.5 | 1 | 2021 | Learning and Planning in Complex Action Spaces · ICML 2021 |
Machine learning › Reinforcement learning
policy optimization |
0.5 | 1 | 2021 | Muesli: Combining Improvements in Policy Optimization · ICML 2021 |
Machine learning › Reinforcement learning
actor-critic methods |
0.4 | 1 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.4 | 1 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Machine learning › Reinforcement learning › actor-critic methods
off-policy actor-critic |
0.4 | 1 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.4 | 1 | 2019 | Multi-Task Deep Reinforcement Learning with PopArt · AAAI 2019 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.4 | 1 | 2019 | Multi-Task Deep Reinforcement Learning with PopArt · AAAI 2019 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.4 | 1 | 2019 | Multi-Task Deep Reinforcement Learning with PopArt · AAAI 2019 |
Systems and software security › memory safety
memory corruption |
0.3 | 1 | 2018 | K-Miner: Uncovering Memory Corruption in Linux · NDSS 2018 |
Systems and software security
vulnerability discovery |
0.3 | 1 | 2018 | K-Miner: Uncovering Memory Corruption in Linux · NDSS 2018 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.1 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Program analysis
static analysis |
0.1 | 1 | 2018 | K-Miner: Uncovering Memory Corruption in Linux · NDSS 2018 |
Methods — techniques the papers use, named apart from their topics
gradient computation · 0.9delta variance · 0.9static analysis · 0.7epistemic uncertainty estimation · 0.7bayesian neural network · 0.7temporal difference learning · 0.6bootstrapping · 0.6online convex optimization · 0.5muzero · 0.5model learning · 0.5data-dependent gating · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | General Uncertainty Estimation with Delta VariancesabstractDecision makers may suffer from uncertainty induced by limited data. This may be mitigated by accounting for epistemic uncertainty, which is however challenging to estimate efficiently for large neural networks. To this extent we investigate Delta Variances, a family of algorithms for epistemic uncertainty quantification, that is computationally efficient and convenient to implement. It can be applied to neural networks and more general functions composed of neural networks. As an example we consider a weather simulator with a neural-network-based step function inside - here Delta Variances empirically obtain competitive results at the cost of a single gradient computation. The approach is convenient as it requires no changes to the neural network architecture or training procedure. We discuss multiple ways to derive Delta Variances theoretically noting that special cases recover popular techniques and present a unified perspective on multiple related methods. Finally we observe that this general perspective gives rise to a natural extension and empirically show its benefit. Simon Schmitt, John Shawe-Taylor, Hado van Hasselt |
AAAI | 1 |
| 2023 | Exploration via Epistemic Value EstimationabstractHow to efficiently explore in reinforcement learning is an open problem. Many exploration algorithms employ the epistemic uncertainty of their own value predictions -- for instance to compute an exploration bonus or upper confidence bound. Unfortunately the required uncertainty is difficult to estimate in general with function approximation. We propose epistemic value estimation (EVE): a recipe that is compatible with sequential decision making and with neural network function approximators. It equips agents with a tractable posterior over all their parameters from which epistemic value uncertainty can be computed efficiently. We use the recipe to derive an epistemic Q-Learning agent and observe competitive performance on a series of benchmarks. Experiments confirm that the EVE recipe facilitates efficient exploration in hard exploration tasks. Simon Schmitt, John Shawe-Taylor, Hado van Hasselt |
AAAI | 1 |
| 2022 | Chaining Value Functions for Off-Policy LearningabstractTo accumulate knowledge and improve its policy of behaviour, a reinforcement learning agent can learn `off-policy' about policies that differ from the policy used to generate its experience. This is important to learn counterfactuals, or because the experience was generated out of its own control. However, off-policy learning is non-trivial, and standard reinforcement-learning algorithms can be unstable and divergent. In this paper we discuss a novel family of off-policy prediction algorithms which are convergent by construction. The idea is to first learn on-policy about the data-generating behaviour, and then bootstrap an off-policy value estimate on this on-policy estimate, thereby constructing a value estimate that is partially off-policy. This process can be repeated to build a chain of value functions, each time bootstrapping a new estimate on the previous estimate in the chain. Each step in the chain is stable and hence the complete algorithm is guaranteed to be stable. Under mild conditions this comes arbitrarily close to the off-policy TD solution when we increase the length of the chain. Hence it can compute the solution even in cases where off-policy TD diverges. We prove that the proposed scheme is convergent and corresponds to an iterative decomposition of the inverse key matrix. Furthermore it can be interpreted as estimating a novel objective -- that we call a `k-step expedition' -- of following the target policy for finitely many steps before continuing indefinitely with the behaviour policy. Empirically we evaluate the idea on challenging MDPs such as Baird's counter example and observe favourable results. Simon Schmitt, John Shawe-Taylor, Hado van Hasselt |
AAAI | 1 |
| 2021 | Gated Linear NetworksabstractThis paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs). What distinguishes GLNs from contemporary neural networks is the distributed and local nature of their credit assignment mechanism; each neuron directly predicts the target, forgoing the ability to learn feature representations in favor of rapid online learning. Individual neurons are able to model nonlinear functions via the use of data-dependent gating in conjunction with online convex optimization. We show that this architecture gives rise to universal learning capabilities in the limit, with effective model capacity increasing as a function of network size in a manner comparable with deep ReLU networks. Furthermore, we demonstrate that the GLN learning mechanism possesses extraordinary resilience to catastrophic forgetting, performing almost on par to an MLP with dropout and Elastic Weight Consolidation on standard benchmarks. Joel Veness, Tor Lattimore, David Budden, Avishkar Bhoopchand, Christopher Mattern, Agnieszka Grabska-Barwinska, Eren Sezener, Peter Toth, Simon Schmitt, Marcus Hutter |
AAAI | 10 |
| 2021 | Muesli: Combining Improvements in Policy OptimizationabstractWe propose a novel policy update that combines regularized policy optimization with model learning as an auxiliary loss. The update (henceforth Muesli) matches MuZero’s state-of-the-art performance on Atari. Notably, Muesli does so without using deep search: it acts directly with a policy network and has computation speed comparable to model-free baselines. The Atari results are complemented by extensive ablations, and by additional results on continuous control and 9x9 Go. Matteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez, Simon Schmitt, Laurent Sifre, Theophane Weber, David Silver 0001, Hado van Hasselt |
ICML | 5 |
| 2021 | Learning and Planning in Complex Action SpacesabstractMany important real-world problems have action spaces that are high-dimensional, continuous or both, making full enumeration of all possible actions infeasible. Instead, only small subsets of actions can be sampled for the purpose of policy evaluation and improvement. In this paper, we propose a general framework to reason in a principled way about policy evaluation and improvement over such sampled action subsets. This sample-based policy iteration framework can in principle be applied to any reinforcement learning algorithm based upon policy iteration. Concretely, we propose Sampled MuZero, an extension of the MuZero algorithm that is able to learn in domains with arbitrarily complex action spaces by planning over sampled actions. We demonstrate this approach on the classical board game of Go and on two continuous control benchmark domains: DeepMind Control Suite and Real-World RL Suite. Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain, Simon Schmitt, David Silver 0001 |
ICML | 5 |
| 2020 | Off-Policy Actor-Critic with Shared Experience ReplayabstractWe investigate the combination of actor-critic reinforcement learning algorithms with a uniform large-scale experience replay and propose solutions for two ensuing challenges: (a) efficient actor-critic learning with experience replay (b) the stability of off-policy learning where agents learn from other agents behaviour. To this end we analyze the bias-variance tradeoffs in V-trace, a form of importance sampling for actor-critic methods. Based on our analysis, we then argue for mixing experience sampled from replay with on-policy experience, and propose a new trust region scheme that scales effectively to data distributions where V-trace becomes unstable. We provide extensive empirical validation of the proposed solutions on DMLab-30 and further show the benefits of this setup in two training regimes for Atari: (1) a single agent is trained up until 200M environment frames per game (2) a population of agents is trained up until 200M environment frames each and may share experience. We demonstrate state-of-the-art data efficiency among model-free agents in both regimes. Simon Schmitt, Matteo Hessel, Karen Simonyan |
ICML | 1 |
| 2019 | Multi-Task Deep Reinforcement Learning with PopArtabstractThe reinforcement learning (RL) community has made great strides in designing algorithms capable of exceeding human performance on specific tasks. These algorithms are mostly trained one task at the time, each new task requiring to train a brand new agent instance. This means the learning algorithm is general, but each solution is not; each agent can only solve the one task it was trained on. In this work, we study the problem of learning to master not one but multiple sequentialdecision tasks at once. A general issue in multi-task learning is that a balance must be found between the needs of multiple tasks competing for the limited resources of a single learning system. Many learning algorithms can get distracted by certain tasks in the set of tasks to solve. Such tasks appear more salient to the learning process, for instance because of the density or magnitude of the in-task rewards. This causes the algorithm to focus on those salient tasks at the expense of generality. We propose to automatically adapt the contribution of each task to the agent’s updates, so that all tasks have a similar impact on the learning dynamics. This resulted in state of the art performance on learning to play all games in a set of 57 diverse Atari games. Excitingly, our method learned a single trained policy - with a single set of weights - that exceeds median human performance. To our knowledge, this was the first time a single agent surpassed human-level performance on this multi-task domain. The same approach also demonstrated state of the art performance on a set of 30 tasks in the 3D reinforcement learning platform DeepMind Lab. Matteo Hessel, Hubert Soyer, Lasse Espeholt, Wojciech Czarnecki 0001, Simon Schmitt, Hado van Hasselt |
AAAI | 5 |
| 2018 | K-Miner: Uncovering Memory Corruption in Linux
David Gens, Simon Schmitt, Lucas Davi, Ahmad-Reza Sadeghi |
NDSS | 2 |
| 2017 | Fast routing graph extraction from floor plansabstractA routing graph allows to find paths in buildings quickly. Raster images of floor plans are simple to obtain but display poor performance. A manually constructed graph is quite optimal if designed by an informed person, but the process is time consuming and expensive. We describe a fast method to calculate a 2D routing graph from raster images. We adapt image processing techniques and apply a conditional erosion technique. A conditional query of every pixel by means of a predefined 3×3 image matrix calculates an approximation of common walkways through corridors and rooms. We translate the image processing idea into propositional logic formulas and simplify them. Compared to our previous work, we reduce the run time by about 90 % and the amount of needed matrices from 29 to eight or even four, depending on the specific application. We also present a parallel version at the cost of redundant edges in the resulting routing graph. Simon Schmitt, Larissa Zech, Katinka Wolter, Thomas Willemsen, Harald Sternberg, Marcel Kyas |
IPIN | 1 |
| 2016 | Conditional erosion to estimate routing graph out of floor plansabstractSystems for indoor navigation differ substantially in implementation and maintenance effort as well as in costs. A system must work on any smart phone to ensure broad adoption and avoid isolated solutions. It must also work as automated as possible. A routing graph is commonly used for path planning. But generally, no routing graph exists and it must be computed. We propose a method to compute a routing graph from floor plans. We use conditional erosion to extract the graph. An approximation to the common routes through corridors and rooms can be calculated by a conditional query of every pixel of the grid based floor plan by means of predefined 3 × 3 image matrices. The grid data is then converted to edges and nodes. We evaluate the method on existing floor plan data of a test building of the HafenCity University of Hamburg. Simon Schmitt, Larissa Zech, Thomas Willemsen, Harald Sternberg, Marcel Kyas |
IPIN | 1 |
| 2015 | A survey of experimental evaluation in indoor localization researchabstractDuring the last decade, research in indoor localization and navigation has focused on techniques, protocols, and algorithms. The first International Conference on Indoor Positioning and Indoor Navigation (IPIN) was held in 2010. Since then, this annual conference showed the progress of research and technology. The variations of evaluation methods are significant in this field: they range from none, to extensive simulations, and real-world experiments under non-lab conditions. We look at the articles published in the proceedings of IPIN by IEEE Xplore from 2010 to 2014, and analyze the development of evaluation methods. We categorized 183 randomly selected papers, in respect to five different aspects. Namely: (1) the underlying system/technology in use, (2) the evaluation method for the proposed technique, (3) the method of ground truth data gathering, (4) the applied metrics, and (5) whether the authors establish a baseline for their work. Stephan Adler, Simon Schmitt, Katinka Wolter, Marcel Kyas |
IPIN | 2 |
| 2014 | Device-free indoor localisation using radio tomography imaging in 800/900 MHz bandabstractRadio tomographic imaging (RTI) can be used as a method for device free localisation of persons in rooms. By measuring the signal strength of all links of a network of sensor nodes, one can estimate the position of an attenuating object with reasonable precision. The sub-GHz is shown to be suitable for an implementation. Such a design is more energy efficient than a 2.4 GHz implementation. We adapt, evaluate, and improve the device-free localization method of Wilson and Patwari to indoor environments using the 800/900 MHz band. The advantage of using 800/900 MHz is reduced reflections compared to 2.4 GHz. At the same time, the signal is attenuated less by objects in its path. Thus, the methods of Wilson and Patwari needed to be refined and parameters needed to be adapted. We evaluate some combinations of the most common choices of norms to perform a Tikhonov regularisation. The first difference gradient operator with H1 norm works best. We equipped a 5 m×5 m room with 20 wireless sensor nodes. We evaluated the influence of the distance to walls and the height of nodes. The radio tomographic image are post-processed by filters to make likely positions of objects more apparent. Additional parameters values suggested by Wilson and Patwari could not be used for the new frequency and hardware. For example, the ellipse excess path length has been experimentally determined to be close to 20 cm instead of 2 cm. In our experiments-up, we achieve an a maximum average localisation error below of 78 cm. With this work, we have reproduced the results of Wilson and Patwari, adapted it to a different frequency in the 800/900 MHz band and developed improvements to the original algorithms. Stephan Adler, Simon Schmitt, Marcel Kyas |
IPIN | 2 |
| 2014 | Experimental evaluation of indoor localization algorithmsabstractIn Radio Frequency (RF)-based indoor localization scenarios, localization algorithms are needed to alleviate the impact of non-line-of-sight and multipath effects on the measurements and thereby estimate the true position precisely. Several resilient lateration algorithms have been proposed in the last couple of years which claim to minimize these effects. However, most of these algorithms were only evaluated using simulations or small static testbeds. We conducted an experiment using 25 anchor nodes and a mobile node installed on top a robotic reference system to collect ranging values. The robot has a localization error of 6.5cm which is an order lower than our range measurement errors. We use this robot to collect range measurements and ground truth positions along a densely grid with approx. 10 cm spacing. The experiment was carried out in a hallway of our office-like building. We collected data on approx. 300 m2. First, we examine the influence of the anchor placement and anchor density on the ranging errors we see. Then, we evaluate and analyze the robustness of localization algorithms on our measured data to decide which one works best for a constellation of anchor placement and building. Our results show, that there are significant differences between the simulations published for lateration algorithms and actual experiments in real-world indoor localization scenarios. As we show in this paper, the distance measurement error distribution has a large influence on these algorithms. Stephan Adler, Simon Schmitt, Yuan Yang 0005, Yubin Zhao, Marcel Kyas |
IPIN | 2 |
| 2014 | The effects of human body shadowing in RF-based indoor localizationabstractIn radio frequency based indoor human localization systems with body mounted sensors, the human body can cause non-line-of-sight (NLOS) effects which might result in severe range estimation and localization errors. However, previous studies on the impact of the human body only conducted static experiments in controlled environments. We confirm known effects and conduct real-world experiments in a typical indoor human tracking scenario using 2.4 GHz time of flight (TOF) range measurements. We analyze the effect on the raw measurements and on the localization results using the localization algorithms Centroid, NLLS, MD-Min-Max, and Geo-n. The experiment design is focused on incident management, where an infrastructure might only be installed in front of the building. We show that these effects have considerable impact on the localization accuracy of the person. Simon Schmitt, Stephan Adler, Marcel Kyas |
IPIN | 1 |
| 2013 | Virtual testbed for indoor localizationabstractWe present a novel, easy to use virtual testbed for the evaluation of localization algorithms. Our testbed enables researchers to easily run tests on a huge body of real world range-based indoor localization data. The data consists of a dense grid of reference points belonging to one or multiple maps. Each point consists of a ground truth value and an arbitrary number of ranging values. Each ranging value belongs to a certain anchor node on a fixed position. The reference data is gathered by a robot which carries (arbitrary) localization devices. The robot stores its location as a ground truth value and simultaneously uses the localization device to measure the distance to a set of anchors in range. The ground truth value is gathered by an optical reference system which is applied to the robot. It is possible to define paths through a map using a web interface. Our system uses our experimental gathered reference points to deliver a dataset of ranging values for the current path. Therefore the researcher can run a virtual experiment by himself and can adjust several parameters. Our system enables other researchers to run reproducible experiments on real word data. The expensive and complex deployment of a dedicated infrastructure and experimental setup can be avoided as well as the error-prone task of modelling a localization system and running a simulation. Our system will be open to the research community and will help to develop a better understanding of the field of range based indoor localization. Stephan Adler, Simon Schmitt, Heiko Will, Thomas Hillebrandt, Marcel Kyas |
IPIN | 2 |
| 2013 | A virtual indoor localization testbed for Wireless Sensor NetworksabstractWe present a novel, easy to use virtual testbed enabling researchers to evaluate their localization algorithms based on distance measurements in indoor environments. We provide precise ground truth information collected by our previously presented reference system, based on a mobile robot, in combination with range measurements from a Wireless Sensor Network (WSN) in multiple buildings. The user can define a virtual experiment using our datasets over the web. This approach separates the process of designing a robust indoor localization algorithm from the need for testing it on actual hardware in different scenarios. Simon Schmitt, Heiko Will, Thomas Hillebrandt, Marcel Kyas |
SECON | 1 |
| 2012 | A reference system for indoor localization testbedsabstractWe present a low-cost robot system capable of performing robust indoor localization while carrying components of another system which shall be evaluated. Using off-the-shelf components, the ground truth positioning data provided by the robot can be used to evaluate a variety of localization systems and algorithms. Not needing any pre-installed components in its environment, it is very easy to setup. The robot system relies on wheel-odometry data of a Roomba robot, and visual distance measurements of two Kinects. The Robot Operating System (ROS) is used for the localization process according to a precise pre-drawn floor plan that may be enhanced with Simultaneous Localization and Mapping (SLAM). The system is able to estimate its position with an average error of 6.7 cm. It records its own positioning data as well as the data from the system under evaluation and provides simple means for analysis. It is also able to re-drive a previous test run if reproducable conditions are needed. Simon Schmitt, Heiko Will, Benjamin Aschenbrenner, Thomas Hillebrandt, Marcel Kyas |
IPIN | 1 |