Ralf Schoknecht

dblp:60/881 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
0since 2021 · last 2004
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 82% Optimization for machine learning · 16% Learning theory · 2%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 50% Software maintenance and evolution · 50%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › function approximation
linear function approximation
0.132004
Convergence of synchronous reinforcement learning with linear function approximation · ICML 2004
Convergent Combinations of Reinforcement Learning with Linear Function Approximation · NIPS 2002
Optimality of Reinforcement Learning Algorithms with Linear Function Approximation · NIPS 2002
Machine learning › Reinforcement learning
value function approximation
0.132002
Optimality of Reinforcement Learning Algorithms with Linear Function Approximation · NIPS 2002
A Necessary Condition of Convergence for Reinforcement Learning with Function Approximation · ICML 2002
Convergent Combinations of Reinforcement Learning with Linear Function Approximation · NIPS 2002
Machine learning › Optimization for machine learning
convergence analysis
0.122003
TD(0) Converges Provably Faster than the Residual Gradient Algorithm · ICML 2003
Convergent Combinations of Reinforcement Learning with Linear Function Approximation · NIPS 2002
Machine learning › Reinforcement learning
temporal difference learning
0.122003
TD(0) Converges Provably Faster than the Residual Gradient Algorithm · ICML 2003
Convergent Combinations of Reinforcement Learning with Linear Function Approximation · NIPS 2002
Machine learning › Reinforcement learning › reinforcement learning theory
convergence of reinforcement learning
0.012004
Convergence of synchronous reinforcement learning with linear function approximation · ICML 2004
Compilers and program optimization
dead code elimination
0.012004
Using Machine Learning for Estimating the Defect Content After an Inspection · IEEE Trans. Software Eng. 2004
Software maintenance and evolution
software inspection
0.012004
Using Machine Learning for Estimating the Defect Content After an Inspection · IEEE Trans. Software Eng. 2004
Machine learning › Reinforcement learning
function approximation
0.012002
A Necessary Condition of Convergence for Reinforcement Learning with Function Approximation · ICML 2002
Machine learning › Reinforcement learning
policy evaluation
0.012002
Optimality of Reinforcement Learning Algorithms with Linear Function Approximation · NIPS 2002

Methods — techniques the papers use, named apart from their topics

linear function approximation · 0.1nonlinear regression · 0.0neural network · 0.0inhomogeneous matrix iteration · 0.0cross-validation · 0.0counterexample construction · 0.0residual gradient algorithm · 0.0synchronous updates · 0.0projection operator · 0.0function approximation · 0.0
YearPublicationVenuePosition
2004 Convergence of synchronous reinforcement learning with linear function approximation
abstract
Synchronous reinforcement learning (RL) algorithms with linear function approximation are representable as inhomogeneous matrix iterations of a special form (Schoknecht & Merke, 2003). In this paper we state conditions of convergence for general inhomogeneous matrix iterations and prove that they are both necessary and sufficient. This result extends the work presented in (Schoknecht & Merke, 2003), where only a sufficient condition of convergence was proved. As the condition of convergence is necessary and sufficient, the new result is suitable to prove convergence and divergence of RL algorithms with function approximation. We use the theorem to deduce a new concise proof of convergence for the synchronous residual gradient algorithm (Baird, 1995). Moreover, we derive a counterexample for which the uniform RL algorithm (Merke & Schoknecht, 2002) diverges. This yields a negative answer to the open question if the uniform RL algorithm converges for arbitrary multiple transitions.
Artur Merke, Ralf Schoknecht
ICML2
2004 Fynesse: An architecture for integrating prior knowledge in autonomously learning agents
Ralf Schoknecht, Martin Spott, Martin A. Riedmiller
Soft Comput.1
2004 Using Machine Learning for Estimating the Defect Content After an Inspection
abstract
We view the problem of estimating the defect content of a document after an inspection as a machine learning problem: The goal is to learn from empirical data the relationship between certain observable features of an inspection (such as the total number of different defects detected) and the number of defects actually contained in the document. We show that some features can carry significant nonlinear information about the defect content. Therefore, we use a nonlinear regression technique, neural networks, to solve the learning problem. To select the best among all neural networks trained on a given data set, one usually reserves part of the data set for later cross-validation; in contrast, we use a technique which leaves the full data set for training. This is an advantage when the data set is small. We validate our approach on a known empirical inspection data set. For that benchmark, our novel approach clearly outperforms both linear regression and the current standard methods in software engineering for estimating the defect content, such as capture-recapture. The validation also shows that our machine learning approach can be successful even when the empirical inspection data set is small.
Frank Padberg, Thomas Ragg, Ralf Schoknecht
IEEE Trans. Software Eng.3
2003 Learning to Control at Multiple Time Scales
Ralf Schoknecht, Martin A. Riedmiller
ICANN1
2003 TD(0) Converges Provably Faster than the Residual Gradient Algorithm
Ralf Schoknecht, Artur Merke
ICML1
2003 Reinforcement learning on explicitly specified time scales
Ralf Schoknecht, Martin A. Riedmiller
Neural Comput. Appl.1
2002 Applying Machine Learning to Solve an Estimation Problem in Software Inspections
Thomas Ragg, Frank Padberg, Ralf Schoknecht
ICANN3
2002 Speeding-up Reinforcement Learning with Multi-step Actions
Ralf Schoknecht, Martin A. Riedmiller
ICANN1
2002 A Necessary Condition of Convergence for Reinforcement Learning with Function Approximation
Artur Merke, Ralf Schoknecht
ICML2
2002 Optimality of Reinforcement Learning Algorithms with Linear Function Approximation
abstract
There are several reinforcement learning algorithms that yield ap(cid:173) proximate solutions for the problem of policy evaluation when the value function is represented with a linear function approximator. In this paper we show that each of the solutions is optimal with respect to a specific objective function. Moreover, we characterise the different solutions as images of the optimal exact value func(cid:173) tion under different projection operations. The results presented here will be useful for comparing the algorithms in terms of the error they achieve relative to the error of the optimal approximate solution.
Ralf Schoknecht
NIPS1
2002 Convergent Combinations of Reinforcement Learning with Linear Function Approximation
abstract
Convergence for iterative reinforcement learning algorithms like TD(O) depends on the sampling strategy for the transitions. How(cid:173) ever, in practical applications it is convenient to take transition data from arbitrary sources without losing convergence. In this paper we investigate the problem of repeated synchronous updates based on a fixed set of transitions. Our main theorem yields suffi(cid:173) cient conditions of convergence for combinations of reinforcement learning algorithms and linear function approximation. This allows to analyse if a certain reinforcement learning algorithm and a cer(cid:173) tain function approximator are compatible. For the combination of the residual gradient algorithm with grid-based linear interpolation we show that there exists a universal constant learning rate such that the iteration converges independently of the concrete transi(cid:173) tion data.
Ralf Schoknecht, Artur Merke
NIPS1