EDBT 2026 Demo / reviewers in the wild / expert
Rahul V. Kulkarni
dblp:14/9367
· DBLP profile ↗
6ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0007-2875-2238ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 87% Representation and self-supervised learning · 13% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reward design › reward shaping
potential-based reward shaping |
1.5 | 2 | 2025 | Bootstrapped Reward Shaping · AAAI 2025 Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement Learning · AAAI 2023 |
Machine learning › Reinforcement learning › reward design
reward shaping |
1.5 | 2 | 2025 | Bootstrapped Reward Shaping · AAAI 2025 Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement Learning · AAAI 2023 |
Machine learning › Reinforcement learning
value function estimation |
0.9 | 1 | 2025 | Bootstrapped Reward Shaping · AAAI 2025 |
Machine learning › Reinforcement learning
maximum entropy reinforcement learning |
0.7 | 1 | 2023 | Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement Learning · AAAI 2023 |
Machine learning › Reinforcement learning
task composition |
0.7 | 1 | 2023 | Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement Learning · AAAI 2023 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.4 | 1 | 2019 | A Free Energy Based Approach for Distance Metric Learning · KDD 2019 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.4 | 1 | 2019 | A Free Energy Based Approach for Distance Metric Learning · KDD 2019 |
Methods — techniques the papers use, named apart from their topics
bootstrapping · 0.9soft value functions · 0.7von neumann entropy · 0.4statistical mechanics · 0.4boltzmann distribution · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bootstrapped Reward ShapingabstractIn reinforcement learning, especially in sparse-reward domains, many environment steps are required to observe reward information. In order to increase the frequency of such observations, "potential-based reward shaping" (PBRS) has been proposed as a method of providing a more dense reward signal while leaving the optimal policy invariant. However, the required potential function must be carefully designed with task-dependent knowledge to not deter training performance. In this work, we propose a bootstrapped method of reward shaping, termed BS-RS, in which the agent's current estimate of the state-value function acts as the potential function for PBRS. We provide convergence proofs for the tabular setting, give insights into training dynamics for deep RL, and show that the proposed method improves training speed in the Atari suite. Jacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V. Kulkarni |
AAAI | 4 |
| 2023 | Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement LearningabstractIn reinforcement learning (RL), the ability to utilize prior knowledge from previously solved tasks can allow agents to quickly solve new problems. In some cases, these new problems may be approximately solved by composing the solutions of previously solved primitive tasks (task composition). Otherwise, prior knowledge can be used to adjust the reward function for a new problem, in a way that leaves the optimal policy unchanged but enables quicker learning (reward shaping). In this work, we develop a general framework for reward shaping and task composition in entropy-regularized RL. To do so, we derive an exact relation connecting the optimal soft value functions for two entropy-regularized RL problems with different reward functions and dynamics. We show how the derived relation leads to a general result for reward shaping in entropy-regularized RL. We then generalize this approach to derive an exact relation connecting optimal value functions for the composition of multiple tasks in entropy-regularized RL. We validate these theoretical contributions with experiments showing that reward shaping and task composition lead to faster learning in various settings. Jacob Adamczyk, Argenis Arriojas, Stas Tiomkin, Rahul V. Kulkarni |
AAAI | 4 |
| 2023 | Bounding the optimal value function in compositional reinforcement learningabstractIn the field of reinforcement learning (RL), agents are often tasked with solving a variety of problems differing only in their reward functions. In order to quickly obtain solutions to unseen problems with new reward functions, a popular approach involves functional composition of previously solved tasks. However, previous work using such functional composition has primarily focused on specific instances of composition functions, whose limiting assumptions allow for exact zero-shot composition. Our work unifies these examples and provides a more general framework for compositionality in both standard and entropy-regularized RL. We find that, for a broad class of functions, the optimal solution for the composite task of interest can be related to the known primitive task solutions. Specifically, we present double-sided inequalities relating the optimal composite value function to the value functions for the primitive tasks. We also show that the regret of using a zero-shot policy can be bounded for this class of functions. The derived bounds can be used to develop clipping approaches for reducing uncertainty during training, allowing agents to quickly adapt to new tasks. Jacob Adamczyk, Volodymyr Makarenko, Argenis Arriojas, Stas Tiomkin, Rahul V. Kulkarni |
UAI | 5 |
| 2023 | Bayesian inference approach for entropy regularized reinforcement learning with stochastic dynamicsabstractWe develop a novel approach to determine the optimal policy in entropy-regularized reinforcement learning (RL) with stochastic dynamics. For deterministic dynamics, the optimal policy can be derived using Bayesian inference in the control-as-inference framework; however, for stochastic dynamics, the direct use of this approach leads to risk-taking optimistic policies. To address this issue, current approaches in entropy-regularized RL involve a constrained optimization procedure which fixes system dynamics to the original dynamics, however this approach is not consistent with the unconstrained Bayesian inference framework. In this work we resolve this inconsistency by developing an exact mapping from the constrained optimization problem in entropy-regularized RL to a different optimization problem which can be solved using the unconstrained Bayesian inference approach. We show that the optimal policies are the same for both problems, thus our results lead to the exact solution for the optimal policy in entropy-regularized RL with stochastic dynamics through Bayesian inference. Argenis Arriojas, Jacob Adamczyk, Stas Tiomkin, Rahul V. Kulkarni |
UAI | 4 |
| 2019 | A Free Energy Based Approach for Distance Metric LearningabstractWe present a reformulation of the distance metric learning problem as a penalized optimization problem, with a penalty term corresponding to the von Neumann entropy of the distance metric. This formulation leads to a mapping to statistical mechanics such that the metric learning optimization problem becomes equivalent to free energy minimization. Correspondingly, our approach leads to an analytical solution of the optimization problem based on the Boltzmann distribution. The mapping established in this work suggests new approaches for dimensionality reduction and provides insights into determination of optimal parameters for the penalty term. Furthermore, we demonstrate that the metric projects the data onto direction of maximum dissimilarity with optimal and tunable separation between classes and thus the transformation can be used for high dimensional data visualization, classification, and clustering tasks. We benchmark our method against previous distance learning methods and provide an efficient implementation in an R package available to download at: https://github.com/kouroshz/fenn. Sho Inaba, Carl Tony Fakhry, Rahul V. Kulkarni, Kourosh Zarringhalam |
KDD | 3 |
| 2015 | Transcriptional Bursting in Gene Expression: Analytical Results for General Stochastic ModelsabstractGene expression in individual cells is highly variable and sporadic, often resulting in the synthesis of mRNAs and proteins in bursts. Such bursting has important consequences for cell-fate decisions in diverse processes ranging from HIV-1 viral infections to stem-cell differentiation. It is generally assumed that bursts are geometrically distributed and that they arrive according to a Poisson process. On the other hand, recent single-cell experiments provide evidence for complex burst arrival processes, highlighting the need for analysis of more general stochastic models. To address this issue, we invoke a mapping between general stochastic models of gene expression and systems studied in queueing theory to derive exact analytical expressions for the moments associated with mRNA/protein steady-state distributions. These results are then used to derive noise signatures, i.e. explicit conditions based entirely on experimentally measurable quantities, that determine if the burst distributions deviate from the geometric distribution or if burst arrival deviates from a Poisson process. For non-Poisson arrivals, we develop approaches for accurate estimation of burst parameters. The proposed approaches can lead to new insights into transcriptional bursting based on measurements of steady-state mRNA/protein distributions. Niraj Kumar 0006, Abhyudai Singh, Rahul V. Kulkarni |
PLoS Comput. Biol. | 3 |