EDBT 2026 Demo / reviewers in the wild / expert
Amirhossein Roknilamouki
dblp:400/5532
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 70% Learning theory · 23% Motion planning and robot control · 7% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.9 | 1 | 2025 | Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces · ICML 2025 |
Machine learning › Reinforcement learning › markov decision process › low-rank MDP
linear MDP |
0.9 | 1 | 2025 | Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces · ICML 2025 |
Machine learning › Learning theory › online learning
regret bounds |
0.9 | 1 | 2025 | Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces · ICML 2025 |
Machine learning › Reinforcement learning › safe reinforcement learning
safe exploration |
0.9 | 1 | 2025 | Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces · ICML 2025 |
Robotics › Motion planning and robot control
collision avoidance |
0.3 | 1 | 2025 | Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
objective-constraint decomposition · 0.9least-squares value iteration · 0.9covering number bounds · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature SpacesabstractIn Reinforcement Learning (RL), tasks with instantaneous hard constraints present significant challenges, particularly when the decision space is non-convex or non-star-convex. This issue is especially relevant in domains like autonomous vehicles and robotics, where constraints such as collision avoidance often take a non-convex form. In this paper, we establish a regret bound of $\tilde{\mathcal{O}}((1 + \tfrac{1}{\tau}) \sqrt{\log(\frac{1}{\tau}) d^3 H^4 K})$, applicable to both star-convex and non-star-convex cases, where $d$ is the feature dimension, $H$ the episode length, $K$ the number of episodes, and $\tau$ the safety threshold. Moreover, the violation of safety constraints is zero with high probability throughout the learning process. A key technical challenge in these settings is bounding the covering number of the value-function class, which is essential for achieving value-aware uniform concentration in model-free function approximation. For the star-convex setting, we develop a novel technique called *Objective–Constraint Decomposition* (OCD) to properly bound the covering number. This result also resolves an error in a previous work on constrained RL. In non-star-convex scenarios, where the covering number can become infinitely large, we propose a two-phase algorithm, Non-Convex Safe Least Squares Value Iteration (NCS-LSVI), which first reduces uncertainty about the safe set by playing a known safe policy. After that, it carefully balances exploration and exploitation to achieve the regret bound. Finally, numerical simulations on an autonomous driving scenario demonstrate the effectiveness of NCS-LSVI. Amirhossein Roknilamouki, Arnob Ghosh, Ming Shi 0003, Fatemeh Nourzad, Eylem Ekici, Ness Shroff |
ICML | 1 |
| 2025 | Safe and Reliable Deep Reinforcement Learning for Covert RoutingabstractReinforcement learning (RL) holds great promise for network control problems, yet its deployment in real-world systems remains limited due to the instability and unpredictability of RL policies during training. To address this challenge, we propose a two-phase conservative RL framework that combines domain expertise from classical network optimization with modern deep RL techniques. Our key idea is to initialize the learning process with a stable base policy, derived from expert knowledge, and then apply conservative fine-tuning under a Kullback–Leibler (KL) divergence constraint to safely explore improved behaviors. We apply this framework to the problem of covert multi-hop routing, where the objective is to optimize data throughput while minimizing detectability by adversaries. In Phase I, we construct a reliable base policy by imitating the back-pressure algorithm, which guarantees throughput-optimal behavior and stable queue dynamics. Phase II fine-tunes this policy to improve covert performance, as measured by the Detection Error Probability (DEP), while preserving training-time stability. Empirical evaluations on a grid network show that our method enables more reliable learning than pure RL. While pure RL (e.g., PPO) can sometimes achieve higher covert performance, it frequently suffers from large queues and collapsed throughput during training. In our experiments, our conservative RL framework reduces the worst-case training-time queue length by over 99% while maintaining comparable covert communication performance. Amirhossein Roknilamouki, Fikadu T. Dagefu, Eylem Ekici, Justin Kong 0001, Terrence J. Moore, Yin Sun 0001, Ness Shroff |
MASS | 1 |