EDBT 2026 Demo / reviewers in the wild / expert
Shyam Sundhar Ramesh
dblp:331/3550
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 31% Trustworthy machine learning · 31% Optimization for machine learning · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Energy systems and smart grids · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › fairness
group robustness |
0.8 | 1 | 2024 | Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024 |
Natural language and speech › Language models and text generation
preference optimization |
0.8 | 1 | 2024 | Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.8 | 1 | 2024 | Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.6 | 1 | 2022 | Movement Penalized Bayesian Optimization with Application to Wind Energy Systems · NeurIPS 2022 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
contextual bayesian optimization |
0.6 | 1 | 2022 | Movement Penalized Bayesian Optimization with Application to Wind Energy Systems · NeurIPS 2022 |
Energy systems and smart grids › renewable energy
wind energy |
0.2 | 1 | 2022 | Movement Penalized Bayesian Optimization with Application to Wind Energy Systems · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
online learning · 1.1mirror descent · 1.1gaussian process · 1.1worst-case group optimization · 0.8direct preference optimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Distributionally Robust Model-based Reinforcement Learning with Large State SpacesabstractThree major challenges in reinforcement learning are the complex dynamical systems with large state spaces, the costly data acquisition processes, and the deviation of real-world dynamics from the training environment deployment. To overcome these issues, we study distributionally robust Markov decision processes with continuous state spaces under the widely used Kullback-Leibler, chi-square, and total variation uncertainty sets. We propose a model-based approach that utilizes Gaussian Processes and the maximum variance reduction algorithm to efficiently learn multi-output nominal transition dynamics, leveraging access to a generative model (i.e., simulator). We further demonstrate the statistical sample complexity of the proposed method for different uncertainty sets. These complexity bounds are independent of the number of states and extend beyond linear dynamics, ensuring the effectiveness of our approach in identifying near-optimal distributionally-robust policies. The proposed method can be further combined with other model-free distributionally robust reinforcement learning methods to obtain a near-optimal robust policy. Experimental results demonstrate the robustness of our algorithm to distributional shifts and its superior performance in terms of the number of samples needed. Shyam Sundhar Ramesh, Pier Giuseppe Sessa, Andreas Krause 0001, Ilija Bogunovic |
AISTATS | 1 |
| 2024 | Group Robust Preference Optimization in Reward-free RLHFabstractAdapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional RLHF approaches adopt a "one-size-fits-all" approach, i.e., they indiscriminately assume and optimize a single preference model, thus not being robust to unique characteristics and needs of the various groups. To address this limitation, we propose a novel Group Robust Preference Optimization (GRPO) method to align LLMs to individual groups' preferences robustly. Our approach builds upon reward-free direct preference optimization methods, but unlike previous approaches, it seeks a robust policy which maximizes the worst-case group performance. To achieve this, GRPO adaptively and sequentially weights the importance of different groups, prioritizing groups with worse cumulative loss. We theoretically study the feasibility of GRPO and analyze its convergence for the log-linear policy class. By fine-tuning LLMs with GRPO using diverse group-based global opinion data, we significantly improved performance for the worst-performing groups, reduced loss imbalances across groups, and improved probability accuracies compared to non-robust baselines. Shyam Sundhar Ramesh, Iason Chaimalas, Viraj Mehta, Pier Giuseppe Sessa, Haitham Bou-Ammar, Ilija Bogunovic |
NeurIPS | 1 |
| 2022 | Movement Penalized Bayesian Optimization with Application to Wind Energy SystemsabstractContextual Bayesian optimization (CBO) is a powerful framework for sequential decision-making given side information, with important applications, e.g., in wind energy systems. In this setting, the learner receives context (e.g., weather conditions) at each round, and has to choose an action (e.g., turbine parameters). Standard algorithms assume no cost for switching their decisions at every round. However, in many practical applications, there is a cost associated with such changes, which should be minimized. We introduce the episodic CBO with movement costs problem and, based on the online learning approach for metrical task systems of Coester and Lee (2019), propose a novel randomized mirror descent algorithm that makes use of Gaussian Process confidence bounds. We compare its performance with the offline optimal sequence for each episode and provide rigorous regret guarantees. We further demonstrate our approach on the important real-world application of altitude optimization for Airborne Wind Energy Systems. In the presence of substantial movement costs, our algorithm consistently outperforms standard CBO algorithms. Shyam Sundhar Ramesh, Pier Giuseppe Sessa, Andreas Krause 0001, Ilija Bogunovic |
NeurIPS | 1 |