Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shyam Sundhar Ramesh

dblp:331/3550 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 31% Trustworthy machine learning · 31% Optimization for machine learning · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Energy systems and smart grids · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.812024
Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024
Machine learning › Trustworthy machine learning
fairness
0.812024
Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024
Machine learning › Trustworthy machine learning › fairness
group robustness
0.812024
Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024
Natural language and speech › Language models and text generation
preference optimization
0.812024
Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Group Robust Preference Optimization in Reward-free RLHF · NeurIPS 2024
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.612022
Movement Penalized Bayesian Optimization with Application to Wind Energy Systems · NeurIPS 2022
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
contextual bayesian optimization
0.612022
Movement Penalized Bayesian Optimization with Application to Wind Energy Systems · NeurIPS 2022
Energy systems and smart grids › renewable energy
wind energy
0.212022
Movement Penalized Bayesian Optimization with Application to Wind Energy Systems · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

online learning · 1.1mirror descent · 1.1gaussian process · 1.1worst-case group optimization · 0.8direct preference optimization · 0.8
YearPublicationVenuePosition
2024 Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
abstract
Three major challenges in reinforcement learning are the complex dynamical systems with large state spaces, the costly data acquisition processes, and the deviation of real-world dynamics from the training environment deployment. To overcome these issues, we study distributionally robust Markov decision processes with continuous state spaces under the widely used Kullback-Leibler, chi-square, and total variation uncertainty sets. We propose a model-based approach that utilizes Gaussian Processes and the maximum variance reduction algorithm to efficiently learn multi-output nominal transition dynamics, leveraging access to a generative model (i.e., simulator). We further demonstrate the statistical sample complexity of the proposed method for different uncertainty sets. These complexity bounds are independent of the number of states and extend beyond linear dynamics, ensuring the effectiveness of our approach in identifying near-optimal distributionally-robust policies. The proposed method can be further combined with other model-free distributionally robust reinforcement learning methods to obtain a near-optimal robust policy. Experimental results demonstrate the robustness of our algorithm to distributional shifts and its superior performance in terms of the number of samples needed.
Shyam Sundhar Ramesh, Pier Giuseppe Sessa, Andreas Krause 0001, Ilija Bogunovic
AISTATS1
2024 Group Robust Preference Optimization in Reward-free RLHF
abstract
Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional RLHF approaches adopt a "one-size-fits-all" approach, i.e., they indiscriminately assume and optimize a single preference model, thus not being robust to unique characteristics and needs of the various groups. To address this limitation, we propose a novel Group Robust Preference Optimization (GRPO) method to align LLMs to individual groups' preferences robustly. Our approach builds upon reward-free direct preference optimization methods, but unlike previous approaches, it seeks a robust policy which maximizes the worst-case group performance. To achieve this, GRPO adaptively and sequentially weights the importance of different groups, prioritizing groups with worse cumulative loss. We theoretically study the feasibility of GRPO and analyze its convergence for the log-linear policy class. By fine-tuning LLMs with GRPO using diverse group-based global opinion data, we significantly improved performance for the worst-performing groups, reduced loss imbalances across groups, and improved probability accuracies compared to non-robust baselines.
Shyam Sundhar Ramesh, Iason Chaimalas, Viraj Mehta, Pier Giuseppe Sessa, Haitham Bou-Ammar, Ilija Bogunovic
NeurIPS1
2022 Movement Penalized Bayesian Optimization with Application to Wind Energy Systems
abstract
Contextual Bayesian optimization (CBO) is a powerful framework for sequential decision-making given side information, with important applications, e.g., in wind energy systems. In this setting, the learner receives context (e.g., weather conditions) at each round, and has to choose an action (e.g., turbine parameters). Standard algorithms assume no cost for switching their decisions at every round. However, in many practical applications, there is a cost associated with such changes, which should be minimized. We introduce the episodic CBO with movement costs problem and, based on the online learning approach for metrical task systems of Coester and Lee (2019), propose a novel randomized mirror descent algorithm that makes use of Gaussian Process confidence bounds. We compare its performance with the offline optimal sequence for each episode and provide rigorous regret guarantees. We further demonstrate our approach on the important real-world application of altitude optimization for Airborne Wind Energy Systems. In the presence of substantial movement costs, our algorithm consistently outperforms standard CBO algorithms.
Shyam Sundhar Ramesh, Pier Giuseppe Sessa, Andreas Krause 0001, Ilija Bogunovic
NeurIPS1