Georgios Tzannetos

dblp:345/8576 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 64% Learning theory · 16% Transfer learning and domain adaptation · 16%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
deep reinforcement learning
0.812024
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning · IJCAI 2024
Machine learning › Reinforcement learning › reinforcement learning from human feedback
learning from human feedback
0.812024
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences · ICML 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.812024
Learning Embeddings for Sequential Tasks Using Population of Agents · IJCAI 2024
Machine learning › Reinforcement learning
population-based learning
0.812024
Learning Embeddings for Sequential Tasks Using Population of Agents · IJCAI 2024
Machine learning › Learning theory
statistical guarantees
0.812024
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences · ICML 2024
Machine learning › Transfer learning and domain adaptation
task embedding
0.812024
Learning Embeddings for Sequential Tasks Using Population of Agents · IJCAI 2024
Natural language and speech › Language models and text generation › preference optimization
preference-based policy optimization
0.212024
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences · ICML 2024

Methods — techniques the papers use, named apart from their topics

task correlations · 0.8proximal curriculum · 0.8minimax analysis · 0.8loglinear policy parametrization · 0.8embedding learning · 0.8agent population · 0.8
YearPublicationVenuePosition
2025 Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
abstract
Training agents to operate under strict constraints during deployment, such as limited resource budgets or stringent safety requirements, presents significant challenges, especially when these constraints render the task complex. In this work, we propose a curriculum learning strategy that gradually tightens constraints during training, enabling the agent to incrementally master the deployment requirements. Inspired by self-paced learning techniques in unconstrained reinforcement learning (RL), our approach facilitates a smoother transition to challenging environments by initially training on simplified versions of the constraints and progressively introducing the full deployment conditions. We provide a theoretical analysis using an RL agent in a binary-tree Markov Decision Process (MDP) to demonstrate that our curriculum strategy can accelerate training relative to a baseline approach that imposes the trajectory constraints from the outset. Moreover, we empirically validate the effectiveness and generality of our method across both RL and large language model (LLM) agents in diverse settings, including a binary-tree MDP, a multi-task navigation domain, and a math reasoning task with two benchmarks. These results highlight the potential of curriculum design in enhancing the efficiency and performance of agents operating under complex trajectory constraints during deployment. Moreover, when applied to LLMs, our strategy enables compression of output chain-of-thought tokens, achieving a substantial inference speedup on consumer hardware, demonstrating its effectiveness for resource-constrained deployment.
Georgios Tzannetos, Parameswaran Kamalaruban, Adish Singla
NeurIPS1
2024 Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
abstract
In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO). We focus our attention on the class of loglinear policy parametrization and linear reward functions. In order to compare the two paradigms, we first derive minimax statistical bounds on the suboptimality gap induced by both RLHF and DPO, assuming access to an oracle that exactly solves the optimization problems. We provide a detailed discussion on the relative comparison between the two paradigms, simultaneously taking into account the sample size, policy and reward class dimensions, and the regularization temperature. Moreover, we extend our analysis to the approximate optimization setting and derive exponentially decaying convergence rates for both RLHF and DPO. Next, we analyze the setting where the ground-truth reward is not realizable and find that, while RLHF incurs a constant additional error, DPO retains its asymptotically decaying gap by just tuning the temperature accordingly. Finally, we extend our comparison to the Markov decision process setting, where we generalize our results with exact optimization. To the best of our knowledge, we are the first to provide such a comparative analysis for RLHF and DPO.
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Georgios Tzannetos, Goran Radanovic, Adish Singla
ICML4
2024 Learning Embeddings for Sequential Tasks Using Population of Agents
Mridul Mahajan, Georgios Tzannetos, Goran Radanovic, Adish Singla
IJCAI2
2024 Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
Georgios Tzannetos, Parameswaran Kamalaruban, Adish Singla
IJCAI1