EDBT 2026 Demo / reviewers in the wild / expert
Sumaita Sadia Rahman
dblp:400/7396
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 50% Planning, search and constraint satisfaction · 25% Learning paradigms · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
curriculum learning |
0.9 | 1 | 2025 | Training a Generally Curious Agent · ICML 2025 |
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | Training a Generally Curious Agent · ICML 2025 |
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context adaptation |
0.9 | 1 | 2025 | Training a Generally Curious Agent · ICML 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
information gathering |
0.9 | 1 | 2025 | Training a Generally Curious Agent · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
synthetic interaction data · 0.9fine-tuning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Training a Generally Curious AgentabstractEfficient exploration is essential for intelligent systems interacting with their environment, but existing language models often fall short in scenarios that require strategic information gathering. In this paper, we present **Paprika**, a fine-tuning approach that enables language models to develop general decision-making capabilities that are not confined to particular environments. By training on synthetic interaction data from different tasks that require diverse strategies, Paprika teaches models to explore and adapt their behavior on a new task based on environment feedback in-context without more gradient updates. Experimental results show that models fine-tuned with Paprika can effectively transfer their learned decision-making capabilities to entirely unseen tasks without additional training. Unlike traditional training, our approach's primary bottleneck lies in sampling useful interaction data instead of model updates. To improve sample efficiency, we propose a curriculum learning strategy that prioritizes sampling trajectories from tasks with high learning potential. These results suggest a promising path towards AI systems that can autonomously solve novel sequential decision-making problems that require interactions with the external world. Fahim Tajwar, Yiding Jiang, Abitha Thankaraj, Sumaita Sadia Rahman, J. Zico Kolter, Jeff G. Schneider, Ruslan Salakhutdinov |
ICML | 4 |