EDBT 2026 Demo / reviewers in the wild / expert
Silviu Pitis
dblp:212/1241
· DBLP profile ↗
13ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 9 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Reinforcement learning · 56% Language models and text generation · 20% Trustworthy machine learning · 8% | |
| Network and information security
1 paper |
Systems and software security · 100% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
counterfactual data augmentation |
1.0 | 2 | 2022 | MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022 Counterfactual Data Augmentation using Locally Factored Dynamics · NeurIPS 2020 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
1.0 | 2 | 2022 | MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022 Counterfactual Data Augmentation using Locally Factored Dynamics · NeurIPS 2020 |
Machine learning › Reinforcement learning
sample efficiency |
1.0 | 2 | 2022 | MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022 Counterfactual Data Augmentation using Locally Factored Dynamics · NeurIPS 2020 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models · NeurIPS 2025 |
Machine learning › Reinforcement learning
temporal difference learning |
0.8 | 2 | 2020 | Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning · AAAI 2020 Source Traces for Temporal Difference Learning · AAAI 2018 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Improving Context-Aware Preference Modeling for Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation
LLM agents |
0.8 | 1 | 2024 | Identifying the Risks of LM Agents with an LM-Emulated Sandbox · ICLR 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning › preference handling
preference modeling |
0.8 | 1 | 2024 | Improving Context-Aware Preference Modeling for Language Models · NeurIPS 2024 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.8 | 1 | 2024 | Improving Context-Aware Preference Modeling for Language Models · NeurIPS 2024 |
Systems and software security
vulnerability discovery |
0.8 | 1 | 2024 | Identifying the Risks of LM Agents with an LM-Emulated Sandbox · ICLR 2024 |
Natural language and speech › Language models and text generation
large language model |
0.7 | 1 | 2023 | Large Language Models are Human-Level Prompt Engineers · ICLR 2023 |
Machine learning › Reinforcement learning
multi-objective reinforcement learning |
0.7 | 1 | 2023 | Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards · NeurIPS 2023 |
Machine learning › Reinforcement learning › reward design
non-markovian reward |
0.7 | 1 | 2023 | Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards · NeurIPS 2023 |
Natural language and speech › Language models and text generation › prompting
prompt engineering |
0.7 | 1 | 2023 | Large Language Models are Human-Level Prompt Engineers · ICLR 2023 |
Machine learning › Reinforcement learning
reward design |
0.7 | 1 | 2023 | Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards · NeurIPS 2023 |
Robotics › Motion planning and robot control
dynamic modeling |
0.6 | 1 | 2022 | MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.6 | 1 | 2022 | MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.6 | 1 | 2022 | MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022 |
Machine learning › Reinforcement learning › exploration › information-theoretic exploration
entropy-based exploration |
0.4 | 1 | 2020 | Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020 |
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2020 | Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020 |
Machine learning › Learning theory
inductive bias |
0.4 | 1 | 2020 | An Inductive Bias for Distances: Neural Nets that Respect the Triangle Inequality · ICLR 2020 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.4 | 1 | 2020 | Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020 |
Machine learning › Reinforcement learning
long-horizon tasks |
0.4 | 1 | 2020 | Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.4 | 1 | 2020 | An Inductive Bias for Distances: Neural Nets that Respect the Triangle Inequality · ICLR 2020 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
multi-goal reinforcement learning |
0.4 | 1 | 2020 | Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.4 | 1 | 2020 | Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning · AAAI 2020 |
Machine learning › Reinforcement learning
value function approximation |
0.4 | 1 | 2020 | Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning · AAAI 2020 |
Machine learning › Reinforcement learning
discount factor |
0.4 | 1 | 2019 | Rethinking the Discount Factor in Reinforcement Learning: A Decision Theoretic Approach · AAAI 2019 |
Machine learning › Reinforcement learning › temporal difference learning
eligibility traces |
0.3 | 1 | 2018 | Source Traces for Temporal Difference Learning · AAAI 2018 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
successor representation |
0.3 | 1 | 2018 | Source Traces for Temporal Difference Learning · AAAI 2018 |
Methods — techniques the papers use, named apart from their topics
multi-turn benchmark · 1.7interactive clinical vignettes · 1.7language model emulation · 1.5automatic safety evaluation · 1.5counterfactual data augmentation · 1.0experience replay · 0.8preference dataset · 0.8fine-tuning · 0.8large language model · 0.7automatic prompt optimization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language ModelsabstractClinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks for large language models (LLMs) primarily assess knowledge recall through single-turn questions, where complete clinical information is provided upfront. To address this gap, we introduce VivaBench, a multi-turn benchmark that evaluates sequential clinical reasoning in LLM agents. Our dataset consists of 1762 physician-curated clinical vignettes structured as interactive scenarios that simulate a $ \textit{viva voce}$ (oral) examination in medical training, requiring agents to actively probe for relevant findings, select appropriate investigations, and synthesize information across multiple steps to reach a diagnosis. While current LLMs demonstrate competence in diagnosing conditions from well-described clinical presentations, their performance degrades significantly when required to navigate iterative diagnostic reasoning under uncertainty in our evaluation. Our analysis identified several failure modes that mirror common cognitive errors in clinical practice, including: (1) fixation on initial hypotheses, (2) inappropriate investigation ordering, (3) premature diagnostic closure, and (4) failing to screen for critical conditions. These patterns reveal fundamental limitations in how current LLMs reason and make decisions under uncertainty. Through VivaBench, we provide a standardized benchmark for evaluating conversational medical AI systems for real-world clinical decision support. Beyond medical applications, we contribute to the larger corpus of research on agentic AI by demonstrating how sequential reasoning trajectories can diverge in complex decision-making environments. Christopher Chiu, Silviu Pitis, Mihaela van der Schaar |
NeurIPS | 2 |
| 2024 | Identifying the Risks of LM Agents with an LM-Emulated SandboxabstractRecent advances in Language Model (LM) agents and tool use, exemplified by applications like ChatGPT Plugins, enable a rich set of capabilities but also amplify potential risks—such as leaking private data or causing financial losses. Identifying these risks is labor-intensive, necessitating implementing the tools, setting up the environment for each test scenario manually, and finding risky cases. As tools and agents become more complex, the high cost of testing these agents will make it increasingly difficult to find high-stakes, long-tail risks. To address these challenges, we introduce ToolEmu: a framework that uses an LM to emulate tool execution and enables scalable testing of LM agents against a diverse range of tools and scenarios. Alongside the emulator, we develop an LM-based automatic safety evaluator that examines agent failures and quantifies associated risks. We test both the tool emulator and evaluator through human evaluation and find that 68.8% of failures identified with ToolEmu would be valid real-world agent failures. Using our curated initial benchmark consisting of 36 high-stakes toolkits and 144 test cases, we provide a quantitative risk analysis of current LM agents and identify numerous failures with potentially severe outcomes. Notably, even the safest LM agent exhibits such failures 23.9% of the time according to our evaluator, underscoring the need to develop safer LM agents for real-world deployment. Yangjun Ruan, Honghua Dong, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, Tatsunori B. Hashimoto |
ICLR | 4 |
| 2024 | Improving Context-Aware Preference Modeling for Language ModelsabstractWhile finetuning language models from pairwise preferences has proven remarkably effective, the underspecified nature of natural language presents critical challenges. Direct preference feedback is uninterpretable, difficult to provide where multidimensional criteria may apply, and often inconsistent, either because it is based on incomplete instructions or provided by diverse principals. To address these challenges, we consider the two-step preference modeling procedure that first resolves the under-specification by selecting a context, and then evaluates preference with respect to the chosen context. We decompose reward modeling error according to these two steps, which suggests that supervising context in addition to context-specific preference may be a viable approach to aligning models with diverse human preferences. For this to work, the ability of models to evaluate context-specific preference is critical. To this end, we contribute context-conditioned preference datasets and accompanying experiments that investigate the ability of language models to evaluate context-specific preference. Unlike past datasets, where context-specific preference is highly correlated with general preference, our "preference reversal" datasets disentangle context-specific and general preferences to isolate context-specific capabilities. We use our datasets to (1) show that existing preference models benefit from, but fail to fully consider, added context, (2) finetune a context-aware reward model with context-specific performance exceeding that of GPT-4 and Llama 3 70B, and (3) investigate the potential value of context-aware preference modeling. Silviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro Sordoni |
NeurIPS | 1 |
| 2023 | Large Language Models are Human-Level Prompt Engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, Jimmy Ba |
ICLR | 5 |
| 2023 | Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsabstractAs the capabilities of artificial agents improve, they are being increasingly deployed to service multiple diverse objectives and stakeholders. However, the composition of these objectives is often performed ad hoc, with no clear justification. This paper takes a normative approach to multi-objective agency: from a set of intuitively appealing axioms, it is shown that Markovian aggregation of Markovian reward functions is not possible when the time preference (discount factor) for each objective may vary. It follows that optimal multi-objective agents must admit rewards that are non-Markovian with respect to the individual objectives. To this end, a practical non-Markovian aggregation scheme is proposed, which overcomes the impossibility with only one additional parameter for each objective. This work offers new insights into sequential, multi-objective agency and intertemporal choice, and has practical implications for the design of AI systems deployed to serve multiple generations of principals with varying time preference. Silviu Pitis |
NeurIPS | 1 |
| 2022 | MoCoDA: Model-based Counterfactual Data AugmentationabstractThe number of states in a dynamic process is exponential in the number of objects, making reinforcement learning (RL) difficult in complex, multi-object domains. For agents to scale to the real world, they will need to react to and reason about unseen combinations of objects. We argue that the ability to recognize and use local factorization in transition dynamics is a key element in unlocking the power of multi-object reasoning. To this end, we show that (1) known local structure in the environment transitions is sufficient for an exponential reduction in the sample complexity of training a dynamics model, and (2) a locally factored dynamics model provably generalizes out-of-distribution to unseen states and actions. Knowing the local structure also allows us to predict which unseen states and actions this dynamics model will generalize to. We propose to leverage these observations in a novel Model-based Counterfactual Data Augmentation (MoCoDA) framework. MoCoDA applies a learned locally factored dynamics model to an augmented distribution of states and actions to generate counterfactual transitions for RL. MoCoDA works with a broader set of local structures than prior work and allows for direct control over the augmented training distribution. We show that MoCoDA enables RL agents to learn policies that generalize to unseen states and actions. We use MoCoDA to train an offline RL agent to solve an out-of-distribution robotics manipulation task on which standard offline RL algorithms fail. Silviu Pitis, Elliot Creager, Ajay Mandlekar, Animesh Garg |
NeurIPS | 1 |
| 2020 | Fixed-Horizon Temporal Difference Methods for Stable Reinforcement LearningabstractWe explore fixed-horizon temporal difference (TD) methods, reinforcement learning algorithms for a new kind of value function that predicts the sum of rewards over a fixed number of future time steps. To learn the value function for horizon h, these algorithms bootstrap from the value function for horizon h−1, or some shorter horizon. Because no value function bootstraps from itself, fixed-horizon methods are immune to the stability problems that plague other off-policy TD methods using function approximation (also known as “the deadly triad”). Although fixed-horizon methods require the storage of additional value functions, this gives the agent additional predictive power, while the added complexity can be substantially reduced via parallel updates, shared weights, and n-step bootstrapping. We show how to use fixed-horizon value functions to solve reinforcement learning problems competitively with methods such as Q-learning that learn conventional value functions. We also prove convergence of fixed-horizon temporal difference methods with linear and general function approximation. Taken together, our results establish fixed-horizon TD methods as a viable new way of avoiding the stability problems of the deadly triad. Kristopher De Asis, Alan Chan 0001, Silviu Pitis, Richard S. Sutton, Daniel Graves |
AAAI | 3 |
| 2020 | An Inductive Bias for Distances: Neural Nets that Respect the Triangle Inequality
Silviu Pitis, Harris Chan, Kiarash Jamali, Jimmy Ba |
ICLR | 1 |
| 2020 | Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningabstractWhat goals should a multi-goal reinforcement learning agent pursue during training in long-horizon tasks? When the desired (test time) goal distribution is too distant to offer a useful learning signal, we argue that the agent should not pursue unobtainable goals. Instead, it should set its own intrinsic goals that maximize the entropy of the historical achieved goal distribution. We propose to optimize this objective by having the agent pursue past achieved goals in sparsely explored areas of the goal space, which focuses exploration on the frontier of the achievable goal set. We show that our strategy achieves an order of magnitude better sample efficiency than the prior state of the art on long-horizon multi-goal tasks including maze navigation and block stacking. Silviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie, Jimmy Ba |
ICML | 1 |
| 2020 | Counterfactual Data Augmentation using Locally Factored DynamicsabstractMany dynamic processes, including common scenarios in robotic control and reinforcement learning (RL), involve a set of interacting subprocesses. Though the subprocesses are not independent, their interactions are often sparse, and the dynamics at any given time step can often be decomposed into locally independent} causal mechanisms. Such local causal structures can be leveraged to improve the sample efficiency of sequence prediction and off-policy reinforcement learning. We formalize this by introducing local causal models (LCMs), which are induced from a global causal model by conditioning on a subset of the state space. We propose an approach to inferring these structures given an object-oriented state representation, as well as a novel algorithm for Counterfactual Data Augmentation (CoDA). CoDA uses local structures and an experience replay to generate counterfactual experiences that are causally valid in the global model. We find that CoDA significantly improves the performance of RL agents in locally factored tasks, including the batch-constrained and goal-conditioned settings. Code available at https://github.com/spitis/mrl. Silviu Pitis, Elliot Creager, Animesh Garg |
NeurIPS | 1 |
| 2019 | Rethinking the Discount Factor in Reinforcement Learning: A Decision Theoretic ApproachabstractReinforcement learning (RL) agents have traditionally been tasked with maximizing the value function of a Markov decision process (MDP), either in continuous settings, with fixed discount factor γ Silviu Pitis |
AAAI | 1 |
| 2018 | Source Traces for Temporal Difference LearningabstractThis paper motivates and develops source traces for temporal difference (TD) learning in the tabular setting. Source traces are like eligibility traces, but model potential histories rather than immediate ones. This allows TD errors to be propagated to potential causal states and leads to faster generalization. Source traces can be thought of as the model-based, backward view of successor representations (SR), and share many of the same benefits. This view, however, suggests several new ideas. First, a TD(λ)-like source learning algorithm is proposed and its convergence is proven. Then, a novel algorithm for learning the source map (or SR matrix) is developed and shown to outperform the previous algorithm. Finally, various approaches to using the source/SR model are explored, and it is shown that source traces can be effectively combined with other model-based methods like Dyna and experience replay. Silviu Pitis |
AAAI | 1 |
| 2017 | Methods for retrieving alternative contract language using a prototypeabstractThis paper addresses the problem of searching for alternative contract language that is similar to, yet different from, a given provision (the prototype). While this is a core task in transactional legal work, generic search solutions do not offer an effective solution. We draw upon modern information retrieval research to propose and validate novel methods for retrieving alternative language using a prototype. Our solution accepts an entire provision as a prototype and retrieves variants on the language from a database of precedent contracts. In designing this solution, we propose two ordered proximity measures and demonstrate their effectiveness relative to existing techniques. Further, we examine the challenge posed by varying definitions of redundant search results and propose to resolve it with a user-tunable, dynamic approach to result clustering. Silviu Pitis |
ICAIL | 1 |