Silviu Pitis

dblp:212/1241 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 9 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Reinforcement learning · 56% Language models and text generation · 20% Trustworthy machine learning · 8%
Network and information security
1 paper
Systems and software security · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
counterfactual data augmentation
1.022022
MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022
Counterfactual Data Augmentation using Locally Factored Dynamics · NeurIPS 2020
Machine learning › Reinforcement learning
model-based reinforcement learning
1.022022
MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022
Counterfactual Data Augmentation using Locally Factored Dynamics · NeurIPS 2020
Machine learning › Reinforcement learning
sample efficiency
1.022022
MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022
Counterfactual Data Augmentation using Locally Factored Dynamics · NeurIPS 2020
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models · NeurIPS 2025
Machine learning › Reinforcement learning
temporal difference learning
0.822020
Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning · AAAI 2020
Source Traces for Temporal Difference Learning · AAAI 2018
Natural language and speech › Language models and text generation
alignment
0.812024
Improving Context-Aware Preference Modeling for Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation
LLM agents
0.812024
Identifying the Risks of LM Agents with an LM-Emulated Sandbox · ICLR 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning › preference handling
preference modeling
0.812024
Improving Context-Aware Preference Modeling for Language Models · NeurIPS 2024
Machine learning › Reinforcement learning › reward learning
reward modeling
0.812024
Improving Context-Aware Preference Modeling for Language Models · NeurIPS 2024
Systems and software security
vulnerability discovery
0.812024
Identifying the Risks of LM Agents with an LM-Emulated Sandbox · ICLR 2024
Natural language and speech › Language models and text generation
large language model
0.712023
Large Language Models are Human-Level Prompt Engineers · ICLR 2023
Machine learning › Reinforcement learning
multi-objective reinforcement learning
0.712023
Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards · NeurIPS 2023
Machine learning › Reinforcement learning › reward design
non-markovian reward
0.712023
Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards · NeurIPS 2023
Natural language and speech › Language models and text generation › prompting
prompt engineering
0.712023
Large Language Models are Human-Level Prompt Engineers · ICLR 2023
Machine learning › Reinforcement learning
reward design
0.712023
Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards · NeurIPS 2023
Robotics › Motion planning and robot control
dynamic modeling
0.612022
MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022
Machine learning › Reinforcement learning
offline reinforcement learning
0.612022
MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.612022
MoCoDA: Model-based Counterfactual Data Augmentation · NeurIPS 2022
Machine learning › Reinforcement learning › exploration › information-theoretic exploration
entropy-based exploration
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Reinforcement learning
exploration
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Learning theory
inductive bias
0.412020
An Inductive Bias for Distances: Neural Nets that Respect the Triangle Inequality · ICLR 2020
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Reinforcement learning
long-horizon tasks
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.412020
An Inductive Bias for Distances: Neural Nets that Respect the Triangle Inequality · ICLR 2020
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
multi-goal reinforcement learning
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Reinforcement learning
value-based reinforcement learning
0.412020
Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning · AAAI 2020
Machine learning › Reinforcement learning
value function approximation
0.412020
Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning · AAAI 2020
Machine learning › Reinforcement learning
discount factor
0.412019
Rethinking the Discount Factor in Reinforcement Learning: A Decision Theoretic Approach · AAAI 2019
Machine learning › Reinforcement learning › temporal difference learning
eligibility traces
0.312018
Source Traces for Temporal Difference Learning · AAAI 2018
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
successor representation
0.312018
Source Traces for Temporal Difference Learning · AAAI 2018

Methods — techniques the papers use, named apart from their topics

multi-turn benchmark · 1.7interactive clinical vignettes · 1.7language model emulation · 1.5automatic safety evaluation · 1.5counterfactual data augmentation · 1.0experience replay · 0.8preference dataset · 0.8fine-tuning · 0.8large language model · 0.7automatic prompt optimization · 0.7
YearPublicationVenuePosition
2025 Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
abstract
Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks for large language models (LLMs) primarily assess knowledge recall through single-turn questions, where complete clinical information is provided upfront. To address this gap, we introduce VivaBench, a multi-turn benchmark that evaluates sequential clinical reasoning in LLM agents. Our dataset consists of 1762 physician-curated clinical vignettes structured as interactive scenarios that simulate a $ \textit{viva voce}$ (oral) examination in medical training, requiring agents to actively probe for relevant findings, select appropriate investigations, and synthesize information across multiple steps to reach a diagnosis. While current LLMs demonstrate competence in diagnosing conditions from well-described clinical presentations, their performance degrades significantly when required to navigate iterative diagnostic reasoning under uncertainty in our evaluation. Our analysis identified several failure modes that mirror common cognitive errors in clinical practice, including: (1) fixation on initial hypotheses, (2) inappropriate investigation ordering, (3) premature diagnostic closure, and (4) failing to screen for critical conditions. These patterns reveal fundamental limitations in how current LLMs reason and make decisions under uncertainty. Through VivaBench, we provide a standardized benchmark for evaluating conversational medical AI systems for real-world clinical decision support. Beyond medical applications, we contribute to the larger corpus of research on agentic AI by demonstrating how sequential reasoning trajectories can diverge in complex decision-making environments.
Christopher Chiu, Silviu Pitis, Mihaela van der Schaar
NeurIPS2
2024 Identifying the Risks of LM Agents with an LM-Emulated Sandbox
abstract
Recent advances in Language Model (LM) agents and tool use, exemplified by applications like ChatGPT Plugins, enable a rich set of capabilities but also amplify potential risks—such as leaking private data or causing financial losses. Identifying these risks is labor-intensive, necessitating implementing the tools, setting up the environment for each test scenario manually, and finding risky cases. As tools and agents become more complex, the high cost of testing these agents will make it increasingly difficult to find high-stakes, long-tail risks. To address these challenges, we introduce ToolEmu: a framework that uses an LM to emulate tool execution and enables scalable testing of LM agents against a diverse range of tools and scenarios. Alongside the emulator, we develop an LM-based automatic safety evaluator that examines agent failures and quantifies associated risks. We test both the tool emulator and evaluator through human evaluation and find that 68.8% of failures identified with ToolEmu would be valid real-world agent failures. Using our curated initial benchmark consisting of 36 high-stakes toolkits and 144 test cases, we provide a quantitative risk analysis of current LM agents and identify numerous failures with potentially severe outcomes. Notably, even the safest LM agent exhibits such failures 23.9% of the time according to our evaluator, underscoring the need to develop safer LM agents for real-world deployment.
Yangjun Ruan, Honghua Dong, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, Tatsunori B. Hashimoto
ICLR4
2024 Improving Context-Aware Preference Modeling for Language Models
abstract
While finetuning language models from pairwise preferences has proven remarkably effective, the underspecified nature of natural language presents critical challenges. Direct preference feedback is uninterpretable, difficult to provide where multidimensional criteria may apply, and often inconsistent, either because it is based on incomplete instructions or provided by diverse principals. To address these challenges, we consider the two-step preference modeling procedure that first resolves the under-specification by selecting a context, and then evaluates preference with respect to the chosen context. We decompose reward modeling error according to these two steps, which suggests that supervising context in addition to context-specific preference may be a viable approach to aligning models with diverse human preferences. For this to work, the ability of models to evaluate context-specific preference is critical. To this end, we contribute context-conditioned preference datasets and accompanying experiments that investigate the ability of language models to evaluate context-specific preference. Unlike past datasets, where context-specific preference is highly correlated with general preference, our "preference reversal" datasets disentangle context-specific and general preferences to isolate context-specific capabilities. We use our datasets to (1) show that existing preference models benefit from, but fail to fully consider, added context, (2) finetune a context-aware reward model with context-specific performance exceeding that of GPT-4 and Llama 3 70B, and (3) investigate the potential value of context-aware preference modeling.
Silviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro Sordoni
NeurIPS1
2023 Large Language Models are Human-Level Prompt Engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, Jimmy Ba
ICLR5
2023 Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards
abstract
As the capabilities of artificial agents improve, they are being increasingly deployed to service multiple diverse objectives and stakeholders. However, the composition of these objectives is often performed ad hoc, with no clear justification. This paper takes a normative approach to multi-objective agency: from a set of intuitively appealing axioms, it is shown that Markovian aggregation of Markovian reward functions is not possible when the time preference (discount factor) for each objective may vary. It follows that optimal multi-objective agents must admit rewards that are non-Markovian with respect to the individual objectives. To this end, a practical non-Markovian aggregation scheme is proposed, which overcomes the impossibility with only one additional parameter for each objective. This work offers new insights into sequential, multi-objective agency and intertemporal choice, and has practical implications for the design of AI systems deployed to serve multiple generations of principals with varying time preference.
Silviu Pitis
NeurIPS1
2022 MoCoDA: Model-based Counterfactual Data Augmentation
abstract
The number of states in a dynamic process is exponential in the number of objects, making reinforcement learning (RL) difficult in complex, multi-object domains. For agents to scale to the real world, they will need to react to and reason about unseen combinations of objects. We argue that the ability to recognize and use local factorization in transition dynamics is a key element in unlocking the power of multi-object reasoning. To this end, we show that (1) known local structure in the environment transitions is sufficient for an exponential reduction in the sample complexity of training a dynamics model, and (2) a locally factored dynamics model provably generalizes out-of-distribution to unseen states and actions. Knowing the local structure also allows us to predict which unseen states and actions this dynamics model will generalize to. We propose to leverage these observations in a novel Model-based Counterfactual Data Augmentation (MoCoDA) framework. MoCoDA applies a learned locally factored dynamics model to an augmented distribution of states and actions to generate counterfactual transitions for RL. MoCoDA works with a broader set of local structures than prior work and allows for direct control over the augmented training distribution. We show that MoCoDA enables RL agents to learn policies that generalize to unseen states and actions. We use MoCoDA to train an offline RL agent to solve an out-of-distribution robotics manipulation task on which standard offline RL algorithms fail.
Silviu Pitis, Elliot Creager, Ajay Mandlekar, Animesh Garg
NeurIPS1
2020 Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning
abstract
We explore fixed-horizon temporal difference (TD) methods, reinforcement learning algorithms for a new kind of value function that predicts the sum of rewards over a fixed number of future time steps. To learn the value function for horizon h, these algorithms bootstrap from the value function for horizon h−1, or some shorter horizon. Because no value function bootstraps from itself, fixed-horizon methods are immune to the stability problems that plague other off-policy TD methods using function approximation (also known as “the deadly triad”). Although fixed-horizon methods require the storage of additional value functions, this gives the agent additional predictive power, while the added complexity can be substantially reduced via parallel updates, shared weights, and n-step bootstrapping. We show how to use fixed-horizon value functions to solve reinforcement learning problems competitively with methods such as Q-learning that learn conventional value functions. We also prove convergence of fixed-horizon temporal difference methods with linear and general function approximation. Taken together, our results establish fixed-horizon TD methods as a viable new way of avoiding the stability problems of the deadly triad.
Kristopher De Asis, Alan Chan 0001, Silviu Pitis, Richard S. Sutton, Daniel Graves
AAAI3
2020 An Inductive Bias for Distances: Neural Nets that Respect the Triangle Inequality
Silviu Pitis, Harris Chan, Kiarash Jamali, Jimmy Ba
ICLR1
2020 Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning
abstract
What goals should a multi-goal reinforcement learning agent pursue during training in long-horizon tasks? When the desired (test time) goal distribution is too distant to offer a useful learning signal, we argue that the agent should not pursue unobtainable goals. Instead, it should set its own intrinsic goals that maximize the entropy of the historical achieved goal distribution. We propose to optimize this objective by having the agent pursue past achieved goals in sparsely explored areas of the goal space, which focuses exploration on the frontier of the achievable goal set. We show that our strategy achieves an order of magnitude better sample efficiency than the prior state of the art on long-horizon multi-goal tasks including maze navigation and block stacking.
Silviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie, Jimmy Ba
ICML1
2020 Counterfactual Data Augmentation using Locally Factored Dynamics
abstract
Many dynamic processes, including common scenarios in robotic control and reinforcement learning (RL), involve a set of interacting subprocesses. Though the subprocesses are not independent, their interactions are often sparse, and the dynamics at any given time step can often be decomposed into locally independent} causal mechanisms. Such local causal structures can be leveraged to improve the sample efficiency of sequence prediction and off-policy reinforcement learning. We formalize this by introducing local causal models (LCMs), which are induced from a global causal model by conditioning on a subset of the state space. We propose an approach to inferring these structures given an object-oriented state representation, as well as a novel algorithm for Counterfactual Data Augmentation (CoDA). CoDA uses local structures and an experience replay to generate counterfactual experiences that are causally valid in the global model. We find that CoDA significantly improves the performance of RL agents in locally factored tasks, including the batch-constrained and goal-conditioned settings. Code available at https://github.com/spitis/mrl.
Silviu Pitis, Elliot Creager, Animesh Garg
NeurIPS1
2019 Rethinking the Discount Factor in Reinforcement Learning: A Decision Theoretic Approach
abstract
Reinforcement learning (RL) agents have traditionally been tasked with maximizing the value function of a Markov decision process (MDP), either in continuous settings, with fixed discount factor γ
Silviu Pitis
AAAI1
2018 Source Traces for Temporal Difference Learning
abstract
This paper motivates and develops source traces for temporal difference (TD) learning in the tabular setting. Source traces are like eligibility traces, but model potential histories rather than immediate ones. This allows TD errors to be propagated to potential causal states and leads to faster generalization. Source traces can be thought of as the model-based, backward view of successor representations (SR), and share many of the same benefits. This view, however, suggests several new ideas. First, a TD(λ)-like source learning algorithm is proposed and its convergence is proven. Then, a novel algorithm for learning the source map (or SR matrix) is developed and shown to outperform the previous algorithm. Finally, various approaches to using the source/SR model are explored, and it is shown that source traces can be effectively combined with other model-based methods like Dyna and experience replay.
Silviu Pitis
AAAI1
2017 Methods for retrieving alternative contract language using a prototype
abstract
This paper addresses the problem of searching for alternative contract language that is similar to, yet different from, a given provision (the prototype). While this is a core task in transactional legal work, generic search solutions do not offer an effective solution. We draw upon modern information retrieval research to propose and validate novel methods for retrieving alternative language using a prototype. Our solution accepts an entire provision as a prototype and retrieves variants on the language from a database of precedent contracts. In designing this solution, we propose two ordered proximity measures and demonstrate their effectiveness relative to existing techniques. Further, we examine the challenge posed by varying definitions of redundant search results and propose to resolve it with a user-tunable, dynamic approach to result clustering.
Silviu Pitis
ICAIL1