Elliot A. Ludvig

dblp:26/1091 · also Elliot Andrew Ludvig · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-0031-6713ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021
YearPublicationVenuePosition
2025 People Consistently Overweight Extreme Outcomes in Risky Choices, Even after Long Delays
Xiaomu Guo, Nick Simonsen, Christopher R. Madan, Marcia Spetch, Elliot A. Ludvig
CogSci5
2025 How descriptions moderate memory biases in experience-based risky choice
Zepeng Sun, Leonardo Weiss-Cohen, Elliot A. Ludvig, Emmanouil Konstantinidis
CogSci3
2024 Risky Decisions from Personal and Observed Experience
Alina Gutoreva, Elliot A. Ludvig
CogSci2
2024 Assimilating human feedback from autonomous vehicle interaction in reinforcement learning models
abstract
A significant challenge for real-world automated vehicles (AVs) is their interaction with human pedestrians. This paper develops a methodology to directly elicit the AV behaviour pedestrians find suitable by collecting quantitative data that can be used to measure and improve an algorithm's performance. Starting with a Deep Q Network (DQN) trained on a simple Pygame/Python-based pedestrian crossing environment, the reward structure was adapted to allow adjustment by human feedback. Feedback was collected by eliciting behavioural judgements collected from people in a controlled environment. The reward was shaped by the inter-action vector, decomposed into feature aspects for relevant behaviours, thereby facilitating both implicit preference selection and explicit task discovery in tandem. Using computational RL and behavioural-science techniques, we harness a formal iterative feedback loop where the rewards were repeatedly adapted based on human behavioural judgments. Experiments were conducted with 124 participants that showed strong initial improvement in the judgement of AV behaviours with the adaptive reward structure. The results indicate that the primary avenue for enhancing vehicle behaviour lies in the predictability of its movements when introduced. More broadly, recognising AV behaviours that receive favourable human judgments can pave the way for enhanced performance.
Richard Fox, Elliot A. Ludvig
Auton. Agents Multi Agent Syst.2
2023 Preferences for descriptiveness and co-explanation in complex explanations
Michael Hattersley, Reed Orchinik, Elliot A. Ludvig, Rahul Bhui
CogSci3
2021 Relationship between Delay Discounting and Risk Preference in Chimpanzees (Pan troglodytes) and Humans
Stefanie Keupp, Sebastian Grueneisen, Felix Warneken, Elliot A. Ludvig, Alicia P. Melis
CogSci4
2017 Information Seeking as Chasing Anticipated Prediction Errors
Jian-Qiao Zhu, Wendi Xiang, Elliot A. Ludvig
CogSci3
2014 Automated Story Selection for Color Commentary in Sports
abstract
Automated sports commentary is a form of automated narrative. Sports commentary exists to keep the viewer informed and entertained. One way to entertain the viewer is by telling brief stories relevant to the game in progress. We present a system called the sports commentary recommendation system (SCoReS) that can automatically suggest stories for commentators to tell during games. Through several user studies, we compared commentary using SCoReS to three other types of commentary and show that SCoReS adds significantly to the broadcast across several enjoyment metrics. We also collected interview data from professional sports commentators who positively evaluated a demonstration of the system. We conclude that SCoReS can be a useful broadcast tool, effective at selecting stories that add to the enjoyment and watchability of sports. SCoReS is a step toward automating sports commentary and, thus, automating narrative.
Greg Lee, Vadim Bulitko, Elliot A. Ludvig
IEEE Trans. Comput. Intell. AI Games3
2008 A computational model of hippocampal function in trace conditioning
abstract
We present a new reinforcement-learning model for the role of the hippocampus in classical conditioning, focusing on the differences between trace and delay conditioning. In the model, all stimuli are represented both as unindividuated wholes and as a series of temporal elements with varying delays. These two stimulus representations interact, producing different patterns of learning in trace and delay conditioning. The model proposes that hippocampal lesions eliminate long-latency temporal elements, but preserve short-latency temporal elements. For trace conditioning, with no contiguity between stimulus and reward, these long-latency temporal elements are vital to learning adaptively timed responses. For delay conditioning, in contrast, the continued presence of the stimulus supports conditioned responding, and the short-latency elements suppress responding early in the stimulus. In accord with the empirical data, simulated hippocampal damage impairs trace conditioning, but not delay conditioning, at medium-length intervals. With longer intervals, learning is impaired in both procedures, and, with shorter intervals, in neither. In addition, the model makes novel predictions about the response topography with extended stimuli or post-training lesions. These results demonstrate how temporal contiguity, as in delay conditioning, changes the timing problem faced by animals, rendering it both easier and less susceptible to disruption by hippocampal lesions.
Elliot A. Ludvig, Richard S. Sutton, Eric Verbeek 0002, E. James Kehoe
NIPS1
2008 Stimulus Representation and the Timing of Reward-Prediction Errors in Models of the Dopamine System
abstract
The phasic firing of dopamine neurons has been theorized to encode a reward-prediction error as formalized by the temporal-difference (TD) algorithm in reinforcement learning. Most TD models of dopamine have assumed a stimulus representation, known as the complete serial compound, in which each moment in a trial is distinctly represented. We introduce a more realistic temporal stimulus representation for the TD model. In our model, all external stimuli, including rewards, spawn a series of internal microstimuli, which grow weaker and more diffuse over time. These microstimuli are used by the TD learning algorithm to generate predictions of future reward. This new stimulus representation injects temporal generalization into the TD model and enhances correspondence between model and data in several experiments, including those when rewards are omitted or received early. This improved fit mostly derives from the absence of large negative errors in the new model, suggesting that dopamine alone can encode the full range of TD errors in these situations.
Elliot A. Ludvig, Richard S. Sutton, E. James Kehoe
Neural Comput.1