EDBT 2026 Demo / reviewers in the wild / expert
Siddharth Suresh
dblp:262/0748
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-4501-8779ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AI-enhanced semantic feature norms for 786 concepts
Siddharth Suresh, Kushin Mukherjee, Tyler Giallanza, Mia Patil, Xizheng Yu, Jonathan D. Cohen 0003, Timothy T. Rogers |
CogSci | 1 |
| 2025 | Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds DecodingabstractYun-Shiuan Chuang, Sameer Narendran, Nikunj Harlalka, Alexander Cheung, Sizhe Gao, Siddharth Suresh, Junjie Hu, Timothy T. Rogers. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yun-Shiuan Chuang, Sameer Narendran, Nikunj Harlalka, Alexander Cheung, Sizhe Gao, Siddharth Suresh, Junjie Hu 0001, Timothy T. Rogers |
EMNLP | 6 |
| 2025 | Effects of Synchronous Movement on Human Trust in Robots*abstractRobot-human trust is an important concern as robots become integrated into human spaces. We tested a method grounded in psychological theory to increase human-robot trust—synchronous motion. Human participants completed a goal-oriented ball-moving task with a robotic arm to sound cues that were synchronous or asynchronous with the robot’s pacing. Participants were instructed to follow sound cues without information about synchrony. We found that participants in the synchrony condition trusted the robot to complete a new task that was comparable to the task they completed, significantly more than the asynchrony condition. However, this effect did not extend to harder tasks. The participants in the synchrony condition also believed that the robot had more influence on the outcomes of the new task compared to the asynchrony condition. On average, participants’ trust increased with the robotic arm after completing the task, regardless of condition. We report findings from a thematic analysis that demonstrate that participants in the synchrony condition found synchrony to be beneficial, while participants in the asynchrony condition found it cognitively taxing to be out-of-sync. Results from this work may be used to improve human-robot interactions in various contexts. Michelle Marji, Megh Vipul Doshi, Siddharth Suresh, Michael R. Zinn, Bilge Mutlu, Paula M. Niedenthal |
RO-MAN | 3 |
| 2024 | Simulating Opinion Dynamics with Networks of LLM-based Agents
Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Dhavan Shah, Junjie Hu 0001, Timothy T. Rogers |
CogSci | 4 |
| 2024 | The Wisdom of Partisan Crowds: Comparing Collective Intelligence in Humans and LLM-based Agents
Yun-Shiuan Chuang, Nikunj Harlalka, Siddharth Suresh, Agam Goyal, Robert Hawkins, Dhavan Shah, Junjie Hu 0001, Timothy T. Rogers |
CogSci | 3 |
| 2024 | Can deep convolutional networks explain the semantic structure that humans see in photographs?
Siddharth Suresh, Wei-Chun Huang, Kushin Mukherjee, Timothy T. Rogers |
CogSci | 1 |
| 2024 | Learning interactions to boost human creativity with bandits and GPT-4
Ara Vartanian, Xiaoxi Sun, Yun-Shiuan Chuang, Siddharth Suresh, Jerry Zhu, Timothy T. Rogers |
CogSci | 4 |
| 2024 | Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon CaptioningabstractWe present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human votes on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker's weekly cartoon caption contest over the past eight years. This unique dataset supports the development and evaluation of multimodal large language models and preference-based fine-tuning algorithms for humorous caption generation. We propose novel benchmarks for judging the quality of model-generated captions, utilizing both GPT4 and human judgments to establish ranking-based evaluation strategies. Our experimental results highlight the limitations of current fine-tuning methods, such as RLHF and DPO, when applied to creative tasks. Furthermore, we demonstrate that even state-of-the-art models like GPT4 and Claude currently underperform top human contestants in generating humorous captions. As we conclude this extensive data collection effort, we release the entire preference dataset to the research community, fostering further advancements in AI humor generation and evaluation. Jifan Zhang, Lalit K. Jain, Kuan Lok Zhou, Siddharth Suresh, Andrew J. Wagenmaker, Scott Sievert, Timothy T. Rogers, Kevin Jamieson 0001, Robert Mankoff, Robert D. Nowak |
NeurIPS | 6 |
| 2023 | Behavioral estimates of conceptual structure are robust across tasks in humans but not large language models
Siddharth Suresh, Kushin Mukherjee, Lisa Padua, Timothy T. Rogers |
CogSci | 1 |
| 2023 | Conceptual structure coheres in human cognition but not in large language modelsabstractNeural network models of language have long been used as a tool for developing hypotheses about conceptual representation in the mind and brain.For many years, such use involved extracting vector-space representations of words and using distances among these to predict or understand human behavior in various semantic tasks.Contemporary large language models (LLMs), however, make it possible to interrogate the latent structure of conceptual representations using experimental methods nearly identical to those commonly used with human participants.The current work utilizes three common techniques borrowed from cognitive psychology to estimate and compare the structure of concepts in humans and a suite of LLMs.In humans, we show that conceptual structure is robust to differences in culture, language, and method of estimation.Structures estimated from LLM behavior, while individually fairly consistent with those estimated from human behavior, vary much more depending upon the particular task used to generate responsesacross tasks, estimates of conceptual structure from the very same model cohere less with one another than do human structure estimates.These results highlight an important difference between contemporary LLMs and human cognition, with implications for understanding some fundamental limitations of contemporary machine language. Siddharth Suresh, Kushin Mukherjee, Xizheng Yu, Wei-Chun Huang, Lisa Padua, Timothy T. Rogers |
EMNLP | 1 |