EDBT 2026 Demo / reviewers in the wild / expert
Cedegao E. Zhang
dblp:245/7546 · also Cedegao Zhang
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2025
0009-0002-3096-9974ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generation and Evaluation in the Human Invention Process through the Lens of Game Design
Katie Collins, Graham Todd, Cedegao E. Zhang, Adrian Weller, Julian Togelius, Junyi Chu, Lionel Wong, Thomas L. Griffiths 0001, Josh Tenenbaum |
CogSci | 3 |
| 2025 | Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
Lionel Wong, Katie Collins, Lance Ying, Cedegao E. Zhang, Adrian Weller, Tobias Gerstenberg, Timothy J. O'Donnell, Alexander K. Lew, Jacob Andreas, Tyler Brooke-Wilson, Josh Tenenbaum |
CogSci | 4 |
| 2025 | Scaling up the think-aloud method
Daniel Wurgaft, Ben Prystawski, Kanishk Gandhi, Cedegao E. Zhang, Josh Tenenbaum, Noah D. Goodman |
CogSci | 4 |
| 2025 | On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad ConceptsabstractLanguage use is shaped by pragmatics-i.e., reasoning about communicative goals and norms in context.As language models (LMs) are increasingly used as conversational agents, it becomes ever more important to understand their pragmatic reasoning abilities.We propose an evaluation framework derived from Wavelength, a popular communication game where a speaker and a listener communicate about a broad range of concepts in a granular manner.We study a range of LMs on both language comprehension and language production using direct and Chain-of-Thought (CoT) prompting, and further explore a Rational Speech Act (RSA) approach to incorporating Bayesian pragmatic reasoning into LM inference.We find that state-of-the-art LMs, but not smaller ones, achieve strong performance on language comprehension, obtaining similar-to-human accuracy and exhibiting high correlations with human judgments even without CoT prompting or RSA.On language production, CoT can outperform direct prompting, and using RSA provides significant improvements over both approaches.Our study helps identify the strengths and limitations in LMs' pragmatic reasoning abilities and demonstrates the potential for improving them with RSA, opening up future avenues for understanding conceptual representation, language understanding, and social reasoning in LMs and humans. 1 Left Concept (0) Target Value Right Concept (100) Human-written Clues Chosen Clue Human Mean Deep thought 10 Shallow thought Evolution, Solving complex problems, Chess, Einstein, Meditation, Quantum mechanics Solv.complex prob. Linlu Qiu, Cedegao E. Zhang, Josh Tenenbaum, Roger Levy |
EMNLP | 2 |
| 2024 | People use fast, goal-directed simulation to reason about novel games
Cedegao E. Zhang, Katie Collins, Lionel Wong, Adrian Weller, Josh Tenenbaum |
CogSci | 1 |
| 2024 | Conditional and Modal Reasoning in Large Language ModelsabstractThe reasoning abilities of large language models (LLMs) are the topic of a growing body of research in AI and cognitive science.In this paper, we probe the extent to which twenty-nine LLMs are able to distinguish logically correct inferences from logically fallacious ones.We focus on inference patterns involving conditionals (e.g., 'If Ann has a queen, then Bob has a jack') and epistemic modals (e.g., 'Ann might have an ace', 'Bob must have a king').These inferences have been of special interest to logicians, philosophers, and linguists, since they play a central role in the fundamental human ability to reason about distal possibilities.Assessing LLMs on these inferences is thus highly relevant to the question of how much the reasoning abilities of LLMs match those of humans.All the LLMs we tested make some basic mistakes with conditionals or modals, though zero-shot chain-of-thought prompting helps them make fewer mistakes.Even the best performing LLMs make basic errors in modal reasoning, display logically inconsistent judgments across inference patterns involving epistemic modals and conditionals, and give answers about complex conditional inferences that do not match reported human judgments.These results highlight gaps in basic logical reasoning in today's LLMs. Wesley H. Holliday, Matthew Mandelkern, Cedegao E. Zhang |
EMNLP | 3 |
| 2023 | Towards a model of confidence judgements in concept learning
Tracey Mills, Cedegao E. Zhang, Tony Chen 0003, Josh Tenenbaum |
CogSci | 2 |
| 2023 | Grounded physical language understanding with probabilistic programs and simulated worlds
Cedegao E. Zhang, Lionel Wong, Gabriel Grand, Josh Tenenbaum |
CogSci | 1 |
| 2023 | LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversabstractLogical reasoning, i.e., deductively inferring the truth value of a conclusion from a set of premises, is an important task for artificial intelligence with wide potential impacts on science, mathematics, and society.While many prompting-based strategies have been proposed to enable Large Language Models (LLMs) to do such reasoning more effectively, they still appear unsatisfactory, often failing in subtle and unpredictable ways.In this work, we investigate the validity of instead reformulating such tasks as modular neurosymbolic programming, which we call LINC: Logical Inference via Neurosymbolic Computation.In LINC, the LLM acts as a semantic parser, translating premises and conclusions from natural language to expressions in first-order logic.These expressions are then offloaded to an external theorem prover, which symbolically performs deductive inference.Leveraging this approach, we observe significant performance gains on FOLIO and a balanced subset of ProofWriter for three different models in nearly all experimental conditions we evaluate.On ProofWriter, augmenting the comparatively small open-source StarCoder+ (15.5B parameters) with LINC even outperforms GPT-3.5 and GPT-4 with Chain-of-Thought (CoT) prompting by an absolute 38% and 10%, respectively.When used with GPT-4, LINC scores 26% higher than CoT on ProofWriter while performing comparatively on FOLIO.Further analysis reveals that although both methods on average succeed roughly equally often on this dataset, they exhibit distinct and complementary failure modes.We thus provide promising evidence for how logical reasoning over natural language can be tackled through jointly leveraging LLMs alongside symbolic provers.All corresponding code is publicly available. Theo X. Olausson, Alex Gu, Benjamin Lipkin, Cedegao E. Zhang, Armando Solar-Lezama, Josh Tenenbaum, Roger Levy |
EMNLP | 4 |
| 2021 | Does Amy Know Ben Knows You Know Your Cards? A Computational Model of Higher-Order Epistemic Reasoning
Cedegao E. Zhang, Ham Huang, Wesley H. Holliday |
CogSci | 1 |
| 2020 | A Model of Temporal Connective Acquisition
Mark Gorenstein, Cedegao E. Zhang, Steve Piantadosi |
CogSci | 2 |