VLDB 2026 Research / reviewers in the wild / expert
Paul McNamara
dblp:40/4366
· DBLP profile ↗
3ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 41% Multi-agent systems · 22% Planning, search and constraint satisfaction · 19% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
1.6 | 2 | 2025 | LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory · NeurIPS 2025 Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent reasoning
strategic reasoning |
0.9 | 1 | 2025 | LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty |
0.8 | 1 | 2024 | Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
thinking chain analysis · 0.9behavioral game-theoretic evaluation · 0.9multiple-choice-list experiment · 0.8behavioral economics theory · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM Strategic Reasoning: Agentic Study through Behavioral Game TheoryabstractWhat does it truly mean for a language model to “reason” strategically, and can scaling up alone guarantee intelligent, context-aware decisions? Strategic decision-making requires adaptive reasoning, where agents anticipate and respond to others’ actions under uncertainty. Yet, most evaluations of large language models (LLMs) for strategic decision-making often rely heavily on Nash Equilibrium (NE) benchmarks, overlook reasoning depth, and fail to reveal the mechanisms behind model behavior. To address this gap, we introduce a behavioral game-theoretic evaluation framework that disentangles intrinsic reasoning from contextual influence. Using this framework, we evaluate 22 state-of-the-art LLMs across diverse strategic scenarios. We find models like GPT-o3-mini, GPT-o1, and DeepSeek-R1 lead in reasoning depth. Through thinking chain analysis, we identify distinct reasoning styles—such as maximin or belief-based strategies—and show that longer reasoning chains do not consistently yield better decisions. Furthermore, embedding demographic personas reveals context-sensitive shifts: some models (e.g., GPT-4o, Claude-3-Opus) improve when assigned female identities, while others (e.g., Gemini 2.0) show diminished reasoning under minority sexuality personas. These findings underscore that technical sophistication alone is insufficient; alignment with ethical standards, human expectations, and situational nuance is essential for the responsible deployment of LLMs in interactive settings. Jingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara, Deming Chen |
NeurIPS | 4 |
| 2024 | Decision-Making Behavior Evaluation Framework for LLMs under Uncertain ContextabstractWhen making decisions under uncertainty, individuals often deviate from rational behavior, which can be evaluated across three dimensions: risk preference, probability weighting, and loss aversion. Given the widespread use of large language models (LLMs) in supporting decision-making processes, it is crucial to assess whether their behavior aligns with human norms and ethical expectations or exhibits potential biases. Although several empirical studies have investigated the rationality and social behavior performance of LLMs, their internal decision-making tendencies and capabilities remain inadequately understood. This paper proposes a framework, grounded in behavioral economics theories, to evaluate the decision-making behaviors of LLMs. With a multiple-choice-list experiment, we initially estimate the degree of risk preference, probability weighting, and loss aversion in a context-free setting for three commercial LLMs: ChatGPT-4.0-Turbo, Claude-3-Opus, and Gemini-1.0-pro. Our results reveal that LLMs generally exhibit patterns similar to humans, such as risk aversion and loss aversion, with a tendency to overweight small probabilities, but there are significant variations in the degree to which these behaviors are expressed across different LLMs. Further, we explore their behavior when embedded with socio-demographic features of human beings, uncovering significant disparities across various demographic characteristics. Jingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara, Deming Chen |
NeurIPS | 4 |
| 2013 | Weight optimisation for iterative distributed model predictive control applied to power networks
Paul McNamara, Rudy R. Negenborn, Bart De Schutter, Gordon Lightbody |
Eng. Appl. Artif. Intell. | 1 |