Shirley Wu

dblp:28/8766 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 39% Question answering and dialogue systems · 18% Reinforcement learning · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 72% Knowledge graphs · 28%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
1.422024
GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts · NeurIPS 2024
D4Explainer: In-distribution Explanations of Graph Neural Network via Discrete Denoising Diffusion · NeurIPS 2023
Machine learning › Trustworthy machine learning
out-of-distribution generalization
1.422024
GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts · NeurIPS 2024
Discover and Cure: Concept-aware Mitigation of Spurious Correlation · ICML 2023
Machine learning › Trustworthy machine learning
interpretability
1.322023
D4Explainer: In-distribution Explanations of Graph Neural Network via Discrete Denoising Diffusion · NeurIPS 2023
Discover and Cure: Concept-aware Mitigation of Spurious Correlation · ICML 2023
Natural language and speech › Question answering and dialogue systems
collaborative dialogue
0.912025
CollabLLM: From Passive Responders to Active Collaborators · ICML 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
human-AI collaboration
0.912025
CollabLLM: From Passive Responders to Active Collaborators · ICML 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning · NeurIPS 2025
Natural language and speech › Question answering and dialogue systems
multi-turn interaction
0.912025
CollabLLM: From Passive Responders to Active Collaborators · ICML 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
CollabLLM: From Passive Responders to Active Collaborators · ICML 2025
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.812024
GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts · NeurIPS 2024
Natural language and speech › Language models and text generation
LLM agents
0.812024
AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive Reasoning · NeurIPS 2024
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering › knowledge-grounded question answering
retrieval-augmented question answering
0.812024
AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive Reasoning · NeurIPS 2024
Information retrieval › evaluation › test collection
retrieval benchmark
0.812024
STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases · NeurIPS 2024
Information retrieval › document retrieval › structured document retrieval
semi-structured retrieval
0.812024
STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
concept-based explanation
0.712023
Discover and Cure: Concept-aware Mitigation of Spurious Correlation · ICML 2023
Machine learning › Trustworthy machine learning › interpretability
graph neural network explanation
0.712023
D4Explainer: In-distribution Explanations of Graph Neural Network via Discrete Denoising Diffusion · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation mitigation
0.712023
Discover and Cure: Concept-aware Mitigation of Spurious Correlation · ICML 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play
0.312025
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning · NeurIPS 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.212024
GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts · NeurIPS 2024
Information retrieval
e-commerce search
0.212024
STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases · NeurIPS 2024
Information retrieval › web search
scholarly search
0.212024
STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

reinforcement fine-tuning · 0.9multiturn-aware rewards · 0.9library augmentation · 0.9experience library · 0.9bootstrapped reasoning · 0.9representation alignment · 0.8prompt optimization · 0.8mixture of experts · 0.8large language model retrieval · 0.8contrastive reasoning · 0.8concept discovery · 0.7
YearPublicationVenuePosition
2025 CollabLLM: From Passive Responders to Active Collaborators
abstract
Large Language Models are typically trained with next-turn rewards, limiting their ability to optimize for long-term interaction. As a result, they often respond passively to ambiguous or open-ended user requests, failing to help users reach their ultimate intents and leading to inefficient conversations. To address these limitations, we introduce CollabLLM, a novel and general training framework that enhances multiturn human-LLM collaboration. Its key innovation is a collaborative simulation that estimates the long-term contribution of responses using Multiturn-aware Rewards. By reinforcement fine-tuning these rewards, CollabLLM goes beyond responding to user requests, and actively uncovers user intent and offers insightful suggestions—a key step towards more human-centered AI. We also devise a multiturn interaction benchmark with three challenging tasks such as document creation. CollabLLM significantly outperforms our baselines with averages of 18.5% higher task performance and 46.3% improved interactivity by LLM judges. Finally, we conduct a large user study with 201 judges, where CollabLLM increases user satisfaction by 17.6% and reduces user spent time by 10.4%.
Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng 0002, Gavin Li, Yao Dou, Weixin Cai, James Zou 0001, Jure Leskovec, Jianfeng Gao 0001
ICML1
2025 SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
abstract
Multi-agent AI systems powered by large language models (LLMs) are increasingly applied to solve complex tasks. However, these systems often rely on fragile, manually designed prompts and heuristics, making optimization difficult. A key challenge in optimizing multi-agent systems is acquiring suitable training data for specialized agents. We introduce SiriuS, a self-improving, reasoning-driven optimization framework for multi-agent systems. Central to our approach is the construction of an experience library: a repository of high-quality reasoning trajectories. The library is built by retaining reasoning steps that lead to successful outcomes, providing a robust training set for optimizing multi-agent system. Additionally, we introduce a library augmentation procedure that refines unsuccessful trajectories, further enriching the library. SiriuS boosts performance by 2.86% to 21.88% on reasoning and biomedical QA and enhances agent negotiation in competitive settings. Our results show that SiriuS enhances multi-agent performance while generating reusable data for self-correction and self-play enhancement in the future.
Wanjia Zhao, Mert Yüksekgönül, Shirley Wu, James Zou 0001
NeurIPS3
2024 GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts
abstract
Graph data are inherently complex and heterogeneous, leading to a high natural diversity of distributional shifts. However, it remains unclear how to build machine learning architectures that generalize to the complex distributional shifts naturally occurring in the real world. Here, we develop GraphMETRO, a Graph Neural Network architecture that models natural diversity and captures complex distributional shifts. GraphMETRO employs a Mixture-of-Experts (MoE) architecture with a gating model and multiple expert models, where each expert model targets a specific distributional shift to produce a referential representation w.r.t. a reference model, and the gating model identifies shift components. Additionally, we design a novel objective that aligns the representations from different expert models to ensure reliable optimization. GraphMETRO achieves state-of-the-art results on four datasets from the GOOD benchmark, which is comprised of complex and natural real-world distribution shifts, improving by 67% and 4.2% on the WebKB and Twitch datasets. Code and data are available at https://github.com/Wuyxin/GraphMETRO.
Shirley Wu, Kaidi Cao, Bruno Ribeiro 0001, James Zou 0001, Jure Leskovec
NeurIPS1
2024 AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive Reasoning
abstract
Large language model (LLM) agents have demonstrated impressive capabilities in utilizing external tools and knowledge to boost accuracy and reduce hallucinations. However, developing prompting techniques that enable LLM agents to effectively use these tools and knowledge remains a heuristic and labor-intensive task. Here, we introduce AvaTaR, a novel and automated framework that optimizes an LLM agent to effectively leverage provided tools, improving performance on a given task. During optimization, we design a comparator module to iteratively deliver insightful and comprehensive prompts to the LLM agent by contrastively reasoning between positive and negative examples sampled from training data. We demon- strate AvaTaR on four complex multimodal retrieval datasets featuring textual, visual, and relational information, and three general question-answering (QA) datasets. We find AvaTaR consistently outperforms state-of-the-art approaches across all seven tasks, exhibiting strong generalization ability when applied to novel cases and achieving an average relative improvement of 14% on the Hit@1 metric for the retrieval datasets and 13% for the QA datasets. Code and dataset are available at https://github.com/zou-group/avatar.
Shirley Wu, Qian Huang 0006, Michihiro Yasunaga, Kaidi Cao, Vassilis N. Ioannidis, Karthik Subbian, Jure Leskovec, James Zou 0001
NeurIPS1
2024 STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases
abstract
Answering real-world complex queries, such as complex product search, often requires accurate retrieval from semi-structured knowledge bases that involve blend of unstructured (e.g., textual descriptions of products) and structured (e.g., entity relations of products) information. However, many previous works studied textual and relational retrieval tasks as separate topics. To address the gap, we develop STARK, a large-scale Semi-structure retrieval benchmark on Textual and Relational Knowledge Bases. Our benchmark covers three domains: product search, academic paper search, and queries in precision medicine. We design a novel pipeline to synthesize realistic user queries that integrate diverse relational information and complex textual properties, together with their ground-truth answers (items). We conduct rigorous human evaluation to validate the quality of our synthesized queries. We further enhance the benchmark with high-quality human-generated queries to provide an authentic reference. STARK serves as a comprehensive testbed for evaluating the performance of retrieval systems driven by large language models (LLMs). Our experiments suggest that STARK presents significant challenges to the current retrieval and LLM systems, highlighting the need for more capable semi-structured retrieval systems.
Shirley Wu, Michihiro Yasunaga, Kaidi Cao, Qian Huang 0006, Vassilis N. Ioannidis, Karthik Subbian, James Zou 0001, Jure Leskovec
NeurIPS1
2023 Discover and Cure: Concept-aware Mitigation of Spurious Correlation
abstract
Deep neural networks often rely on spurious correlations to make predictions, which hinders generalization beyond training environments. For instance, models that associate cats with bed backgrounds can fail to predict the existence of cats in other environments without beds. Mitigating spurious correlations is crucial in building trustworthy models. However, the existing works lack transparency to offer insights into the mitigation process. In this work, we propose an interpretable framework, Discover and Cure (DISC), to tackle the issue. With human-interpretable concepts, DISC iteratively 1) discovers unstable concepts across different environments as spurious attributes, then 2) intervenes on the training data using the discovered concepts to reduce spurious correlation. Across systematic experiments, DISC provides superior generalization ability and interpretability than the existing approaches. Specifically, it outperforms the state-of-the-art methods on an object recognition task and a skin-lesion classification task by 7.5% and 9.6%, respectively. Additionally, we offer theoretical analysis and guarantees to understand the benefits of models trained by DISC. Code and data are available at https://github.com/Wuyxin/DISC.
Shirley Wu, Mert Yüksekgönül, Linjun Zhang, James Zou 0001
ICML1
2023 D4Explainer: In-distribution Explanations of Graph Neural Network via Discrete Denoising Diffusion
abstract
The widespread deployment of Graph Neural Networks (GNNs) sparks significant interest in their explainability, which plays a vital role in model auditing and ensuring trustworthy graph learning. The objective of GNN explainability is to discern the underlying graph structures that have the most significant impact on model predictions. Ensuring that explanations generated are reliable necessitates consideration of the in-distribution property, particularly due to the vulnerability of GNNs to out-of-distribution data. Unfortunately, prevailing explainability methods tend to constrain the generated explanations to the structure of the original graph, thereby downplaying the significance of the in-distribution property and resulting in explanations that lack reliability. To address these challenges, we propose D4Explainer, a novel approach that provides in-distribution GNN explanations for both counterfactual and model-level explanation scenarios. The proposed D4Explainer incorporates generative graph distribution learning into the optimization objective, which accomplishes two goals: 1) generate a collection of diverse counterfactual graphs that conform to the in-distribution property for a given instance, and 2) identify the most discriminative graph patterns that contribute to a specific class prediction, thus serving as model-level explanations. It is worth mentioning that D4Explainer is the first unified framework that combines both counterfactual and model-level explanations. Empirical evaluations conducted on synthetic and real-world datasets provide compelling evidence of the state-of-the-art performance achieved by D4Explainer in terms of explanation accuracy, faithfulness, diversity, and robustness.
Shirley Wu, Abhijit Gupta, Rex Ying
NeurIPS2
2009 Microblogging the ISMB: A New Approach to Conference Reporting
abstract
Microblogging platforms and other tools for videos, podcasts, and virtual environments provide an untapped potential for science conferences. Our experiment using FriendFeed to cover ISMB 2008 was educational and surprisingly successful. We found that it enhanced our note-taking skills, allowed us to compile notes from parallel sessions, attracted wider interest from non-attendees, and, in addition to the “live” aspect, generated a permanent archive of the meeting. ISMB/ECCB 2009 will be held in Stockholm. We look forward to the new developments in Web usage by scientists that are sure to emerge between now and then. We also anticipate new and exciting ways to report from Stockholm as it happens; perhaps the ISMB/ECCB 2009 Web site will look something like this: http://www.bork.embl.de/̃jensen/ismb2008/keynotes.php.html?
Neil F. W. Saunders, Pedro Beltrão, Lars Juhl Jensen, Daniel Jurczak, Roland Krause, Michael Kuhn 0004, Shirley Wu
PLoS Comput. Biol.7