EDBT 2026 Demo / reviewers in the wild / expert
Xingchen Wan
dblp:255/7214
· DBLP profile ↗
20ranked-venue papers
8as first author
20since 2021 · last 2025
0000-0003-4142-0426ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language ModelsabstractRetrieval-augmented generation (RAG), while effective in integrating external knowledge to enhance large language models (LLMs), can be undermined by imperfect retrieval, which may introduce irrelevant, misleading, or even malicious information.Despite its importance, previous studies have rarely explored the behavior of RAG with errors from imperfect retrieval, and how potential conflicts arise between the LLMs' internal knowledge and external sources.We show that imperfect retrieval augmentation might be inevitable and quite harmful, through controlled analysis under realistic conditions.Knowledge conflicts between LLM-internal and external knowledge from retrieval is a bottleneck to overcome in the post-retrieval stage of RAG.To render LLMs resilient to imperfect retrieval, we propose ASTUTE RAG, a novel RAG approach that adaptively elicits essential information from LLMs' internal knowledge, iteratively consolidates internal and external knowledge with source-awareness, and finalizes the answer according to information reliability.Our experiments with Gemini and Claude demonstrate that ASTUTE RAG significantly outperforms previous robustness-enhanced RAG methods.Notably, ASTUTE RAG is the only approach that matches or exceeds the performance of LLMs without RAG under worst-case scenarios.ASTUTE RAG effectively resolves knowledge conflicts, improving the reliability and trustworthiness of RAG systems.45.4% LLM correct RAG correct LLM incorrect RAG incorrect Both sides are wrong.It is hard to improve, but combining internal and external knowledge may help.Previous work leverages RAG to address LLMs' knowledge gap.Zonia receives from Reuben a letter in the play.Zonia receives from Reuben a kiss in the play Xingchen Wan, Ruoxi Sun 0002, Jiefeng Chen 0001, Sercan Ö. Arik |
ACL (1) | 2 |
| 2025 | From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and GenerationabstractRecent advances in long-context large language models (LLMs) have led to the emerging paradigm of many-shot in-context learning (ICL), where it is observed that scaling many more demonstrating examples beyond the conventional few-shot setup in the context can lead to performance benefits. However, despite its promise, it is unclear what aspects dominate the benefits and whether simply scaling to more examples is the most effective way of improving many-shot ICL. In this work, we first provide an analysis on the factors driving many-shot ICL, and we find that 1) many-shot performance can still be attributed to often a few disproportionately influential examples and 2) identifying such influential examples ("optimize") and using them as demonstrations to regenerate new examples ("generate") can lead to further improvements. Inspired by the findings, we propose BRIDGE, an algorithm that alternates between the optimize step with Bayesian optimization to discover the influential sets of examples and the generate step to reuse this set to expand the reasoning paths of the examples back to the many-shot regime automatically. On Gemini, Claude, and Mistral LLMs of different sizes, we show BRIDGE led to significant improvements across a diverse set of tasks including symbolic reasoning, numerical reasoning and code generation. Xingchen Wan, Han Zhou 0010, Ruoxi Sun 0002, Sercan Ö. Arik |
ICLR | 1 |
| 2024 | Working Memory Capacity of ChatGPT: An Empirical StudyabstractWorking memory is a critical aspect of both human intelligence and artificial intelligence, serving as a workspace for the temporary storage and manipulation of information. In this paper, we systematically assess the working memory capacity of ChatGPT, a large language model developed by OpenAI, by examining its performance in verbal and spatial n-back tasks under various conditions. Our experiments reveal that ChatGPT has a working memory capacity limit strikingly similar to that of humans. Furthermore, we investigate the impact of different instruction strategies on ChatGPT's performance and observe that the fundamental patterns of a capacity limit persist. From our empirical findings, we propose that n-back tasks may serve as tools for benchmarking the working memory capacity of large language models and hold potential for informing future efforts aimed at enhancing AI working memory. Dongyu Gong, Xingchen Wan, Dingmin Wang |
AAAI | 2 |
| 2024 | Adaptive Batch Sizes for Active Learning: A Probabilistic Numerics ApproachabstractActive learning parallelization is widely used, but typically relies on fixing the batch size throughout experimentation. This fixed approach is inefficient because of a dynamic trade-off between cost and speed—larger batches are more costly, smaller batches lead to slower wall-clock run-times—and the trade-off may change over the run (larger batches are often preferable earlier). To address this trade-off, we propose a novel Probabilistic Numerics framework that adaptively changes batch sizes. By framing batch selection as a quadrature task, our integration-error-aware algorithm facilitates the automatic tuning of batch sizes to meet predefined quadrature precision objectives, akin to how typical optimizers terminate based on convergence thresholds. This approach obviates the necessity for exhaustive searches across all potential batch sizes. We also extend this to scenarios with constrained active learning and constrained optimization, interpreting constraint violations as reductions in the precision requirement, to subsequently adapt batch construction. Through extensive experiments, we demonstrate that our approach significantly enhances learning efficiency and flexibility in diverse Bayesian batch active learning and Bayesian optimization applications. Masaki Adachi, Satoshi Hayakawa, Martin Jørgensen, Xingchen Wan, Harald Oberhauser, Michael A. Osborne |
AISTATS | 4 |
| 2024 | Fairer Preferences Elicit Improved Human-Aligned Large Language Model JudgmentsabstractLarge language models (LLMs) have shown promising abilities as cost-effective and reference-free evaluators for assessing language generation quality.In particular, pairwise LLM evaluators, which compare two generated texts and determine the preferred one, have been employed in a wide range of applications.However, LLMs exhibit preference biases and worrying sensitivity to prompt designs.In this work, we first reveal that the predictive preference of LLMs can be highly brittle and skewed, even with semantically equivalent instructions.We find that fairer predictive preferences from LLMs consistently lead to judgments that are better aligned with humans.Motivated by this phenomenon, we propose an automatic Zero-shot Evaluation-oriented Prompt Optimization framework, ZEPO, which aims to produce fairer preference decisions and improve the alignment of LLM evaluators with human judgments.To this end, we propose a zeroshot learning objective based on the preference decision fairness.ZEPO demonstrates substantial performance improvements over stateof-the-art LLM evaluators, without requiring labeled data, on representative meta-evaluation benchmarks.Our findings underscore the critical correlation between preference fairness and human alignment, positioning ZEPO as an efficient prompt optimizer for bridging the gap between LLM evaluators and human judgments.* Now at Google.Code is available at https://github. com/cambridgeltl/zepo.Generate new prompts Zero-shot Fairness ZEPO Biased Preference Fairer Preference Optimized Prompt Initial Prompt LLM Optimizer LLM Evaluator Which summary candidate has better coherence?If the candidate A is better, please return 'A'.If the candidate B is better, please return 'B'.Which one exhibits better coherence?Return 'A' for the rst summary or 'B' for the second.Only provide the letter of your choice. Han Zhou 0010, Xingchen Wan, Yinhong Liu, Nigel Collier, Ivan Vulic, Anna Korhonen |
EMNLP | 2 |
| 2024 | Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt EngineeringabstractPrompting and in-context learning (ICL) have become efficient learning paradigms for large language models (LLMs). However, LLMs suffer from prompt brittleness and various bias factors in the prompt, including but not limited to the formatting, the choice verbalizers, and the ICL examples. To address this problem that results in unexpected performance degradation, calibration methods have been developed to mitigate the effects of these biases while recovering LLM performance. In this work, we first conduct a systematic analysis of the existing calibration methods, where we both provide a unified view and reveal the failure cases. Inspired by these analyses, we propose Batch Calibration (BC), a simple yet intuitive method that controls the contextual bias from the batched input, unifies various prior approaches and effectively addresses the aforementioned issues. BC is zero-shot, inference-only, and incurs negligible additional costs. In the few-shot setup, we further extend BC to allow it to learn the contextual bias from labeled data. We validate the effectiveness of BC with PaLM 2-(S, M, L) and CLIP models and demonstrate state-of-the-art performance over previous calibration baselines across more than 10 natural language understanding and image classification tasks. Han Zhou 0010, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine A. Heller, Subhrajit Roy |
ICLR | 2 |
| 2024 | UQE: A Query Engine for Unstructured DatabasesabstractAnalytics on structured data is a mature field with many successful methods.
However, most real world data exists in unstructured form, such as images and conversations.
We investigate the potential of Large Language Models (LLMs) to enable unstructured data analytics.
In particular, we propose a new Universal Query Engine (UQE) that directly interrogates and draws insights from unstructured data collections.
This engine accepts queries in a Universal Query Language (UQL), a dialect of SQL that provides full natural language flexibility in specifying conditions and operators.
The new engine leverages the ability of LLMs to conduct analysis of unstructured data, while also allowing us to exploit advances in sampling and optimization techniques to achieve efficient and accurate query execution.
In addition, we borrow techniques from classical compiler theory to better orchestrate the workflow between sampling methods and foundation model calls.
We demonstrate the efficiency of UQE on data analytics across different modalities, including images, dialogs and reviews, across a range of useful query types, including conditional aggregation, semantic retrieval and abstraction aggregation. Hanjun Dai, Bethany Wang, Xingchen Wan, Bo Dai 0001, Sherry Yang 0001, Azade Nova, Phitchaya Mangpo Phothilimthana, Charles Sutton, Dale Schuurmans |
NeurIPS | 3 |
| 2024 | Bayesian Optimization of Functions over Node Subsets in GraphsabstractWe address the problem of optimizing over functions defined on node subsets in a graph. The optimization of such functions is often a non-trivial task given their combinatorial, black-box and expensive-to-evaluate nature. Although various algorithms have been introduced in the literature, most are either task-specific or computationally inefficient and only utilize information about the graph structure without considering the characteristics of the function. To address these limitations, we utilize Bayesian Optimization (BO), a sample-efficient black-box solver, and propose a novel framework for combinatorial optimization on graphs. More specifically, we map each $k$-node subset in the original graph to a node in a new combinatorial graph and adopt a local modeling approach to efficiently traverse the latter graph by progressively sampling its subgraphs using a recursive algorithm. Extensive experiments under both synthetic and real-world setups demonstrate the effectiveness of the proposed BO framework on various types of graphs and optimization tasks, where its behavior is analyzed in detail with ablation studies. Huidong Liang, Xingchen Wan, Xiaowen Dong 0001 |
NeurIPS | 2 |
| 2024 | Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt OptimizationabstractLarge language models have demonstrated remarkable capabilities but their performance is heavily reliant on effective prompt engineering. Automatic prompt optimization (APO) methods are designed to automate this and can be broadly categorized into those targeting instructions (instruction optimization, IO) vs. those targeting exemplars (exemplar optimization, EO). Despite their shared objective, these have evolved rather independently, with IO receiving more research attention recently. This paper seeks to bridge this gap by comprehensively comparing the performance of representative IO and EO techniques both isolation and combination on a diverse set of challenging tasks. Our findings reveal that intelligently reusing model-generated input-output pairs obtained from evaluating prompts on the validation set as exemplars, consistently improves performance on top of IO methods but is currently under-investigated. We also find that despite the recent focus on IO, how we select exemplars can outweigh how we optimize instructions, with EO strategies as simple as random search outperforming state-of-the-art IO methods with seed instructions without any optimization. Moreover, we observe a synergy between EO and IO, with optimal combinations surpassing the individual contributions. We conclude that studying exemplar optimization both as a standalone method and its optimal combination with instruction optimization remain a crucial aspect of APO and deserve greater consideration in future research, even in the era of highly capable instruction-following models. Xingchen Wan, Ruoxi Sun 0002, Hootan Nakhost, Sercan Ö. Arik |
NeurIPS | 1 |
| 2024 | Iterate Averaging in the Quest for Best Test ErrorabstractWe analyse and explain the increased generalisation performance of iterate averaging using a Gaussian process perturbation model between the true and batch risk surface on the high dimensional quadratic. We derive three phenomena from our theoretical results: (1) The importance of combining iterate averaging (IA) with large learning rates and regularisation for improved generalisation. (2) Justification for less frequent averaging. (3) That we expect adaptive gradient methods to work equally well, or better, with iterate averaging than their non-adaptive counterparts. Inspired by these results, together with empirical investigations of the importance of appropriate regularisation for the solution diversity of the iterates, we propose two adaptive algorithms with iterate averaging. These give significantly better results compared to stochastic gradient descent (SGD), require less tuning and do not require early stopping or validation set monitoring. We showcase the efficacy of our approach on the CIFAR-10/100, ImageNet and Penn Treebank datasets on a variety of modern and classical network architectures. Diego Granziol, Nicholas P. Baskerville, Xingchen Wan, Samuel Albanie, Stephen J. Roberts |
J. Mach. Learn. Res. | 3 |
| 2024 | AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-TuningabstractAbstract Large pretrained language models are widely used in downstream NLP tasks via task- specific fine-tuning, but such procedures can be costly. Recently, Parameter-Efficient Fine-Tuning (PEFT) methods have achieved strong task performance while updating much fewer parameters than full model fine-tuning (FFT). However, it is non-trivial to make informed design choices on the PEFT configurations, such as their architecture, the number of tunable parameters, and even the layers in which the PEFT modules are inserted. Consequently, it is highly likely that the current, manually designed configurations are suboptimal in terms of their performance-efficiency trade-off. Inspired by advances in neural architecture search, we propose AutoPEFT for automatic PEFT configuration selection: We first design an expressive configuration search space with multiple representative PEFT modules as building blocks. Using multi-objective Bayesian optimization in a low-cost setup, we then discover a Pareto-optimal set of configurations with strong performance-cost trade-offs across different numbers of parameters that are also highly transferable across different tasks. Empirically, on GLUE and SuperGLUE tasks, we show that AutoPEFT-discovered configurations significantly outperform existing PEFT methods and are on par or better than FFT without incurring substantial training efficiency costs. Han Zhou 0010, Xingchen Wan, Ivan Vulic, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Universal Self-Adaptive PromptingabstractA hallmark of modern large language models (LLMs) is their impressive general zero-shot and few-shot abilities, often elicited through in-context learning (ICL) via prompting.However, while highly coveted and being the most general, zero-shot performances in LLMs are still typically weaker due to the lack of guidance and the difficulty of applying existing automatic prompt design methods in general tasks when ground-truth labels are unavailable.In this study, we address this by presenting Universal Self-Adaptive Prompting (USP), an automatic prompt design approach specifically tailored for zero-shot learning (while compatible with few-shot).Requiring only a small amount of unlabeled data and an inferenceonly LLM, USP is highly versatile: to achieve universal prompting, USP categorizes a possible NLP task into one of the three possible task types and then uses a corresponding selector to select the most suitable queries and zero-shot model-generated responses as pseudo-demonstrations, thereby generalizing ICL to the zero-shot setup in a fully automated way.We evaluate USP with PaLM and PaLM 2 models and demonstrate performances that are considerably stronger than standard zero-shot baselines and often comparable to or even superior to few-shot baselines across more than 40 natural language understanding, natural language generation, and reasoning tasks. Xingchen Wan, Ruoxi Sun 0002, Hootan Nakhost, Hanjun Dai, Julian Martin Eisenschlos, Sercan Ö. Arik, Tomas Pfister |
EMNLP | 1 |
| 2023 | Bayesian Optimisation of Functions on GraphsabstractThe increasing availability of graph-structured data motivates the task of optimising over functions defined on the node set of graphs. Traditional graph search algorithms can be applied in this case, but they may be sample-inefficient and do not make use of information about the function values; on the other hand, Bayesian optimisation is a class of promising black-box solvers with superior sample efficiency, but it has scarcely been applied to such novel setups. To fill this gap, we propose a novel Bayesian optimisation framework that optimises over functions defined on generic, large-scale and potentially unknown graphs. Through the learning of suitable kernels on graphs, our framework has the advantage of adapting to the behaviour of the target function. The local modelling approach further guarantees the efficiency of our method. Extensive experiments on both synthetic and real-world graphs demonstrate the effectiveness of the proposed optimisation framework. Xingchen Wan, Pierre Osselin, Henry Kenlay, Binxin Ru, Michael A. Osborne, Xiaowen Dong 0001 |
NeurIPS | 1 |
| 2022 | BOiLS: Bayesian Optimisation for Logic SynthesisabstractOptimising the quality-of-results (QoR) of circuits during logic synthesis is a formidable challenge necessitating the exploration of exponentially sized search spaces. While expert-designed operations aid in uncovering effective sequences, the increase in complexity of logic circuits favours automated procedures. To enable efficient and scalable solvers, we propose BOiLS, the first algorithm adapting Bayesian optimisation to navigate the space of synthesis operations. BOiLS requires no human intervention and trades-off exploration versus exploitation through novel Gaussian process kernels and trust-region constrained acquisitions. In a set of experiments on EPFL benchmarks, we demonstrate BOiLS's superior performance compared to state-of-the-art in terms of both sample efficiency and QoR values. Antoine Grosnit, Cédric Malherbe, Rasul Tutunov, Xingchen Wan, Jun Wang 0012, Haitham Bou-Ammar |
DATE | 4 |
| 2022 | On Redundancy and Diversity in Cell-based Neural Architecture Search
Xingchen Wan, Binxin Ru, Pedro M. Esperança, Zhenguo Li |
ICLR | 1 |
| 2022 | Bayesian Optimization over Discrete and Mixed Spaces via Probabilistic ReparameterizationabstractOptimizing expensive-to-evaluate black-box functions of discrete (and potentially continuous) design parameters is a ubiquitous problem in scientific and engineering applications. Bayesian optimization (BO) is a popular, sample-efficient method that leverages a probabilistic surrogate model and an acquisition function (AF) to select promising designs to evaluate. However, maximizing the AF over mixed or high-cardinality discrete search spaces is challenging standard gradient-based methods cannot be used directly or evaluating the AF at every point in the search space would be computationally prohibitive. To address this issue, we propose using probabilistic reparameterization (PR). Instead of directly optimizing the AF over the search space containing discrete parameters, we instead maximize the expectation of the AF over a probability distribution defined by continuous parameters. We prove that under suitable reparameterizations, the BO policy that maximizes the probabilistic objective is the same as that which maximizes the AF, and therefore, PR enjoys the same regret bounds as the original BO policy using the underlying AF. Moreover, our approach provably converges to a stationary point of the probabilistic objective under gradient ascent using scalable, unbiased estimators of both the probabilistic objective and its gradient. Therefore, as the number of starting points and gradient steps increase, our approach will recover of a maximizer of the AF (an often-neglected requisite for commonly used BO regret bounds). We validate our approach empirically and demonstrate state-of-the-art optimization performance on a wide range of real-world applications. PR is complementary to (and benefits) recent work and naturally generalizes to settings with multiple objectives and black-box constraints. Samuel Daulton, Xingchen Wan, David Eriksson, Maximilian Balandat, Michael A. Osborne, Eytan Bakshy |
NeurIPS | 2 |
| 2022 | Approximate Neural Architecture Search via Operation Distribution LearningabstractThe standard paradigm in Neural Architecture Search (NAS) is to search for a fully deterministic architecture with specific operations and connections. In this work, we instead propose to search for the optimal operation distribution, thus providing a stochastic and approximate solution, which can be used to sample architectures of arbitrary length. We propose and show, that given an architectural cell, its performance largely depends on the ratio of used operations, rather than any specific connection pattern in typical search spaces; that is, small changes in the ordering of the operations are often irrelevant. This intuition is orthogonal to any specific search strategy and can be applied to a diverse set of NAS algorithms. Through extensive validation on 4 data-sets and 4 NAS techniques (Bayesian optimisation, differentiable search, local search and random search), we show that the operation distribution (1) holds enough discriminating power to reliably identify a solution and (2) is significantly easier to optimise than traditional encodings, leading to large speed-ups at little to no cost in performance. Indeed, this simple intuition significantly reduces the cost of current approaches and potentially enable NAS to be used in a broader range of applications. Xingchen Wan, Binxin Ru, Pedro M. Esperança, Fabio Maria Carlucci |
WACV | 1 |
| 2021 | Interpretable Neural Architecture Search via Bayesian Optimisation with Weisfeiler-Lehman Kernels
Bin Xin Ru, Xingchen Wan, Xiaowen Dong 0001, Michael A. Osborne |
ICLR | 2 |
| 2021 | Think Global and Act Local: Bayesian Optimisation over High-Dimensional Categorical and Mixed Search SpacesabstractHigh-dimensional black-box optimisation remains an important yet notoriously challenging problem. Despite the success of Bayesian optimisation methods on continuous domains, domains that are categorical, or that mix continuous and categorical variables, remain challenging. We propose a novel solution—we combine local optimisation with a tailored kernel design, effectively handling high-dimensional categorical and mixed search spaces, whilst retaining sample efficiency. We further derive convergence guarantee for the proposed approach. Finally, we demonstrate empirically that our method outperforms the current baselines on a variety of synthetic and real-world tasks in terms of performance, computational costs, or both. Xingchen Wan, Huong Ha 0001, Bin Xin Ru, Cong Lu, Michael A. Osborne |
ICML | 1 |
| 2021 | Adversarial Attacks on Graph Classifiers via Bayesian OptimisationabstractGraph neural networks, a popular class of models effective in a wide range of graph-based learning tasks, have been shown to be vulnerable to adversarial attacks. While the majority of the literature focuses on such vulnerability in node-level classification tasks, little effort has been dedicated to analysing adversarial attacks on graph-level classification, an important problem with numerous real-life applications such as biochemistry and social network analysis. The few existing methods often require unrealistic setups, such as access to internal information of the victim models, or an impractically-large number of queries. We present a novel Bayesian optimisation-based attack method for graph classification models. Our method is black-box, query-efficient and parsimonious with respect to the perturbation applied. We empirically validate the effectiveness and flexibility of the proposed method on a wide range of graph classification tasks involving varying graph properties, constraints and modes of attack. Finally, we analyse common interpretable patterns behind the adversarial samples produced, which may shed further light on the adversarial robustness of graph classification models. Xingchen Wan, Henry Kenlay, Robin Ru, Arno Blaas, Michael A. Osborne, Xiaowen Dong 0001 |
NeurIPS | 1 |