Xinnuo Xu

dblp:211/7908 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 39% Knowledge representation and reasoning · 23% Probabilistic and Bayesian machine learning · 13%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
1.722025
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation · ICML 2025
Reasoning Elicitation in Language Models via Counterfactual Feedback · ICLR 2025
Natural language and speech › Language models and text generation
text generation
1.022021
AugNLG: Few-shot Natural Language Generation using Self-trained Data Augmentation · ACL/IJCNLP (1) 2021
AggGen: Ordering and Aggregating while Generating · ACL/IJCNLP (1) 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.912025
Compositional Causal Reasoning Evaluation in Language Models · ICML 2025
Natural language and speech › Language models and text generation › evaluation of language models › reasoning evaluation
causal reasoning evaluation
0.912025
Compositional Causal Reasoning Evaluation in Language Models · ICML 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Compositional Causal Reasoning Evaluation in Language Models · ICML 2025
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation
0.912025
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.812024
A Bayesian Approach to Data Point Selection · NeurIPS 2024
Machine learning › Efficient and distributed learning
data selection
0.812024
A Bayesian Approach to Data Point Selection · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference
0.812024
A Bayesian Approach to Data Point Selection · NeurIPS 2024
Machine learning › Deep learning architectures and training
data augmentation
0.512021
AugNLG: Few-shot Natural Language Generation using Self-trained Data Augmentation · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › text generation
data-to-text generation
0.512021
AggGen: Ordering and Aggregating while Generating · ACL/IJCNLP (1) 2021
Machine learning › Generative modeling › image generation › data-efficient image generation
few-shot image generation
0.512021
AugNLG: Few-shot Natural Language Generation using Self-trained Data Augmentation · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
text summarization
0.412020
Fact-based Content Weighting for Evaluating Abstractive Summarisation · ACL 2020
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder
0.312018
Better Conversations by Modeling, Filtering, and Optimizing for Coherence and Diversity · EMNLP 2018
Natural language and speech › Question answering and dialogue systems
open-domain dialogue
0.312018
Better Conversations by Modeling, Filtering, and Optimizing for Coherence and Diversity · EMNLP 2018
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.312025
Reasoning Elicitation in Language Models via Counterfactual Feedback · ICLR 2025
Machine learning › Optimization for machine learning
bilevel optimization
0.212024
A Bayesian Approach to Data Point Selection · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

symbolic problem generation · 0.9probability of necessity and sufficiency · 0.9ladder of causation · 0.9fine-tuning · 0.9counterfactual feedback · 0.9average treatment effect · 0.9stochastic gradient langevin dynamics · 0.8bayesian inference · 0.8MCMC · 0.8aggregation · 0.5
YearPublicationVenuePosition
2026 Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages
abstract
A key consideration when training an LLM is whether the target language is more or less resourced, for example English compared to Welsh, or Python compared to Excel. Typical training data for programming languages consists of real program demonstrations coupled with explanatory human-written comments. In this work we present a novel approach to the creation of such data for low resource programming languages, which lack naturally occurring data. Our process generates synthetic, textbook-quality demonstrations of how to use library functions, which we show makes for good model finetuning data. We demonstrate in an example domain of Excel Formulas. First, we collate language documentation, then we use this to augment a powerful teacher model which generates synthetic training data, and finally finetune student models on the demonstrations. Our technique improves student performance on 2 question-answering datasets: WikiTQ and TAT-QA. We also show advantages of finetuning over standard RAG approaches, which can offer only modest improvement due to the unfamiliarity of the target domain to student models.
Nick McKenna, Xinnuo Xu, Jack Williams 0001, Nicholas C. Wilson, Benjamin Van Durme, Christian Pölitz
LREC2
2025 Reasoning Elicitation in Language Models via Counterfactual Feedback
abstract
Despite the increasing effectiveness of language models, their reasoning capabilities remain underdeveloped. In particular, causal reasoning through counterfactual question answering is lacking. This work aims to bridge this gap. We first derive novel metrics that balance accuracy in factual and counterfactual questions, capturing a more complete view of the reasoning abilities of language models than traditional factual-only based metrics. Second, we propose several fine-tuning approaches that aim to elicit better reasoning mechanisms, in the sense of the proposed metrics. Finally, we evaluate the performance of the fine-tuned language models in a variety of realistic scenarios. In particular, we investigate to what extent our fine-tuning approaches systemically achieve better generalization with respect to the base models in several problems that require, among others, inductive and deductive reasoning capabilities.
Alihan Hüyük, Xinnuo Xu, Jacqueline R. M. A. Maasch, Aditya V. Nori, Javier González 0002
ICLR2
2025 Compositional Causal Reasoning Evaluation in Language Models
abstract
Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring the extent of these behaviors requires principled evaluation methods. We explore a unified perspective that considers both behaviors simultaneously, termed compositional causal reasoning (CCR): the ability to infer how causal measures compose and, equivalently, how causal quantities propagate through graphs. We instantiate a framework for the systematic evaluation of CCR for the average treatment effect and the probability of necessity and sufficiency. As proof of concept, we demonstrate CCR evaluation for language models in the Llama, Phi, and GPT families. On a math word problem, our framework revealed a range of taxonomically distinct error patterns. CCR errors increased with the complexity of causal paths for all models except o1.
Jacqueline R. M. A. Maasch, Alihan Hüyük, Xinnuo Xu, Aditya V. Nori, Javier González 0002
ICML3
2025 RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
abstract
Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true “reasoning” or from statistical recall of the training set. Inspired by the ladder of causation (Pearl, 2009) and its three levels (associations, interventions and counterfactuals), this paper introduces RE-IMAGINE: a framework to characterize a hierarchy of reasoning ability in LLMs, alongside an automated pipeline to generate problem variations at different levels of the hierarchy. By altering problems in an intermediate symbolic representation, RE-IMAGINE generates arbitrarily many problems that are not solvable using memorization alone. Moreover, the framework is general and can work across reasoning domains, including math, code, and logic. We demonstrate our framework on four widely-used benchmarks to evaluate several families of LLMs, and observe reductions in performance when the models are queried with problem variations. These assessments indicate a degree of reliance on statistical recall for past performance, and open the door to further research targeting skills across the reasoning hierarchy.
Xinnuo Xu, Rachel Lawrence, Kshitij Dubey, Atharva Pandey, Risa Ueno, Fabian Falck, Aditya V. Nori, Rahul Sharma 0001, Amit Sharma 0007, Javier González 0002
ICML1
2024 Graph Guided Question Answer Generation for Procedural Question-Answering
abstract
Hai Pham, Isma Hadji, Xinnuo Xu, Ziedune Degutyte, Jay Rainey, Evangelos Kazakos, Afsaneh Fazly, Georgios Tzimiropoulos, Brais Martinez. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hai X. Pham, Isma Hadji, Xinnuo Xu, Ziedune Degutyte, Jay Rainey, Evangelos Kazakos, Afsaneh Fazly, Georgios Tzimiropoulos, Brais Martínez
EACL (1)3
2024 A Bayesian Approach to Data Point Selection
abstract
Data point selection (DPS) is becoming a critical topic in deep learning due to the ease of acquiring uncurated training data compared to the difficulty of obtaining curated or processed data. Existing approaches to DPS are predominantly based on a bi-level optimisation (BLO) formulation, which is demanding in terms of memory and computation, and exhibits some theoretical defects regarding minibatches. Thus, we propose a novel Bayesian approach to DPS. We view the DPS problem as posterior inference in a novel Bayesian model where the posterior distributions of the instance-wise weights and the main neural network parameters are inferred under a reasonable prior and likelihood model. We employ stochastic gradient Langevin MCMC sampling to learn the main network and instance-wise weights jointly, ensuring convergence even with minibatches. Our update equation is comparable to the widely used SGD and much more efficient than existing BLO-based methods. Through controlled experiments in both the vision and language domains, we present the proof-of-concept. Additionally, we demonstrate that our method scales effectively to large language models and facilitates automated per-task optimization for instruction fine-tuning datasets.
Xinnuo Xu, Minyoung Kim 0001, Royson Lee, Brais Martínez, Timothy M. Hospedales
NeurIPS1
2021 AggGen: Ordering and Aggregating while Generating
abstract
Xinnuo Xu, Ondřej Dušek, Verena Rieser, Ioannis Konstas. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xinnuo Xu, Ondrej Dusek, Verena Rieser, Ioannis Konstas
ACL/IJCNLP (1)1
2021 AugNLG: Few-shot Natural Language Generation using Self-trained Data Augmentation
abstract
Xinnuo Xu, Guoyin Wang, Young-Bum Kim, Sungjin Lee. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xinnuo Xu, Guoyin Wang 0002, Young-Bum Kim
ACL/IJCNLP (1)1
2020 Fact-based Content Weighting for Evaluating Abstractive Summarisation
abstract
ive summarisation is notoriously hard to evaluate since standard word-overlap-based metrics are insufficient. We introduce a new evaluation metric which is based on fact-level content weighting, i.e. relating the facts of the document to the facts of the summary. We fol- low the assumption that a good summary will reflect all relevant facts, i.e. the ones present in the ground truth (human-generated refer- ence summary). We confirm this hypothe- sis by showing that our weightings are highly correlated to human perception and compare favourably to the recent manual highlight- based metric of Hardy et al. (2019).
Xinnuo Xu, Ondrej Dusek, Verena Rieser, Ioannis Konstas
ACL1
2020 Datasets and Benchmarks for Task-Oriented Log Dialogue Ranking Task
Xinnuo Xu, Yizhe Zhang 0002, Lars Liden
INTERSPEECH1
2019 Unsupervised Dialogue Spectrum Generation for Log Dialogue Ranking
abstract
Although the data-driven approaches of some recent bot building platforms make it possible for a wide range of users to easily create dialogue systems, those platforms don't offer tools for quickly identifying which log dialogues contain problems.This is important since corrections to log dialogues provide a means to improve performance after deployment.A log dialogue ranker, which ranks problematic dialogues higher, is an essential tool due to the sheer volume of log dialogues that could be generated.However, training a ranker typically requires labelling a substantial amount of data, which is not feasible for most users.In this paper, we present a novel unsupervised approach for dialogue ranking using GANs and release a corpus of labelled dialogues for evaluation and comparison with supervised methods.The evaluation result shows that our method compares favorably to supervised methods without any labelled data.
Xinnuo Xu, Yizhe Zhang 0002, Lars Liden
SIGdial1
2018 Better Conversations by Modeling, Filtering, and Optimizing for Coherence and Diversity
abstract
We present three enhancements to existing encoder-decoder models for open-domain conversational agents, aimed at effectively modeling coherence and promoting output diversity: (1) We introduce a measure of coherence as the GloVe embedding similarity between the dialogue context and the generated response, (2) we filter our training corpora based on the measure of coherence to obtain topically coherent and lexically diverse context-response pairs, (3) we then train a response generator using a conditional variational autoencoder model that incorporates the measure of coherence as a latent variable and uses a context gate to guarantee topical consistency with the context and promote lexical diversity.Experiments on the OpenSubtitles corpus show a substantial improvement over competitive neural models in terms of BLEU score as well as metrics of coherence and diversity.
Xinnuo Xu, Ondrej Dusek, Ioannis Konstas, Verena Rieser
EMNLP1