Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jon Ander Campos

dblp:262/3907 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 65% Reinforcement learning · 13% Question answering and dialogue systems · 12%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › prompt tuning
adversarial prompt tuning
0.912025
Reverse Engineering Human Preferences with Reinforcement Learning · NeurIPS 2025
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting
0.912025
No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025
Machine learning › Reinforcement learning
learning from failure
0.912025
No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.912025
Reverse Engineering Human Preferences with Reinforcement Learning · NeurIPS 2025
Natural language and speech › Language models and text generation
prompting
0.912025
No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering
0.412020
DoQA - Accessing Domain-Specific FAQs via Conversational QA · ACL 2020
Natural language and speech › Question answering and dialogue systems
dialogue evaluation
0.412020
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems · EMNLP (1) 2020
Natural language and speech › Language models and text generation
mathematical reasoning
0.312025
No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025
Natural language and speech › Question answering and dialogue systems
conversational agents
0.112020
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems · EMNLP (1) 2020
Information retrieval › question answering
FAQ retrieval
0.112020
DoQA - Accessing Domain-Specific FAQs via Conversational QA · ACL 2020
Information retrieval › question answering
question answering retrieval
0.112020
DoQA - Accessing Domain-Specific FAQs via Conversational QA · ACL 2020

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.9prompting · 0.9preamble generation · 0.9judge language model · 0.9in-context learning · 0.9wizard of oz crowdsourcing · 0.9transfer learning · 0.9fine-tuning · 0.9human evaluation · 0.4bot detection · 0.4
YearPublicationVenuePosition
2025 From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions
abstract
Nathanaël Carraz Rakotonirina, Mohammed Hamdy, Jon Ander Campos, Lucas Weber, Alberto Testoni, Marzieh Fadaee, Sandro Pezzelle, Marco Del Tredici. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nathanaël Carraz Rakotonirina, Mohammed Hamdy, Jon Ander Campos, Lucas Weber, Alberto Testoni, Marzieh Fadaee, Sandro Pezzelle, Marco Del Tredici
ACL (1)3
2025 No Need for Explanations: LLMs can implicitly learn from mistakes in-context
abstract
Showing incorrect answers to Large Language Models (LLMs) is a popular strategy to improve their performance in reasoning-intensive tasks.It is widely assumed that, in order to be helpful, the incorrect answers must be accompanied by comprehensive rationales, explicitly detailing where the mistakes are and how to correct them.However, in this work we present a counterintuitive finding: we observe that LLMs perform better in math reasoning tasks when these rationales are eliminated from the context and models are left to infer on their own what makes an incorrect answer flawed.This approach also substantially outperforms chainof-thought prompting in our evaluations.These results are consistent across LLMs of different sizes and varying reasoning abilities.To gain an understanding of why LLMs learn from mistakes more effectively without explicit corrective rationales, we perform a thorough analysis, investigating changes in context length and answer diversity between different prompting strategies, and their effect on performance.We also examine evidence of overfitting to the in-context rationales when these are provided, and study the extent to which LLMs are able to autonomously infer high-quality corrective rationales given only incorrect answers as input.We find evidence that, while incorrect answers are more beneficial for LLM learning than additional diverse correct answers, explicit corrective rationales over-constrain the model, thus limiting those benefits.
Lisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan, Marek Rei, Max Bartolo
EMNLP3
2025 Reverse Engineering Human Preferences with Reinforcement Learning
abstract
The capabilities of Large Language Models (LLMs) are routinely evaluated by other LLMs trained to predict human preferences. This framework—known as *LLM-as-a-judge*—is highly scalable and relatively low cost. However, it is also vulnerable to malicious exploitation, as LLM responses can be tuned to overfit the preferences of the judge. Previous work shows that the answers generated by a candidate-LLM can be edited *post hoc* to maximise the score assigned to them by a judge-LLM. In this study, we adopt a different approach and use the signal provided by judge-LLMs as a reward to adversarially tune models that generate text preambles designed to boost downstream performance. We find that frozen LLMs pipelined with these models attain higher LLM-evaluation scores than existing frameworks. Crucially, unlike other frameworks which intervene directly on the model's response, our method is virtually undetectable. We also demonstrate that the effectiveness of the tuned preamble generator transfers when the candidate-LLM and the judge-LLM are replaced with models that are not used during training. These findings raise important questions about the design of more reliable LLM-as-a-judge evaluation settings. They also demonstrate that human preferences can be reverse engineered effectively, by pipelining LLMs to optimise upstream preambles via reinforcement learning—an approach that could find future applications in diverse tasks and domains beyond adversarial attacks.
Lisa Alazraki, Yi Chern Tan, Jon Ander Campos, Maximilian Mozes, Marek Rei, Max Bartolo
NeurIPS3
2020 DoQA - Accessing Domain-Specific FAQs via Conversational QA
abstract
The goal of this work is to build conversational Question Answering (QA) interfaces for the large body of domain-specific information available in FAQ sites.We present DoQA, a dataset with 2,437 dialogues and 10,917 QA pairs.The dialogues are collected from three Stack Exchange sites using the Wizard of Oz method with crowdsourcing.Compared to previous work, DoQA comprises well-defined information needs, leading to more coherent and natural conversations with less factoid questions and is multi-domain.In addition, we introduce a more realistic information retrieval (IR) scenario where the system needs to find the answer in any of the FAQ documents.The results of an existing, strong, system show that, thanks to transfer learning from a Wikipedia QA dataset and fine tuning on a single FAQ domain, it is possible to build high quality conversational QA systems for FAQs without indomain training data.The good results carry over into the more challenging IR scenario.In both cases, there is still ample room for improvement, as indicated by the higher human upperbound.
Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, Eneko Agirre
ACL1
2020 Improving Conversational Question Answering Systems after Deployment using Feedback-Weighted Learning
abstract
The interaction of conversational systems with users poses an exciting opportunity for improving them after deployment, but little evidence has been provided of its feasibility.In most applications, users are not able to provide the correct answer to the system, but they are able to provide binary (correct, incorrect) feedback.In this paper we propose feedback-weighted learning based on importance sampling to improve upon an initial supervised system using binary user feedback.We perform simulated experiments on document classification (for development) and Conversational Question Answering datasets like QuAC and DoQA, where binary user feedback is derived from gold annotations.The results show that our method is able to improve over the initial supervised system, getting close to a fully-supervised system that has access to the same labeled examples in in-domain experiments (QuAC), and even matching in out-of-domain experiments (DoQA).Our work opens the prospect to exploit interactions with real users and improve conversational systems after deployment.
Jon Ander Campos, Kyunghyun Cho, Arantxa Otegi, Aitor Soroa, Eneko Agirre, Gorka Azkune
COLING1
2020 Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems
abstract
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Álvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak
EMNLP (1)4
2020 Give your Text Representation Models some Love: the Case for Basque
abstract
Word embeddings and pre-trained language models allow to build rich representations of text and have enabled improvements across most NLP tasks. Unfortunately they are very expensive to train, and many small companies and research groups tend to use models that have been pre-trained and made available by third parties, rather than building their own. This is suboptimal as, for many languages, the models have been trained on smaller (or lower quality) corpora. In addition, monolingual pre-trained models for non-English languages are not always available. At best, models for those languages are included in multilingual versions, where each language shares the quota of substrings and parameters with the rest of the languages. This is particularly true for smaller languages such as Basque. In this paper we show that a number of monolingual models (FastText word embeddings, FLAIR and BERT language models) trained with larger Basque corpora produce much better results than publicly available versions in downstream NLP tasks, including topic classification, sentiment classification, PoS tagging and NER. This work sets a new state-of-the-art in those tasks for Basque. All benchmarks and models used in this work are publicly available.
Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, Eneko Agirre
LREC3
2020 Conversational Question Answering in Low Resource Scenarios: A Dataset and Case Study for Basque
abstract
Conversational Question Answering (CQA) systems meet user information needs by having conversations with them, where answers to the questions are retrieved from text. There exist a variety of datasets for English, with tens of thousands of training examples, and pre-trained language models have allowed to obtain impressive results. The goal of our research is to test the performance of CQA systems under low-resource conditions which are common for most non-English languages: small amounts of native annotations and other limitations linked to low resource languages, like lack of crowdworkers or smaller wikipedias. We focus on the Basque language, and present the first non-English CQA dataset and results. Our experiments show that it is possible to obtain good results with low amounts of native data thanks to cross-lingual transfer, with quality comparable to those obtained for English. We also discovered that dialogue history models are not directly transferable to another language, calling for further research. The dataset is publicly available.
Arantxa Otegi, Aitor Gonzalez-Agirre, Jon Ander Campos, Aitor Soroa, Eneko Agirre
LREC3