VLDB 2026 Research / reviewers in the wild / expert
Maximilian Mozes
dblp:225/7660
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8138-3792ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 63% Reinforcement learning · 11% Knowledge representation and reasoning · 9% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 77% Usable security · 23% |
Topics — the 13 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › prompt tuning
adversarial prompt tuning |
0.9 | 1 | 2025 | Reverse Engineering Human Preferences with Reinforcement Learning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.9 | 1 | 2025 | Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models · ICLR 2025 |
Machine learning › Reinforcement learning
learning from failure |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | Reverse Engineering Human Preferences with Reinforcement Learning · NeurIPS 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge structures
procedural knowledge |
0.9 | 1 | 2025 | Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation
prompting |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
text classification |
0.5 | 1 | 2021 | Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text Classification · EMNLP (1) 2021 |
Security and privacy of machine learning
adversarial attack |
0.5 | 1 | 2021 | Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text Classification · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.3 | 1 | 2018 | Identifying the narrative styles of YouTube's vloggers · EMNLP 2018 |
Web and social media mining
social media analysis |
0.3 | 1 | 2018 | Identifying the narrative styles of YouTube's vloggers · EMNLP 2018 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.3 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
word substitution · 1.0crowdsourcing · 1.0unsupervised clustering · 1.0dynamic intra-textual sentiment analysis · 1.0reinforcement learning · 0.9prompting · 0.9preamble generation · 0.9judge language model · 0.9influence functions · 0.9in-context learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-contextabstractShowing incorrect answers to Large Language Models (LLMs) is a popular strategy to improve their performance in reasoning-intensive tasks.It is widely assumed that, in order to be helpful, the incorrect answers must be accompanied by comprehensive rationales, explicitly detailing where the mistakes are and how to correct them.However, in this work we present a counterintuitive finding: we observe that LLMs perform better in math reasoning tasks when these rationales are eliminated from the context and models are left to infer on their own what makes an incorrect answer flawed.This approach also substantially outperforms chainof-thought prompting in our evaluations.These results are consistent across LLMs of different sizes and varying reasoning abilities.To gain an understanding of why LLMs learn from mistakes more effectively without explicit corrective rationales, we perform a thorough analysis, investigating changes in context length and answer diversity between different prompting strategies, and their effect on performance.We also examine evidence of overfitting to the in-context rationales when these are provided, and study the extent to which LLMs are able to autonomously infer high-quality corrective rationales given only incorrect answers as input.We find evidence that, while incorrect answers are more beneficial for LLM learning than additional diverse correct answers, explicit corrective rationales over-constrain the model, thus limiting those benefits. Lisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan, Marek Rei, Max Bartolo |
EMNLP | 2 |
| 2025 | Procedural Knowledge in Pretraining Drives Reasoning in Large Language ModelsabstractThe capabilities and limitations of Large Language Models (LLMs) have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies. The sheer volume of data used in the design of LLMs has precluded us from applying the method traditionally used to measure generalisation: train-test set separation. To overcome this, we study what kind of generalisation strategies LLMs employ when performing reasoning tasks by investigating the pretraining data they rely on. For two models of different sizes (7B and 35B) and 2.5B of their pretraining tokens, we identify what documents influence the model outputs for three simple mathematical reasoning tasks and contrast this to the data that are influential for answering factual questions. We find that, while the models rely on mostly distinct sets of data for each factual question, a document often has a similar influence across different reasoning questions within the same task, indicating the presence of procedural knowledge. We further find that the answers to factual questions often show up in the most influential data. However, for reasoning questions the answers usually do not show up as highly influential, nor do the answers to the intermediate reasoning steps. When we characterise the top ranked documents for the reasoning questions qualitatively, we confirm that the influential documents often contain procedural knowledge, like demonstrating how to obtain a solution using formulae or code. Our findings indicate that the approach to reasoning the models use is unlike retrieval, and more like a generalisable strategy that synthesises procedural knowledge from documents doing a similar form of reasoning. Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara, Dwaraknath Gnaneshwar, Acyr Locatelli, Robert Kirk, Tim Rocktäschel, Edward Grefenstette, Max Bartolo |
ICLR | 2 |
| 2025 | Reverse Engineering Human Preferences with Reinforcement LearningabstractThe capabilities of Large Language Models (LLMs) are routinely evaluated by other LLMs trained to predict human preferences. This framework—known as *LLM-as-a-judge*—is highly scalable and relatively low cost. However, it is also vulnerable to malicious exploitation, as LLM responses can be tuned to overfit the preferences of the judge. Previous work shows that the answers generated by a candidate-LLM can be edited *post hoc* to maximise the score assigned to them by a judge-LLM. In this study, we adopt a different approach and use the signal provided by judge-LLMs as a reward to adversarially tune models that generate text preambles designed to boost downstream performance. We find that frozen LLMs pipelined with these models attain higher LLM-evaluation scores than existing frameworks. Crucially, unlike other frameworks which intervene directly on the model's response, our method is virtually undetectable. We also demonstrate that the effectiveness of the tuned preamble generator transfers when the candidate-LLM and the judge-LLM are replaced with models that are not used during training. These findings raise important questions about the design of more reliable LLM-as-a-judge evaluation settings. They also demonstrate that human preferences can be reverse engineered effectively, by pipelining LLMs to optimise upstream preambles via reinforcement learning—an approach that could find future applications in diverse tasks and domains beyond adversarial attacks. Lisa Alazraki, Yi Chern Tan, Jon Ander Campos, Maximilian Mozes, Marek Rei, Max Bartolo |
NeurIPS | 4 |
| 2021 | Frequency-Guided Word Substitutions for Detecting Textual Adversarial ExamplesabstractRecent efforts have shown that neural text processing models are vulnerable to adversarial examples, but the nature of these examples is poorly understood. In this work, we show that adversarial attacks against CNN, LSTM and Transformer-based classification models perform word substitutions that are identifiable through frequency differences between replaced words and their corresponding substitutions. Based on these findings, we propose frequency-guided word substitutions (FGWS), a simple algorithm exploiting the frequency properties of adversarial word substitutions for the detection of adversarial examples. FGWS achieves strong performance by accurately detecting adversarial examples on the SST-2 and IMDb sentiment datasets, with F1 detection scores of up to 91.4% against RoBERTa-based classification models. We compare our approach against a recently proposed perturbation discrimination framework and show that we outperform it by up to 13.0% F1. Maximilian Mozes, Pontus Stenetorp, Bennett Kleinberg, Lewis D. Griffin |
EACL | 1 |
| 2021 | Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text ClassificationabstractResearch shows that natural language processing models are generally considered to be vulnerable to adversarial attacks; but recent work has drawn attention to the issue of validating these adversarial inputs against certain criteria (e.g., the preservation of semantics and grammaticality). Enforcing constraints to uphold such criteria may render attacks unsuccessful, raising the question of whether valid attacks are actually feasible. In this work, we investigate this through the lens of human language ability. We report on crowdsourcing studies in which we task humans with iteratively modifying words in an input text, while receiving immediate model feedback, with the aim of causing a sentiment classification model to misclassify the example. Our findings suggest that humans are capable of generating a substantial amount of adversarial examples using semantics-preserving word substitutions. We analyze how human-generated adversarial examples compare to the recently proposed TextFooler, Genetic, BAE and SememePSO attack algorithms on the dimensions naturalness, preservation of sentiment, grammaticality and substitution rate. Our findings suggest that human-generated adversarial examples are not more able than the best algorithms to generate natural-reading, sentiment-preserving examples, though they do so by being much more computationally efficient. Maximilian Mozes, Max Bartolo, Pontus Stenetorp, Bennett Kleinberg, Lewis D. Griffin |
EMNLP (1) | 1 |
| 2018 | Identifying the narrative styles of YouTube's vloggersabstractVlogs provide a rich public source of data in a novel setting.This paper examined the continuous sentiment styles employed in 27,333 vlogs using a dynamic intra-textual approach to sentiment analysis.Using unsupervised clustering, we identified seven distinct continuous sentiment trajectories characterized by fluctuations of sentiment throughout a vlog's narrative time.We provide a taxonomy of these seven continuous sentiment styles and found that vlogs whose sentiment builds up towards a positive ending are the most prevalent in our sample.Gender was associated with preferences for different continuous sentiment trajectories.This paper discusses the findings with respect to previous work and concludes with an outlook towards possible uses of the corpus, method and findings of this paper for related areas of research. Bennett Kleinberg, Maximilian Mozes, Isabelle van der Vegt |
EMNLP | 2 |