VLDB 2026 Research / reviewers in the wild / expert
Alexander Miserlis Hoyle
dblp:297/8769
· DBLP profile ↗
14ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0004-3375-0470ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsabstractAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ďurech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabolčec, Yixuan Xu, Michael Aerni, Badr AlKhamissi, Inés Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Milan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao, Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert i Llaquet, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Durech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan Eghlidi, Skander Moalla, Tiancheng Chen, Vinko Sabolcec, Yixuan Even Xu, Michael Aerni, Badr AlKhamissi, Ines Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein 0002, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush K. Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Alexander Ilic, Ana Klimovic, Andreas Krause 0001, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag |
ACL (1) | 44 |
| 2025 | ProxAnn: Use-Oriented Evaluations of Topic Models and Document ClusteringabstractAlexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan Lee Boyd-Graber, Philip Resnik. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Alexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 1 |
| 2025 | Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic ModelsabstractA common use of NLP by social scientists is to understand large document collections. Recent data exploration and content analysis have shifted from probabilistic topic models to Large Language Models (LLMs). Yet their effectiveness in helping users understand content in real-world applications remains under explored. This study compares the knowledge users gain from unsupervised LLMs, supervised LLMs, and traditional topic models across two datasets. While unsupervised LLMs generate more human-readable topics, their topics are overly generic for domain-specific datasets and do not help users learn much about the documents. Adding human supervision to LLM generation improves data exploration by mitigating hallucination and over-genericity but requires greater human effort. Traditional topic models, such as Latent Dirichlet Allocation (LDA), remain effective for exploration but are less user-friendly. LLMs struggle to describe the haystack of large corpora without human help, particularly domain-specific data, and face scaling and hallucination limitations due to context length constraints. Zongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu, Daniel Kofi Stephens, Juan Francisco Fung, Alden Dima, Jordan L. Boyd-Graber |
ACL (1) | 3 |
| 2025 | The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept ErasureabstractEmbedding-based similarity metrics between text sequences can be influenced not just by the content dimensions we most care about, but can also be biased by spurious attributes like the text's source or language.These document confounders cause problems for many applications, but especially those that need to pool texts from different corpora.This paper shows that a debiasing algorithm that removes information about observed confounders from the encoder representations substantially reduces these biases at a minimal computational cost.Document similarity and clustering metrics improve across every embedding variant and task we evaluate-often dramatically.Interestingly, performance on out-of-distribution benchmarks is not impacted, indicating that the embeddings are not otherwise degraded.1 Yu Fan 0007, Shauli Ravfogel, Mrinmaya Sachan, Elliott Ash, Alexander Miserlis Hoyle |
EMNLP | 6 |
| 2025 | Measuring scalar constructs in social science with LLMsabstractHauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel, Niklas Stoehr, Elliott Ash, Alexander Miserlis Hoyle. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel 0001, Niklas Stoehr, Elliott Ash, Alexander Miserlis Hoyle |
EMNLP | 7 |
| 2025 | How Persuasive Is Your Context?abstractTwo central capabilities of language models (LMs) are: (i) drawing on prior knowledge about entities, which allows them to answer queries such as What's the official language of Austria?, and (ii) adapting to new information provided in context, e.g., Pretend the official language of Austria is Tagalog., that is pre-pended to the question.In this article, we introduce targeted persuasion score (TPS), designed to quantify how persuasive a given context is to an LM where persuasion is operationalized as the ability of the context to alter the LM's answer to the question.In contrast to evaluating persuasiveness only by inspecting the greedily decoded answer under the model, TPS provides a more fine-grained view of model behavior.Based on the Wasserstein distance, TPS measures how much a context shifts a model's original answer distribution toward a target distribution.Empirically, through a series of experiments, we show that TPS captures a more nuanced notion of persuasiveness than previously proposed metrics. Kevin Du, Alexander Miserlis Hoyle, Ryan Cotterell |
EMNLP | 3 |
| 2024 | A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning StickabstractNishant Balepur, Matthew Shu, Alexander Hoyle, Alison Robey, Shi Feng, Seraphina Goldfarb-Tarrant, Jordan Lee Boyd-Graber. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nishant Balepur, Matthew Shu, Alexander Miserlis Hoyle, Alison Robey, Shi Feng 0005, Seraphina Goldfarb-Tarrant, Jordan L. Boyd-Graber |
EMNLP | 3 |
| 2024 | TopicGPT: A Prompt-based Topic Modeling FrameworkabstractChau Minh Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Alexander Miserlis Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer |
NAACL-HLT | 2 |
| 2023 | Natural Language Decompositions of Implicit Content Enable Better Text RepresentationsabstractWhen people interpret text, they rely on inferences that go beyond the observed language itself.Inspired by this observation, we introduce a method for the analysis of text that takes implicitly communicated content explicitly into account.We use a large language model to produce sets of propositions that are inferentially related to the text that has been observed, then validate the plausibility of the generated content via human judgments.Incorporating these explicit representations of implicit content proves useful in multiple problem settings that involve the human interpretation of utterances: assessing the similarity of arguments, making sense of a body of opinion data, and modeling legislative behavior.Our results suggest that modeling the meanings behind observed language, rather than the literal text alone, is a valuable direction for NLP and particularly its applications to social science. 1 * Equal contribution. Alexander Miserlis Hoyle, Rupak Sarkar, Pranav Goel 0001, Philip Resnik |
EMNLP | 1 |
| 2023 | Revisiting Automated Topic Model Evaluation with Large Language ModelsabstractTopic models help make sense of large text collections.Automatically evaluating their output and determining the optimal number of topics are both longstanding challenges, with no effective automated solutions to date.This paper evaluates the effectiveness of large language models (LLMs) for these tasks.We find that LLMs appropriately assess the resulting topics, correlating more strongly with human judgments than existing automated metrics.However, the type of evaluation task matters -LLMs correlate better with coherence ratings of word sets than on a word intrusion task.We find that LLMs can also guide users toward a reasonable number of topics.In actual applications, topic models are typically used to answer a research question related to a collection of texts.We can incorporate this research question in the prompt to the LLM, which helps estimate the optimal number of topics.github.com/dominiksinsaarland/ evaluating-topic-model-output Dominik Stammbach, Vilém Zouhar, Alexander Miserlis Hoyle, Mrinmaya Sachan, Elliott Ash |
EMNLP | 3 |
| 2021 | Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?abstractPedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, Jordan Boyd-Graber. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pedro Rodríguez 0001, Joe Barrow, Alexander Miserlis Hoyle, John Lalor, Robin Jia, Jordan L. Boyd-Graber |
ACL/IJCNLP (1) | 3 |
| 2021 | Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceabstractTopic model evaluation, like evaluation of other unsupervised methods, can be contentious. However, the field has coalesced around automated estimates of topic coherence, which rely on the frequency of word co-occurrences in a reference corpus. Contemporary neural topic models surpass classical ones according to these metrics. At the same time, topic model evaluation suffers from a validation gap: automated coherence, developed for classical models, has not been validated using human experimentation for neural models. In addition, a meta-analysis of topic modeling literature reveals a substantial standardization gap in automated topic modeling benchmarks. To address the validation gap, we compare automated coherence with the two most widely accepted human judgment tasks: topic rating and word intrusion. To address the standardization gap, we systematically evaluate a dominant classical model and two state-of-the-art neural models on two commonly used datasets. Automated evaluations declare a winning model when corresponding human evaluations do not, calling into question the validity of fully automatic evaluations independent of human judgments. Alexander Miserlis Hoyle, Pranav Goel 0001, Andrew Hian-Cheong, Denis Peskov, Jordan L. Boyd-Graber, Philip Resnik |
NeurIPS | 1 |
| 2020 | Improving Neural Topic Models using Knowledge DistillationabstractTopic models are often used to identify humaninterpretable topics to help make sense of large document collections.We use knowledge distillation to combine the best attributes of probabilistic topic models and pretrained transformers.Our modular method can be straightforwardly applied with any neural topic model to improve topic quality, which we demonstrate using two models having disparate architectures, obtaining state-of-the-art topic coherence.We show that our adaptable framework not only improves performance in the aggregate over all estimated topics, as is commonly reported, but also in head-to-head comparisons of aligned topics. Alexander Miserlis Hoyle, Pranav Goel 0001, Philip Resnik |
EMNLP (1) | 1 |
| 2019 | Unsupervised Discovery of Gendered Language through Latent-Variable ModelingabstractStudying the ways in which language is gendered has long been an area of interest in sociolinguistics.Studies have explored, for example, the speech of male and female characters in film and the language used to describe male and female politicians.In this paper, we aim not to merely study this phenomenon qualitatively, but instead to quantify the degree to which the language used to describe men and women is different and, moreover, different in a positive or negative way.To that end, we introduce a generative latent-variable model that jointly represents adjective (or verb) choice, with its sentiment, given the natural gender of a head (or dependent) noun.We find that there are significant differences between descriptions of male and female nouns and that these differences align with common gender stereotypes: Positive adjectives used to describe women are more often related to their bodies than adjectives used to describe men. Alexander Miserlis Hoyle, Lawrence Wolf-Sonkin, Hanna M. Wallach, Isabelle Augenstein, Ryan Cotterell |
ACL (1) | 1 |