EDBT 2026 Demo / reviewers in the wild / expert
Alejandro Benito-Santos
dblp:190/6901 · also Alejandro Benito
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-5317-6390ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 61% Information extraction and text analysis · 30% Language models and text generation · 9% | |
| Human-computer interaction and pervasive computing
1 paper |
Design research and methods · 77% Usability and user experience research · 23% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 3 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visualization and visual analytics
visualization evaluation |
1.0 | 1 | 2026 | Chasing Meaning and/or Insight? A Survey on Evaluation Practices at the Intersection of Visualization and the Humanities · CHI 2026 |
Machine learning › Trustworthy machine learning
annotator disagreement |
0.9 | 1 | 2025 | Beyond Averages: Learning with Annotator Disagreement in STS · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity |
0.9 | 1 | 2025 | Beyond Averages: Learning with Annotator Disagreement in STS · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
systematic survey · 2.0truncated gaussian · 0.9temperature scaling · 0.9ordinal distribution prediction · 0.9linear mixed-effects models · 0.9cross-encoder · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chasing Meaning and/or Insight? A Survey on Evaluation Practices at the Intersection of Visualization and the HumanitiesabstractThe intersection of visualization and the humanities (VIS*H) is marked by a tension between chasing analytical "insight" and interpretive "meaning." The effectiveness of visualization techniques hinges on established evaluation frameworks that assess both analytical utility and communicative efficacy, creating a potential mismatch with the non-positivist, interpretive aims of humanities scholarship. To examine how this tension manifests in practice, we systematically surveyed 171 VIS*H design studies to analyze their evaluation workflows and rigor according to standard practice. Our findings reveal recurring flaws, such as an over-reliance on monomethod approaches, and show that higher-quality evaluations emerge from workflows that effectively triangulate diverse evidence. From these findings, we derive recommendations to refine quality and validation criteria for humanities visualizations, and juxtapose them to ongoing critical debates in the field, ultimately arguing for a paradigm shift that can reconcile the advantages of established validation techniques with the interpretive depth required for humanistic inquiry. Alejandro Benito-Santos, Florian Windhager, Aida Horaniet Ibañez, Rabea Kleymann, Alfie Abdul-Rahman, Eva Mayr |
CHI | 1 |
| 2025 | Robust Estimation of Population-Level Effects in Repeated-Measures NLP Experimental DesignsabstractNLP research frequently grapples with multiple sources of variability-spanning runs, datasets, annotators, and more-yet conventional analysis methods often neglect these hierarchical structures, threatening the reproducibility of findings.To address this gap, we contribute a case study illustrating how linear mixed-effects models (LMMs) can rigorously capture systematic language-dependent differences (i.e., population-level effects) in a population of monolingual and multilingual language models.In the context of a bilingual hate speech detection task, we demonstrate that LMMs can uncover significant population-level effects-even under low-resource (small-N) experimental designs-while mitigating confounds and random noise.By setting out a transparent blueprint for repeated-measures experimentation, we encourage the NLP community to embrace variability as a feature, rather than a nuisance, in order to advance more robust, reproducible, and ultimately trustworthy results. Alejandro Benito-Santos, Adrián Ghajari, Víctor Fresno-Fernández |
ACL (1) | 1 |
| 2025 | Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal ContaminationabstractIn this article we present UNED-ACCESS 2024, a bilingual dataset that consists of 1003 multiple-choice questions of university entrance level exams in Spanish and English. Questions are originally formulated in Spanish and manually translated into English, and have not ever been publicly released, ensuring minimal contamination when evaluating Large Language Models with this dataset. A selection of current open-source and proprietary models are evaluated in a uniform zero-shot experimental setting both on the UNED-ACCESS 2024 dataset and on an equivalent subset of MMLU questions. Results show that (i) Smaller models not only perform worse than the largest models, but also degrade faster in Spanish than in English. The performance gap between both languages is negligible for the best models, but grows up to 37% for smaller models; (ii) Model ranking on UNED-ACCESS 2024 is almost identical (0.98 Pearson correlation) to the one obtained with MMLU (a similar, but publicly available benchmark), suggesting that contamination affects similarly to all models, and (iii) As in publicly available datasets, reasoning questions in UNED-ACCESS are more challenging for models of all sizes. Eva Sánchez-Salido, Roser Morante, Julio Gonzalo 0001, Guillermo Marco, Jorge Carrillo de Albornoz, Laura Plaza, Enrique Amigó, Andrés Fernández García, Alejandro Benito-Santos, Adrián Ghajari Espinosa, Víctor Fresno-Fernández |
COLING | 9 |
| 2025 | Beyond Averages: Learning with Annotator Disagreement in STSabstractThis work investigates capturing and modeling disagreement in Semantic Textual Similarity (STS), where sentence pairs are assigned ordinal similarity labels (0-5).Conventional STS systems average multiple annotator scores and focus on a single numeric estimate, overlooking label dispersion.By leveraging the disaggregated SemEval-2015 dataset (Soft-STS-15), this paper proposes and compares two disagreement-aware strategies that treat STS as an ordinal distribution prediction problem: a lightweight truncated Gaussian head for standard regression models, and a cross-encoder trained with a distance-aware objective, refined with temperature scaling.Results show improved performance in distance-based metrics, with the calibrated soft-label model proving best overall and notably more accurate on the most ambiguous pairs.This demonstrates that modeling disagreement benefits both calibration and ranking accuracy, highlighting the value of retaining and modeling full annotation distributions rather than collapsing them to a single mean label. Alejandro Benito-Santos, Adrián Ghajari |
EMNLP | 1 |
| 2025 | Test-driving information theory-based compositional distributional semantics: A case study on Spanish song lyricsabstractThe registered version of this article, first published in “Knowledge-Based Systems, vol. 319, 2025", is available online at the publisher's website: Elsevier, https://doi.org/10.1016/j.knosys.2025.113549 La versión registrada de este artículo, publicado por primera vez en “Knowledge-Based Systems, vol. 319, 2025", está disponible en línea en el sitio web del editor: Elsevier, https://doi.org/10.1016/j.knosys.2025.113549 Adrián Ghajari, Alejandro Benito-Santos, Salvador Ros 0001, Víctor Fresno-Fernández, Elena González-Blanco |
Knowl. Based Syst. | 2 |