Alejandro Benito-Santos

dblp:190/6901 · also Alejandro Benito · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-5317-6390ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 61% Information extraction and text analysis · 30% Language models and text generation · 9%
Human-computer interaction and pervasive computing
1 paper
Design research and methods · 77% Usability and user experience research · 23%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 3 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics
visualization evaluation
1.012026
Chasing Meaning and/or Insight? A Survey on Evaluation Practices at the Intersection of Visualization and the Humanities · CHI 2026
Machine learning › Trustworthy machine learning
annotator disagreement
0.912025
Beyond Averages: Learning with Annotator Disagreement in STS · EMNLP 2025
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity
0.912025
Beyond Averages: Learning with Annotator Disagreement in STS · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

systematic survey · 2.0truncated gaussian · 0.9temperature scaling · 0.9ordinal distribution prediction · 0.9linear mixed-effects models · 0.9cross-encoder · 0.9
YearPublicationVenuePosition
2026 Chasing Meaning and/or Insight? A Survey on Evaluation Practices at the Intersection of Visualization and the Humanities
abstract
The intersection of visualization and the humanities (VIS*H) is marked by a tension between chasing analytical "insight" and interpretive "meaning." The effectiveness of visualization techniques hinges on established evaluation frameworks that assess both analytical utility and communicative efficacy, creating a potential mismatch with the non-positivist, interpretive aims of humanities scholarship. To examine how this tension manifests in practice, we systematically surveyed 171 VIS*H design studies to analyze their evaluation workflows and rigor according to standard practice. Our findings reveal recurring flaws, such as an over-reliance on monomethod approaches, and show that higher-quality evaluations emerge from workflows that effectively triangulate diverse evidence. From these findings, we derive recommendations to refine quality and validation criteria for humanities visualizations, and juxtapose them to ongoing critical debates in the field, ultimately arguing for a paradigm shift that can reconcile the advantages of established validation techniques with the interpretive depth required for humanistic inquiry.
Alejandro Benito-Santos, Florian Windhager, Aida Horaniet Ibañez, Rabea Kleymann, Alfie Abdul-Rahman, Eva Mayr
CHI1
2025 Robust Estimation of Population-Level Effects in Repeated-Measures NLP Experimental Designs
abstract
NLP research frequently grapples with multiple sources of variability-spanning runs, datasets, annotators, and more-yet conventional analysis methods often neglect these hierarchical structures, threatening the reproducibility of findings.To address this gap, we contribute a case study illustrating how linear mixed-effects models (LMMs) can rigorously capture systematic language-dependent differences (i.e., population-level effects) in a population of monolingual and multilingual language models.In the context of a bilingual hate speech detection task, we demonstrate that LMMs can uncover significant population-level effects-even under low-resource (small-N) experimental designs-while mitigating confounds and random noise.By setting out a transparent blueprint for repeated-measures experimentation, we encourage the NLP community to embrace variability as a feature, rather than a nuisance, in order to advance more robust, reproducible, and ultimately trustworthy results.
Alejandro Benito-Santos, Adrián Ghajari, Víctor Fresno-Fernández
ACL (1)1
2025 Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination
abstract
In this article we present UNED-ACCESS 2024, a bilingual dataset that consists of 1003 multiple-choice questions of university entrance level exams in Spanish and English. Questions are originally formulated in Spanish and manually translated into English, and have not ever been publicly released, ensuring minimal contamination when evaluating Large Language Models with this dataset. A selection of current open-source and proprietary models are evaluated in a uniform zero-shot experimental setting both on the UNED-ACCESS 2024 dataset and on an equivalent subset of MMLU questions. Results show that (i) Smaller models not only perform worse than the largest models, but also degrade faster in Spanish than in English. The performance gap between both languages is negligible for the best models, but grows up to 37% for smaller models; (ii) Model ranking on UNED-ACCESS 2024 is almost identical (0.98 Pearson correlation) to the one obtained with MMLU (a similar, but publicly available benchmark), suggesting that contamination affects similarly to all models, and (iii) As in publicly available datasets, reasoning questions in UNED-ACCESS are more challenging for models of all sizes.
Eva Sánchez-Salido, Roser Morante, Julio Gonzalo 0001, Guillermo Marco, Jorge Carrillo de Albornoz, Laura Plaza, Enrique Amigó, Andrés Fernández García, Alejandro Benito-Santos, Adrián Ghajari Espinosa, Víctor Fresno-Fernández
COLING9
2025 Beyond Averages: Learning with Annotator Disagreement in STS
abstract
This work investigates capturing and modeling disagreement in Semantic Textual Similarity (STS), where sentence pairs are assigned ordinal similarity labels (0-5).Conventional STS systems average multiple annotator scores and focus on a single numeric estimate, overlooking label dispersion.By leveraging the disaggregated SemEval-2015 dataset (Soft-STS-15), this paper proposes and compares two disagreement-aware strategies that treat STS as an ordinal distribution prediction problem: a lightweight truncated Gaussian head for standard regression models, and a cross-encoder trained with a distance-aware objective, refined with temperature scaling.Results show improved performance in distance-based metrics, with the calibrated soft-label model proving best overall and notably more accurate on the most ambiguous pairs.This demonstrates that modeling disagreement benefits both calibration and ranking accuracy, highlighting the value of retaining and modeling full annotation distributions rather than collapsing them to a single mean label.
Alejandro Benito-Santos, Adrián Ghajari
EMNLP1
2025 Test-driving information theory-based compositional distributional semantics: A case study on Spanish song lyrics
abstract
The registered version of this article, first published in “Knowledge-Based Systems, vol. 319, 2025", is available online at the publisher's website: Elsevier, https://doi.org/10.1016/j.knosys.2025.113549 La versión registrada de este artículo, publicado por primera vez en “Knowledge-Based Systems, vol. 319, 2025", está disponible en línea en el sitio web del editor: Elsevier, https://doi.org/10.1016/j.knosys.2025.113549
Adrián Ghajari, Alejandro Benito-Santos, Salvador Ros 0001, Víctor Fresno-Fernández, Elena González-Blanco
Knowl. Based Syst.2