Martín Santillán Cooper

dblp:324/1684 · also Martin Santillan Cooper · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0004-3808-0509ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
3 papers
Human-AI interaction · 58% Usability and user experience research · 42%
Artificial intelligence
1 paper
Language models and text generation · 100%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 5 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
EvalAssist: LLM-as-a-Judge Simplified · AAAI 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.912025
EvalAssist: LLM-as-a-Judge Simplified · AAAI 2025
Human-AI interaction › AI-assisted decision-making
AI-assisted evaluation
0.912025
EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences · UIST 2025
Usability and user experience research › evaluation methodology
evaluation framework
0.612022
A Simulation-Based Evaluation Framework for Interactive AI Systems and Its Application · AAAI 2022
Human-AI interaction
simulation-based evaluation
0.612022
InteractEva: A Simulation-Based Evaluation Framework for Interactive AI Systems · AAAI 2022

Methods — techniques the papers use, named apart from their topics

user simulation · 1.1token-probability based judgement · 0.9positional bias checking · 0.9direct assessment evaluation · 0.9certainty estimation · 0.9simulation-based evaluation · 0.6
YearPublicationVenuePosition
2025 EvalAssist: LLM-as-a-Judge Simplified
abstract
We present EvalAssist, a framework that simplifies the LLM- as-a-judge workflow. The system provides an online criteria development environment, where users can interactively build, test, and share custom evaluation criteria in a structured and portable format. A library of LLM based evaluators is made available that incorporates various algorithmic innovations such as token-probability based judgement, positional bias checking, and certainty estimation that help to engender trust in the evaluation process. We have computed extensive benchmarks and also deployed the system internally in our organization with several hundreds of users.
Michael Desmond, Zahra Ashktorab, Werner Geyer, Elizabeth Daly, Martín Santillán Cooper, Rahul Nair 0004, Nico Wagner, Tejaswini Pedapati
AAAI5
2025 EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
abstract
User flow diagram for EvalAssist in the direct assessment evaluation, illustrating criteria definition, test data input, annotation, AI evaluator selection, result review, iterative adjustments, and criteria export for dataset-wide evaluation via SDK.
Zahra Ashktorab, Michael Desmond, James M. Johnson, Martín Santillán Cooper, Elizabeth Daly, Rahul Nair 0004, Tejaswini Pedapati, Hyo Jin Do, Werner Geyer
UIST5
2022 A Simulation-Based Evaluation Framework for Interactive AI Systems and Its Application
abstract
Interactive AI (IAI) systems are increasingly popular as the human-centered AI design paradigm is gaining strong traction. However, evaluating IAI systems, a key step in building such systems, is particularly challenging, as their output highly depends on the performed user actions. Developers often have to rely on limited and mostly qualitative data from ad-hoc user testing to assess and improve their systems. In this paper, we present InteractEva; a systematic evaluation framework for IAI systems. We also describe how we have applied InteractEva to evaluate a commercial IAI system, leading to both quality improvements and better data-driven design decisions.
Maeda F. Hanafi, Yannis Katsis, Martín Santillán Cooper, Yunyao Li 0001
AAAI3
2022 InteractEva: A Simulation-Based Evaluation Framework for Interactive AI Systems
abstract
Evaluating interactive AI (IAI) systems is a challenging task, as their output highly depends on the performed user actions. As a result, developers often depend on limited and mostly qualitative data derived from user testing to improve their systems. In this paper, we present InteractEva; a systematic evaluation framework for IAI systems. InteractEva employs (a) a user simulation backend to test the system against different use cases and user interactions at scale with (b) an interactive frontend allowing developers to perform important quantitative evaluation tasks, including acquiring a performance overview, performing error analysis, and conducting what-if studies. The framework has supported the evaluation and improvement of an industrial IAI text extraction system, results of which will be presented during our demonstration.
Yannis Katsis, Maeda F. Hanafi, Martín Santillán Cooper, Yunyao Li 0001
AAAI3
2022 Predicting future sedentary behaviour using wearable and mobile devices
Martín Santillán Cooper, Marcelo Gabriel Armentano
Inf. Process. Manag.1