Dimitrios Rontogiannis

dblp:350/4384 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 96% Language models and text generation · 4%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
1.722026
GLANCE: Global Actions in a Nutshell for Counterfactual Explainability · AAAI 2026
Fairness Aware Counterfactuals for Subgroups · NeurIPS 2023
Machine learning › Trustworthy machine learning
fairness
1.722026
GLANCE: Global Actions in a Nutshell for Counterfactual Explainability · AAAI 2026
Fairness Aware Counterfactuals for Subgroups · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
1.722026
GLANCE: Global Actions in a Nutshell for Counterfactual Explainability · AAAI 2026
Fairness Aware Counterfactuals for Subgroups · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability › counterfactual explanation
algorithmic recourse
1.012026
GLANCE: Global Actions in a Nutshell for Counterfactual Explainability · AAAI 2026
Program synthesis and code generation
code generation with language models
1.012026
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks · AAAI 2026
Machine learning › Trustworthy machine learning › fairness › group fairness
subgroup fairness
0.712023
Fairness Aware Counterfactuals for Subgroups · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.312026
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks · AAAI 2026

Methods — techniques the papers use, named apart from their topics

requirement dependency graph · 2.0feedback-driven dialogue · 2.0expert annotation · 2.0model-agnostic auditing · 1.3counterfactual explanation · 1.3counterfactual generation · 1.0agglomerative clustering · 1.0
YearPublicationVenuePosition
2026 GLANCE: Global Actions in a Nutshell for Counterfactual Explainability
abstract
The widespread deployment of machine learning systems in critical real-world decision-making applications has highlighted the urgent need for counterfactual explainability methods that operate effectively. Global counterfactual explanations, expressed as actions to offer recourse, aim to provide succinct explanations and insights applicable to large population subgroups. High effectiveness, measured by the fraction of the population that is provided recourse, ensures that the actions benefit as many individuals as possible. Keeping the cost of actions low ensures the proposed recourse actions remain practical and actionable. Limiting the number of actions that provide global counterfactuals is essential to maximize interpretability. The primary challenge, therefore, is to balance these trade-offs—maximizing effectiveness, minimizing cost, while maintaining a small number of actions. We introduce GLANCE, a versatile and adaptive algorithm that employs a novel agglomerative approach, jointly considering both the feature space and the space of counterfactual actions, thereby accounting for the distribution of points in a way that aligns with the model's structure. This design enables the careful balancing of the trade-offs among the three key objectives, with the size objective functioning as a tunable parameter to keep the actions few and easy to interpret. Our extensive experimental evaluation demonstrates that GLANCE consistently shows greater robustness and performance compared to existing methods across various datasets and models.
Loukas Kavouras, Eleni Psaroudaki, Konstantinos Tsopelas, Dimitrios Rontogiannis, Nikolas Theologitis, Dimitris Sacharidis, Giorgos Giannopoulos, Dimitrios Tomaras, Kleopatra Markou, Dimitrios Gunopulos, Dimitris Fotakis 0001, Ioannis Z. Emiris
AAAI4
2026 Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
abstract
Standard single-turn, static benchmarks fall short in evaluating the nuanced capabilities of Large Language Models (LLMs) on complex tasks such as software engineering. In this work, we propose a novel interactive evaluation framework that assesses LLMs on multi-requirement programming tasks through structured, feedback-driven dialogue. Each task is modeled as a requirement dependency graph, and an "interviewer" LLM, aware of the ground-truth solution, provides minimal, targeted hints to an "interviewee" model to help correct errors and fulfill target constraints. This dynamic protocol enables fine-grained diagnostic insights into model behavior, uncovering strengths and systematic weaknesses that static benchmarks fail to measure. We build on DevAI, a benchmark of 55 curated programming tasks, by adding ground-truth solutions and evaluating the relevance and utility of interviewer hints through expert annotation. Our results highlight the importance of dynamic evaluation in advancing the development of collaborative code-generating agents.
Dimitrios Rontogiannis, Maxime Peyrard, Nicolas Mario Baldwin, Martin Josifoski, Robert West 0001, Dimitrios Gunopulos
AAAI1
2023 Fairness Aware Counterfactuals for Subgroups
abstract
In this work, we present Fairness Aware Counterfactuals for Subgroups (FACTS), a framework for auditing subgroup fairness through counterfactual explanations. We start with revisiting (and generalizing) existing notions and introducing new, more refined notions of subgroup fairness. We aim to (a) formulate different aspects of the difficulty of individuals in certain subgroups to achieve recourse, i.e. receive the desired outcome, either at the micro level, considering members of the subgroup individually, or at the macro level, considering the subgroup as a whole, and (b) introduce notions of subgroup fairness that are robust, if not totally oblivious, to the cost of achieving recourse. We accompany these notions with an efficient, model-agnostic, highly parameterizable, and explainable framework for evaluating subgroup fairness. We demonstrate the advantages, the wide applicability, and the efficiency of our approach through a thorough experimental evaluation on different benchmark datasets.
Loukas Kavouras, Konstantinos Tsopelas, Giorgos Giannopoulos, Dimitris Sacharidis, Eleni Psaroudaki, Nikolas Theologitis, Dimitrios Rontogiannis, Dimitris Fotakis 0001, Ioannis Z. Emiris
NeurIPS7