Salim I. Amoukou

dblp:289/1335 · also Salim Ibrahim Amoukou · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 73% Language models and text generation · 17% Question answering and dialogue systems · 8%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.232025
Regional Explanations: Bridging Local and Global Variable Importance · NeurIPS 2025
Counterfactual Metarules for Local and Global Recourse · ICML 2024
Consistent Sufficient Explanations and Minimal Local Rules for explaining the decision of any classifier or regressor · NeurIPS 2022
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
0.912025
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models · ICML 2025
Natural language and speech › Question answering and dialogue systems
answer aggregation
0.912025
Representation Consistency for Accurate and Coherent LLM Answer Aggregation · NeurIPS 2025
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.912025
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models · ICML 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Representation Consistency for Accurate and Coherent LLM Answer Aggregation · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
0.812024
Counterfactual Metarules for Local and Global Recourse · ICML 2024
Machine learning › Trustworthy machine learning › robustness › distribution shift
distribution shift detection
0.812024
Sequential Harmful Shift Detection Without Labels · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Sequential Harmful Shift Detection Without Labels · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › logic-based explanation
rule-based explanation
0.812024
Counterfactual Metarules for Local and Global Recourse · ICML 2024
Machine learning › Trustworthy machine learning › interpretability › logic-based explanation
abductive explanation
0.612022
Consistent Sufficient Explanations and Minimal Local Rules for explaining the decision of any classifier or regressor · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability › logic-based explanation › rule-based explanation
local rule-based explanation
0.612022
Consistent Sufficient Explanations and Minimal Local Rules for explaining the decision of any classifier or regressor · NeurIPS 2022
Machine learning › Learning theory › statistical estimation
error estimation
0.212024
Sequential Harmful Shift Detection Without Labels · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

sparse autoencoder · 0.9shapley value · 0.9representation similarity · 0.9mechanistic interpretability · 0.9calibration · 0.9LIME · 0.9tree-based surrogate models · 0.8sequential testing · 0.8proxy error estimator · 0.8meta-rules · 0.8
YearPublicationVenuePosition
2025 To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
abstract
We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or oversteering, MERA addresses these limitations by (i) optimising the intervention direction, and (ii) calibrating when and how much to steer, thereby provably improving performance or abstaining when no confident correction is possible. Experiments across diverse datasets and LM families demonstrate safe, effective, non-degrading error correction and that MERA outperforms existing baselines. Moreover, MERA can be applied on top of existing steering techniques to further enhance their performance, establishing it as a general-purpose and efficient approach to mechanistic activation steering.
Anna Hedström, Salim I. Amoukou, Tom Bewley, Saumitra Mishra, Manuela M. Veloso
ICML2
2025 Regional Explanations: Bridging Local and Global Variable Importance
abstract
We analyze two widely used local attribution methods, Local Shapley Values and LIME, which aim to quantify the contribution of a feature value $x_i$ to a specific prediction $f(x_1, \dots, x_p)$. Despite their widespread use, we identify fundamental limitations in their ability to reliably detect locally important features, even under ideal conditions with exact computations and independent features. We argue that a sound local attribution method should not assign importance to features that neither influence the model output (e.g., features with zero coefficients in a linear model) nor exhibit statistical dependence with functionality-relevant features. We demonstrate that both Local SV and LIME violate this fundamental principle. To address this, we propose R-LOCO (Regional Leave Out COvariates), which bridges the gap between local and global explanations and provides more accurate attributions. R-LOCO segments the input space into regions with similar feature importance characteristics. It then applies global attribution methods within these regions, deriving an instance's feature contributions from its regional membership. This approach delivers more faithful local attributions while avoiding local explanation instability and preserving instance-specific detail often lost in global methods.
Salim I. Amoukou, Nicolas J.-B. Brunel
NeurIPS1
2025 Representation Consistency for Accurate and Coherent LLM Answer Aggregation
abstract
Test-time scaling improves large language models' (LLMs) performance by allocating more compute budget during inference. To achieve this, existing methods often require intricate modifications to prompting and sampling strategies. In this work, we introduce representation consistency (RC), a test-time scaling method for aggregating answers drawn from multiple candidate responses of an LLM regardless of how they were generated, including variations in prompt phrasing and sampling strategy. RC enhances answer aggregation by not only considering the number of occurrences of each answer in the candidate response set, but also the consistency of the model's internal activations while generating the set of responses leading to each answer. These activations can be either dense (raw model activations) or sparse (encoded via pretrained sparse autoencoders). Our rationale is that if the model's representations of multiple responses converging on the same answer are highly variable, this answer is more likely to be the result of incoherent reasoning and should be down-weighted during aggregation. Importantly, our method only uses cached activations and lightweight similarity computations and requires no additional model queries. Through experiments with four open-source LLMs and four reasoning datasets, we validate the effectiveness of RC for improving task performance during inference, with consistent accuracy improvements (up to 4\%) over strong test-time scaling baselines. We also show that consistency in the sparse activation signals aligns well with the common notion of coherent reasoning.
Junqi Jiang, Tom Bewley, Salim I. Amoukou, Francesco Leofante, Antonio Rago 0001, Saumitra Mishra, Francesca Toni
NeurIPS3
2024 PICE: Polyhedral Complex Informed Counterfactual Explanations
abstract
Polyhedral geometry can be used to shed light on the behaviour of piecewise linear neural networks, such as ReLU-based architectures. Counterfactual explanations are a popular class of methods for examining model behaviour by comparing a query to the closest point with a different label, subject to constraints. We present a new algorithm, Polyhedral-complex Informed Counterfactual Explanations (PICE), which leverages the decomposition of the piecewise linear neural network into a polyhedral complex to find counterfactuals that are provably minimal in the Euclidean norm and exactly on the decision boundary for any given query. Moreover, we develop variants of the algorithm that target popular counterfactual desiderata such as sparsity, robustness, speed, plausibility, and actionability. We empirically show on four publicly available real-world datasets that our method outperforms other popular techniques to find counterfactuals and adversarial attacks by distance to decision boundary and distance to query. Moreover, we successfully improve our baseline method in the dimensions of the desiderata we target, as supported by experimental evaluations.
Mattia J. Villani, Emanuele Albini, Saumitra Mishra, Salim I. Amoukou, Daniele Magazzeni, Manuela M. Veloso
AIES (1)5
2024 Counterfactual Metarules for Local and Global Recourse
abstract
We introduce T-CREx, a novel model-agnostic method for local and global counterfactual explanation (CE), which summarises recourse options for both individuals and groups in the form of generalised rules. It leverages tree-based surrogate models to learn the counterfactual rules, alongside metarules denoting their regimes of optimality, providing both a global analysis of model behaviour and diverse recourse options for users. Experiments indicate that T-CREx achieves superior aggregate performance over existing rule-based baselines on a range of CE desiderata, while being orders of magnitude faster to run.
Tom Bewley, Salim I. Amoukou, Saumitra Mishra, Daniele Magazzeni, Manuela M. Veloso
ICML2
2024 Sequential Harmful Shift Detection Without Labels
abstract
We introduce a novel approach for detecting distribution shifts that negatively impact the performance of machine learning models in continuous production environments, which requires no access to ground truth data labels. It builds upon the work of Podkopaev and Ramdas [2022], who address scenarios where labels are available for tracking model errors over time. Our solution extends this framework to work in the absence of labels, by employing a proxy for the true error. This proxy is derived using the predictions of a trained error estimator. Experiments show that our method has high power and false alarm control under various distribution shifts, including covariate and label shifts and natural shifts over geography and time.
Salim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué, Daniele Magazzeni, Manuela M. Veloso
NeurIPS1
2022 Accurate Shapley Values for explaining tree-based models
abstract
Although Shapley Values (SV) are widely used in explainable AI, they can be poorly understood and estimated, implying that their analysis may lead to spurious inferences and explanations. As a starting point, we remind an invariance principle for SV and derive the correct approach for computing the SV of categorical variables that are particularly sensitive to the encoding used. In the case of tree-based models, we introduce two estimators of Shapley Values that exploit the tree structure efficiently and are more accurate than state-of-the-art methods. Simulations and comparisons are performed with state-of-the-art algorithms and show the practical gain of our approach. Finally, we discuss the ability of SV to provide reliable local explanations. We also provide a Python package that compute our estimators at https://github.com/salimamoukou/acv00.
Salim I. Amoukou, Tangi Salaün, Nicolas J.-B. Brunel
AISTATS1
2022 Consistent Sufficient Explanations and Minimal Local Rules for explaining the decision of any classifier or regressor
abstract
To explain the decision of any regression and classification model, we extend the notion of probabilistic sufficient explanations (P-SE). For each instance, this approach selects the minimal subset of features that is sufficient to yield the same prediction with high probability, while removing other features. The crux of P-SE is to compute the conditional probability of maintaining the same prediction. Therefore, we introduce an accurate and fast estimator of this probability via random Forests for any data $(\boldsymbol{X}, Y)$ and show its efficiency through a theoretical analysis of its consistency. As a consequence, we extend the P-SE to regression problems. In addition, we deal with non-discrete features, without learning the distribution of $\boldsymbol{X}$ nor having the model for making predictions. Finally, we introduce local rule-based explanations for regression/classification based on the P-SE and compare our approaches w.r.t other explainable AI methods. These methods are available as a Python Package.
Salim I. Amoukou, Nicolas J.-B. Brunel
NeurIPS1