Plamen Pasliev

dblp:270/8189 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
explanation robustness
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020
Machine learning › Trustworthy machine learning
fairness
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020
Machine learning › Trustworthy machine learning › fairness
fairwashing
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020
Machine learning › Trustworthy machine learning
interpretability
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020

Methods — techniques the papers use, named apart from their topics

differential geometry · 0.4
YearPublicationVenuePosition
2020 Fairwashing explanations with off-manifold detergent
abstract
Explanation methods promise to make black-box classifiers more transparent. As a result, it is hoped that they can act as proof for a sensible, fair and trustworthy decision-making process of the algorithm and thereby increase its acceptance by the end-users. In this paper, we show both theoretically and experimentally that these hopes are presently unfounded. Specifically, we show that, for any classifier $g$, one can always construct another classifier $\tilde{g}$ which has the same behavior on the data (same train, validation, and test error) but has arbitrarily manipulated explanation maps. We derive this statement theoretically using differential geometry and demonstrate it experimentally for various explanation methods, architectures, and datasets. Motivated by our theoretical insights, we then propose a modification of existing explanation methods which makes them significantly more robust.
Christopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, Pan Kessel
ICML2