Ammar Shaker

dblp:08/8419 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 A Robust Prototype-Based Network with Interpretable RBF Classifier Foundations
abstract
Prototype-based classification learning methods are known to be inherently interpretable. However, this paradigm suffers from major limitations compared to deep models, such as lower performance. This led to the development of the so-called deep Prototype-Based Networks (PBNs), also known as prototypical parts models. In this work, we analyze these models with respect to different properties, including interpretability. In particular, we focus on the Classification-by-Components (CBC) approach, which uses a probabilistic model to ensure interpretability and can be used as a shallow or deep architecture. We show that this model has several shortcomings, like creating contradicting explanations. Based on these findings, we propose an extension of CBC that solves these issues. Moreover, we prove that this extension has robustness guarantees and derive a loss that optimizes robustness. Additionally, our analysis shows that most (deep) PBNs are related to (deep) RBF classifiers, which implies that our robustness guarantees generalize to shallow RBF classifiers. The empirical evaluation demonstrates that our deep PBN yields state-of-the-art classification accuracy on different benchmarks while resolving the interpretability shortcomings of other approaches. Further, our shallow PBN variant outperforms other shallow PBNs while being inherently interpretable and exhibiting provable robustness guarantees.
Sascha Saralajew, Ashish Rana, Thomas Villmann, Ammar Shaker
AAAI4
2024 A Human-Centric Assessment of the Usefulness of Attribution Methods in Computer Vision
Wiem Ben Rim, Ammar Shaker, Zhao Xu 0001, Kiril Gashteovski, Bhushan Kotnis, Carolin Lawrence, Jürgen Quittek, Sascha Saralajew
ECML/PKDD (5)2
2023 Multi-Source Survival Domain Adaptation
abstract
Survival analysis is the branch of statistics that studies the relation between the characteristics of living entities and their respective survival times, taking into account the partial information held by censored cases. A good analysis can, for example, determine whether one medical treatment for a group of patients is better than another. With the rise of machine learning, survival analysis can be modeled as learning a function that maps studied patients to their survival times. To succeed with that, there are three crucial issues to be tackled. First, some patient data is censored: we do not know the true survival times for all patients. Second, data is scarce, which led past research to treat different illness types as domains in a multi-task setup. Third, there is the need for adaptation to new or extremely rare illness types, where little or no labels are available. In contrast to previous multi-task setups, we want to investigate how to efficiently adapt to a new survival target domain from multiple survival source domains. For this, we introduce a new survival metric and the corresponding discrepancy measure between survival distributions. These allow us to define domain adaptation for survival analysis while incorporating censored data, which would otherwise have to be dropped. Our experiments on two cancer data sets reveal a superb performance on target domains, a better treatment recommendation, and a weight matrix with a plausible explanation.
Ammar Shaker, Carolin Lawrence
AAAI1
2022 Learning to Transfer with von Neumann Conditional Divergence
abstract
The similarity of feature representations plays a pivotal role in the success of problems related to domain adaptation. Feature similarity includes both the invariance of marginal distributions and the closeness of conditional distributions given the desired response y (e.g., class labels). Unfortunately, traditional methods always learn such features without fully taking into consideration the information in y, which in turn may lead to a mismatch of the conditional distributions or the mixup of discriminative structures underlying data distributions. In this work, we introduce the recently proposed von Neumann conditional divergence to improve the transferability across multiple domains. We show that this new divergence is differentiable and eligible to easily quantify the functional dependence between features and y. Given multiple source tasks, we integrate this divergence to capture discriminative information in y and design novel learning objectives assuming those source tasks are observed either simultaneously or sequentially. In both scenarios, we obtain favorable performance against state-of-the-art methods in terms of smaller generalization error on new tasks and less catastrophic forgetting on source tasks (in the sequential setup).
Ammar Shaker, Shujian Yu, Daniel Oñoro-Rubio
AAAI1
2022 MILIE: Modular & Iterative Multilingual Open Information Extraction
abstract
Bhushan Kotnis, Kiril Gashteovski, Daniel Rubio, Ammar Shaker, Vanesa Rodriguez-Tembras, Makoto Takamoto, Mathias Niepert, Carolin Lawrence. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Bhushan Kotnis, Kiril Gashteovski, Daniel Oñoro-Rubio, Ammar Shaker, Vanesa Rodriguez-Tembras, Makoto Takamoto, Mathias Niepert, Carolin Lawrence
ACL (1)4
2022 Uncertainty Propagation in Node Classification
abstract
Quantifying predictive uncertainty of neural networks has recently attracted increasing attention. In this work, we focus on measuring uncertainty of graph neural networks (GNNs) for the task of node classification. Most existing GNNs model message passing among nodes. The messages are often deterministic. Questions naturally arise: Does there exist uncertainty in the messages? How could we propagate such uncertainty over a graph together with messages? To address these issues, we propose a Bayesian uncertainty propagation (BUP) method, which embeds GNNs in a Bayesian modeling framework, and models predictive uncertainty of node classification with Bayesian confidence of predictive probability and uncertainty of messages. Our method proposes a novel uncertainty propagation mechanism inspired by Gaussian models. Moreover, we present an uncertainty oriented loss for node classification that allows the GNNs to clearly integrate predictive uncertainty in learning procedure. Consequently, the training examples with large predictive uncertainty will be penalized. We demonstrate the BUP with respect to prediction reliability and out-of-distribution (OOD) predictions. The learned uncertainty is also analyzed in depth. The relations between uncertainty and graph topology, as well as predictive uncertainty in the OOD cases are investigated with extensive experiments. The empirical results with popular benchmark datasets demonstrate the superior performance of the proposed method.
Zhao Xu 0001, Carolin Lawrence, Ammar Shaker, Raman Siarheyeu
ICDM3
2022 Modular-Relatedness for Continual Learning
Ammar Shaker, Francesco Alesiani, Shujian Yu
IDA1
2021 Bilevel Continual Learning
abstract
Continual Learning (CL) studies the problem of learning a sequence of tasks, one at a time, such that the learning of each new task does not lead to the deterioration in performance on the previously seen ones while exploiting previously learned features. This paper presents Bilevel Continual Learning (BiCL), a general framework for continual learning that fuses bilevel optimization and recent advances in meta-learning for deep neural networks. BiCL is able to train both deep discriminative and generative models under the conservative setting of the online continual learning. Experimental results show that BiCL provides competitive performance in terms of accuracy for the current task while reducing the effect of catastrophic forgetting.
Ammar Shaker, Francesco Alesiani, Shujian Yu, Wenzhe Yin
IJCNN1
2021 TSK-Streams: learning TSK fuzzy systems for regression on data streams
abstract
Abstract The problem of adaptive learning from evolving and possibly non-stationary data streams has attracted a lot of interest in machine learning in the recent past, and also stimulated research in related fields, such as computational intelligence and fuzzy systems. In particular, several rule-based methods for the incremental induction of regression models have been proposed. In this paper, we develop a method that combines the strengths of two existing approaches rooted in different learning paradigms. More concretely, our method adopts basic principles of the state-of-the-art learning algorithm AMRules and enriches them by the representational advantages of fuzzy rules. In a comprehensive experimental study, TSK-Streams is shown to be highly competitive in terms of performance.
Ammar Shaker, Eyke Hüllermeier
Data Min. Knowl. Discov.1
2020 Measuring the Discrepancy between Conditional Distributions: Methods, Properties and Applications
abstract
We propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlying distributions in high-dimensional space and it operates on the cone of symmetric positive semidefinite (SPS) matrix using the Bregman matrix divergence. Moreover, it inherits the merits of the correntropy function to explicitly incorporate high-order statistics in the data. We present the properties of our new statistic and illustrate its connections to prior art. We finally show the applications of our new statistic on three different machine learning problems, namely the multi-task learning over graphs, the concept drift detection, and the information-theoretic feature selection, to demonstrate its utility and advantage. Code of our statistic is available at https://bit.ly/BregmanCorrentropy.
Shujian Yu, Ammar Shaker, Francesco Alesiani, José C. Príncipe
IJCAI2
2020 Online Meta-Forest for Regression Data Streams
abstract
Stream learning is essential when there is limited memory, time and computational power. However, existing streaming methods are mostly designed for classification with only a few exceptions for regression problems. Although being fast, the performance of these online regression methods is inadequate due to their dependence on merely linear models. Besides, only a few stream methods are based on meta-learning that aims at facilitating the dynamic choice of the right model. Nevertheless, these approaches are restricted to recommend learners on a window and not on the instance level. In this paper, we present a novel approach, named Online Meta-Forest, that incrementally induces an ensemble of meta-learners that selects the best set of predictors for each test example. Each meta-learner has the ability to find a non-linear mapping of the input space to the set of induced models. We conduct a series of experiments demonstrating that Online Meta-Forest outperforms related methods on 16 out of 25 evaluated benchmark and domain datasets in transportation.
Ammar Shaker, Christoph Gärtner, Shujian Yu
IJCNN1
2020 Towards Interpretable Multi-task Learning Using Bilevel Programming
Francesco Alesiani, Shujian Yu, Ammar Shaker, Wenzhe Yin
ECML/PKDD (2)3
2019 Efficient and Scalable Multi-Task Regression on Massive Number of Tasks
abstract
Many real-world large-scale regression problems can be formulated as Multi-task Learning (MTL) problems with a massive number of tasks, as in retail and transportation domains. However, existing MTL methods still fail to offer both the generalization performance and the scalability for such problems. Scaling up MTL methods to problems with a tremendous number of tasks is a big challenge. Here, we propose a novel algorithm, named Convex Clustering Multi-Task regression Learning (CCMTL), which integrates with convex clustering on the k-nearest neighbor graph of the prediction models. Further, CCMTL efficiently solves the underlying convex problem with a newly proposed optimization method. CCMTL is accurate, efficient to train, and empirically scales linearly in the number of tasks. On both synthetic and real-world datasets, the proposed CCMTL outperforms seven state-of-the-art (SoA) multi-task learning methods in terms of prediction accuracy as well as computational efficiency. On a real-world retail dataset with 23,812 tasks, CCMTL requires only around 30 seconds to train on a single thread, while the SoA methods need up to hours or even days.
Francesco Alesiani, Ammar Shaker
AAAI3
2018 MetaBags: Bagged Meta-Decision Trees for Regression
Jihed Khiari, Luís Moreira-Matias, Ammar Shaker, Bernard Zenko, Saso Dzeroski
ECML/PKDD (1)3
2017 Learning TSK Fuzzy Rules from Data Streams
Ammar Shaker, Waleri Heldt, Eyke Hüllermeier
ECML/PKDD (2)1
2017 Imprecise Matching of Requirements Specifications for Software Services Using Fuzzy Logic
abstract
Today, software components are provided by global markets in the form of services. In order to optimally satisfy service requesters and service providers, adequate techniques for automatic service matching are needed. However, a requester's requirements may be vague and the information available about a provided service may be incomplete. As a consequence, fuzziness is induced into the matching procedure. The contribution of this paper is the development of a systematic matching procedure that leverages concepts and techniques from fuzzy logic and possibility theory based on our formal distinction between different sources and types of fuzziness in the context of service matching. In contrast to existing methods, our approach is able to deal with imprecision and incompleteness in service specifications and to inform users about the extent of induced fuzziness in order to improve the user's decision-making. We demonstrate our approach on the example of specifications for service reputation based on ratings given by previous users. Our evaluation based on real service ratings shows the utility and applicability of our approach.
Marie Platenius-Mohr, Ammar Shaker, Matthias Becker 0001, Eyke Hüllermeier, Wilhelm Schäfer
IEEE Trans. Software Eng.2
2015 Recovery analysis for adaptive learning from non-stationary data streams: Experimental design and case study
Ammar Shaker, Eyke Hüllermeier
Neurocomputing1
2013 Evolving fuzzy pattern trees for binary classification on data streams
Ammar Shaker, Robin Senge, Eyke Hüllermeier
Inf. Sci.1