VLDB 2026 Research / reviewers in the wild / expert
Alexandros Kalousis
dblp:68/6004
· DBLP profile ↗
66ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0001-9734-195XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 25 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DARO: An Auction-Based Multi-Agent Reinforcement Learning Framework for Task Scheduling in the Cloud Continuum
Syed Mafooq Ul Hassan, Marios Touloupou, Jacopo Castellini, Pablo Strasser, Alexandros Kalousis, Herodotos Herodotou |
CLOSER | 5 |
| 2025 | GLAD: Improving Latent Graph Generative Modeling with Simple QuantizationabstractLearning graph generative models over latent spaces has received less attention compared to models that operate on the original data space and has so far demonstrated lacklustre performance. We present GLAD a latent space graph generative model. Unlike most previous latent space graph generative models, GLAD operates on a discrete latent space that preserves to a significant extent the discrete nature of the graph structures making no unnatural assumptions such as latent space continuity. We learn the prior of our discrete latent space by adapting diffusion bridges to its structure. By operating over an appropriately constructed latent space we avoid relying on decompositions that are often used in models that operate in the original data space. We present experiments on a series of graph benchmark datasets that demonstrates GLAD as the first equivariant latent graph generative method achieves competitive performance with the state of the art baselines. Van Khoa Nguyen, Yoann Boget, Frantzeska Lavda, Alexandros Kalousis |
AAAI | 4 |
| 2025 | MING: A Functional Approach to Learning Molecular Generative ModelsabstractTraditional molecule generation methods often rely on sequence- or graph-based representations, which can limit their expressive power or require complex permutation-equivariant architectures. This paper introduces a novel paradigm for learning molecule generative models based on functional representations. Specifically, we propose Molecular Implicit Neural Generation (MING), a diffusion-based model that learns molecular distributions in the function space. Unlike standard diffusion processes in the data space, MING employs a novel functional denoising probabilistic process, which jointly denoises information in both the function’s input and output spaces by leveraging an expectation-maximization procedure for latent implicit neural representations of data. This approach enables a simple yet effective model design that accurately captures underlying function distributions. Experimental results on molecule-related datasets demonstrate MING’s superior performance and ability to generate plausible molecular samples, surpassing state-of-the-art data-space methods while offering a more streamlined architecture and significantly faster generation times. The code is available at \url{https://github.com/v18nguye/MING.} Van Khoa Nguyen, Maciej Falkiewicz, Giangiacomo Mercatali, Alexandros Kalousis |
AISTATS | 4 |
| 2024 | Mimicking Better by Matching the Approximate Action DistributionabstractIn this paper, we introduce MAAD, a novel, sample-efficient on-policy algorithm for Imitation Learning from Observations. MAAD utilizes a surrogate reward signal, which can be derived from various sources such as adversarial games, trajectory matching objectives, or optimal transport criteria. To compensate for the non-availability of expert actions, we rely on an inverse dynamics model that infers plausible actions distribution given the expert’s state-state transitions; we regularize the imitator’s policy by aligning it to the inferred action distribution. MAAD leads to significantly improved sample efficiency and stability. We demonstrate its effectiveness in a number of MuJoCo environments, both int the OpenAI Gym and the DeepMind Control Suite. We show that it requires considerable fewer interactions to achieve expert performance, outperforming current state-of-the-art on-policy methods. Remarkably, MAAD often stands out as the sole method capable of attaining expert performance levels, underscoring its simplicity and efficacy. João A. Cândido Ramos, Lionel Blondé, Naoya Takeishi, Alexandros Kalousis |
ICML | 4 |
| 2023 | Deep Grey-Box Modeling With Adaptive Data-Driven Models Toward Trustworthy Estimation of Theory-Driven ModelsabstractThe combination of deep neural nets and theory-driven models (deep grey-box models) can be advantageous due to the inherent robustness and interpretability of the theory-driven part. Deep grey-box models are usually learned with a regularized risk minimization to prevent a theory-driven part from being overwritten and ignored by a deep neural net. However, an estimation of the theory-driven part obtained by uncritically optimizing a regularizer can hardly be trustworthy if we are not sure which regularizer is suitable for the given data, which may affect the interpretability. Toward a trustworthy estimation of the theory-driven part, we should analyze the behavior of regularizers to compare different candidates and to justify a specific choice. In this paper, we present a framework that allows us to empirically analyze the behavior of a regularizer with a slight change in the architecture of the neural net and the training objective. Naoya Takeishi, Alexandros Kalousis |
AISTATS | 2 |
| 2023 | Calibrating Neural Simulation-Based Inference with Differentiable Coverage ProbabilityabstractBayesian inference allows expressing the uncertainty of posterior belief under a probabilistic model given prior information and the likelihood of the evidence. Predominantly, the likelihood function is only implicitly established by a simulator posing the need for simulation-based inference (SBI). However, the existing algorithms can yield overconfident posteriors (Hermans *et al.*, 2022) defeating the whole purpose of credibility if the uncertainty quantification is inaccurate. We propose to include a calibration term directly into the training objective of the neural model in selected amortized SBI techniques. By introducing a relaxation of the classical formulation of calibration error we enable end-to-end backpropagation. The proposed method is not tied to any particular neural model and brings moderate computational overhead compared to the profits it introduces. It is directly applicable to existing computational pipelines allowing reliable black-box posterior inference. We empirically show on six benchmark problems that the proposed method achieves competitive or better results in terms of coverage and expected posterior density than the previously existing approaches. Maciej Falkiewicz, Naoya Takeishi, Imahn Shekhzadeh, Antoine Wehenkel, Arnaud Delaunoy, Gilles Louppe, Alexandros Kalousis |
NeurIPS | 7 |
| 2022 | Graph annotation generative adversarial networks
Yoann Boget, Magda Gregorová, Alexandros Kalousis |
ACML | 3 |
| 2022 | Lipschitzness is all you need to tame off-policy generative adversarial imitation learningabstractAbstract Despite the recent success of reinforcement learning in various domains, these approaches remain, for the most part, deterringly sensitive to hyper-parameters and are often riddled with essential engineering feats allowing their success. We consider the case of off-policy generative adversarial imitation learning, and perform an in-depth review, qualitative and quantitative, of the method. We show that forcing the learned reward function to be local Lipschitz-continuous is asine qua noncondition for the method to perform well. We then study the effects of this necessary condition and provide several theoretical results involving the local Lipschitzness of the state-value function. We complement these guarantees with empirical evidence attesting to the strong positive effect that the consistent satisfaction of the Lipschitzness constraint on the reward has on imitation performance. Finally, we tackle a generic pessimistic reward preconditioning add-on spawning a large class of reward shaping methods, which makes the base method it is plugged into provably more robust, as shown in several additional theoretical guarantees. We then discuss these through a fine-grained lens and share our insights. Crucially, the guarantees derived and reported in this work are valid foranyreward satisfying the Lipschitzness condition, nothing is specific to imitation. As such, these may be of independent interest. Lionel Blondé, Pablo Strasser, Alexandros Kalousis |
Mach. Learn. | 3 |
| 2021 | Analysing the Data-Driven Approach of Dynamically Estimating Positioning AccuracyabstractThe primary expectation from positioning systems is for them to provide the users with reliable estimates of their position. An additional piece of information that can greatly help the users utilize position estimates is the level of uncertainty that a positioning system assigns to the position estimate it produced. The concept of dynamically estimating the accuracy of position estimates of fingerprinting positioning systems has been sporadically discussed over the last decade in the literature of the field, where mainly handcrafted rules based on domain knowledge have been proposed. The emergence of IoT devices and the proliferation of data from Low Power Wide Area Networks (LPWANs) have facilitated the conceptualization of data-driven methods of determining the estimated certainty over position estimates. In this work, we analyze the data-driven approach of determining the Dynamic Accuracy Estimation (DAE), considering it in the broader context of a positioning system. More specifically, with the use of a public LoRaWAN dataset, the current work analyses: the repartition of the available training set between the tasks of determining the location estimates and the DAE, the concept of selecting a subset of the most reliable estimates, and the impact that the spatial distribution of the data has to the accuracy of the DAE. The work provides a wide overview of the data-driven approach of DAE determination in the context of the overall design of a positioning system. Grigorios G. Anagnostopoulos, Alexandros Kalousis |
ICC | 2 |
| 2021 | Kanerva++: Extending the Kanerva Machine With Differentiable, Locally Block Allocated Latent Memory
Jason Ramapuram, Alexandros Kalousis |
ICLR | 3 |
| 2021 | ProxyFAUG: Proximity-based Fingerprint AugmentationabstractThe proliferation of data-demanding machine learning methods has brought to light the necessity for methodologies which can enlarge the size of training datasets, with simple, rule-based methods. In-line with this concept, the fingerprint augmentation scheme proposed in this work aims to augment fingerprint datasets which are used to train positioning models. The proposed method utilizes fingerprints which are recorded in spacial proximity, in order to perform fingerprint augmentation, creating new fingerprints which combine the features of the original ones. The proposed method of composing the new, augmented fingerprints is inspired by the crossover and mutation operators of genetic algorithms. The ProxyFAUG method aims to improve the achievable positioning accuracy of fingerprint datasets, by introducing a rule-based, stochastic, proximity-based method of fingerprint augmentation. The performance of ProxyFAUG is evaluated in an outdoor Sigfox setting using a public dataset. The best performing published positioning method on this dataset is improved by 40% in terms of median error and 6% in terms of mean error, with the use of the augmented dataset. The analysis of the results indicates a systematic and significant performance improvement at the lower error quartiles, as indicated by the impressive improvement of the median error. Grigorios G. Anagnostopoulos, Alexandros Kalousis |
IPIN | 2 |
| 2021 | Physics-Integrated Variational Autoencoders for Robust and Interpretable Generative ModelingabstractIntegrating physics models within machine learning models holds considerable promise toward learning robust models with improved interpretability and abilities to extrapolate. In this work, we focus on the integration of incomplete physics models into deep generative models. In particular, we introduce an architecture of variational autoencoders (VAEs) in which a part of the latent space is grounded by physics. A key technical challenge is to strike a balance between the incomplete physics and trainable components such as neural networks for ensuring that the physics part is used in a meaningful manner. To this end, we propose a regularized learning method that controls the effect of the trainable components and preserves the semantics of the physics-based latent variables as intended. We not only demonstrate generative performance improvements over a set of synthetic and real-world datasets, but we also show that we learn robust models that can consistently extrapolate beyond the training distribution in a meaningful manner. Moreover, we show that we can control the generative process in an interpretable manner. Naoya Takeishi, Alexandros Kalousis |
NeurIPS | 2 |
| 2021 | Conditional Neural Relational Inference for Interacting Systems
João A. Cândido Ramos, Lionel Blondé, Stéphane Armand, Alexandros Kalousis |
ECML/PKDD (5) | 4 |
| 2020 | Improving VAE Generations of Multimodal Data Through Data-Dependent Conditional PriorsabstractOne of the major shortcomings of variational autoencoders is the inability to produce generations from the individual modalities of data originating from mixture distributions. This is primarily due to the use of a simple isotropic Gaussian as the prior for the latent code in the ancestral sampling procedure for the data generations. We propose a novel formulation of variational autoencoders, conditional prior VAE (CP-VAE), which learns to differentiate between the individual mixture components and therefore allows for generations from the distributional data clusters. We assume a two-level generative process with a continuous (Gaussian) latent variable sampled conditionally on a discrete (categorical) latent component. The new variational objective naturally couples the learning of the posterior and prior conditionals, and the learning of the latent categories encoding the multimodality of the original data in an unsupervised manner. The data-dependent conditional priors are then used to sample the continuous latent code when generating new samples from the individual mixture components corresponding to the multimodal structure of the original data. Our experimental results illustrate the generative performance of our new model comparing to multiple baselines. Frantzeska Lavda, Magda Gregorová, Alexandros Kalousis |
ECAI | 3 |
| 2020 | Hyperbolic Knowledge Graph Embeddings for Knowledge Base Completion
Prodromos Kolyvakis, Alexandros Kalousis, Dimitris Kiritsis |
ESWC | 2 |
| 2020 | Goal-directed Generation of Discrete Structures with Conditional Generative ModelsabstractDespite recent advances, goal-directed generation of structured discrete data remains challenging. For problems such as program synthesis (generating source code) and materials design (generating molecules), finding examples which satisfy desired constraints or exhibit desired properties is difficult. In practice, expensive heuristic search or reinforcement learning algorithms are often employed. In this paper, we investigate the use of conditional generative models which directly attack this inverse problem, by modeling the distribution of discrete structures given properties of interest. Unfortunately, the maximum likelihood training of such models often fails with the samples from the generative model inadequately respecting the input properties. To address this, we introduce a novel approach to directly optimize a reinforcement learning objective, maximizing an expected reward. We avoid high-variance score-function estimators that would otherwise be required by sampling from an approximation to the normalized rewards, allowing simple Monte Carlo estimation of model gradients. We test our methodology on two tasks: generating molecules with user-defined properties and identifying short python expressions which evaluate to a given target value. In both cases, we find improvements over maximum likelihood estimation and other baselines. Amina Mollaysa, Brooks Paige, Alexandros Kalousis |
NeurIPS | 3 |
| 2020 | Lifelong generative modeling
Jason Ramapuram, Magda Gregorová, Alexandros Kalousis |
Neurocomputing | 3 |
| 2019 | Learning to Augment with Feature Side-informationabstractNeural networks typically need huge amounts of data to train in order to get reasonable generalizable results. A common approach is to artificially generate samples by using prior knowledge of the data properties or other relevant domain knowledge. However, if the assumptions on the data properties are not accurate or the domain knowledge is irrelevant to the task at hand, one may end up degenerating learning performance by using such augmented data in comparison to simply training on the limited available dataset. We propose a critical data augmentation method using feature side-information, which is obtained from domain knowledge and provides detailed information about features' intrinsic properties. Most importantly, we introduce an instance wise quality checking procedure on the augmented data. It filters out irrelevant or harmful augmented data prior to entering the model. We validated this approach on both synthetic and real-world datasets, specifically in a scenario where the data augmentation is done based on a task independent, unreliable source of information. The experiments show that the introduced critical data augmentation scheme helps avoid performance degeneration resulting from incorporating wrong augmented data. Amina Mollaysa, Alexandros Kalousis, Eric Bruno, Maurits Diephuis |
ACML | 2 |
| 2019 | Sample-Efficient Imitation Learning via Generative Adversarial NetsabstractGAIL is a recent successful imitation learning architecture that exploits the adversarial training procedure introduced in GANs. Albeit successful at generating behaviours similar to those demonstrated to the agent, GAIL suffers from a high sample complexity in the number of interactions it has to carry out in the environment in order to achieve satisfactory performance. We dramatically shrink the amount of interactions with the environment necessary to learn well-behaved imitation policies, by up to several orders of magnitude. Our framework, operating in the model-free regime, exhibits a significant increase in sample-efficiency over previous methods by simultaneously a) learning a self-tuned adversarially-trained surrogate reward and b) leveraging an off-policy actor-critic architecture. We show that our approach is simple to implement and that the learned agents remain remarkably stable, as shown in our experiments that span a variety of continuous control tasks. Video visualisations available at: \url{https://youtu.be/-nCsqUJnRKU}. Lionel Blondé, Alexandros Kalousis |
AISTATS | 2 |
| 2019 | Variational Saccading: Efficient Inference for Large Resolution Images
Jason Ramapuram, Maurits Diephuis, Frantzeska Lavda, Russell Webb, Alexandros Kalousis |
BMVC | 5 |
| 2019 | A Reproducible Analysis of RSSI Fingerprinting for Outdoor Localization Using Sigfox: Preprocessing and Hyperparameter TuningabstractFingerprinting techniques, which are a common method for indoor localization, have been recently applied with success into outdoor settings. Particularly, the communication signals of Low Power Wide Area Networks (LPWAN) such as Sigfox, have been used for localization. In this rather recent field of study, not many publicly available datasets, which would facilitate the consistent comparison of different positioning systems, exist so far. In the current study, a published dataset of RSSI measurements on a Sigfox network deployed in Antwerp, Belgium is used to analyse the appropriate selection of preprocessing steps and to tune the hyperparameters of a kNN fingerprinting method. Initially, the tuning of hyperparameter k for a variety of distance metrics, and the selection of efficient data transformation schemes, proposed by relevant works, is presented. In addition, accuracy improvements are achieved in this study, by a detailed examination of the appropriate adjustment of the parameters of the data transformation schemes tested, and of the handling of out of range values. With the appropriate tuning of these factors, the achieved mean localization error was 298 meters, and the median error was 109 meters. To facilitate the reproducibility of tests and comparability of results, the code and train/validation/test split used in this study are available. Grigorios G. Anagnostopoulos, Alexandros Kalousis |
IPIN | 2 |
| 2018 | DeepAlignment: Unsupervised Ontology Matching with Refined Word VectorsabstractProdromos Kolyvakis, Alexandros Kalousis, Dimitris Kiritsis. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Prodromos Kolyvakis, Alexandros Kalousis, Dimitris Kiritsis |
NAACL-HLT | 2 |
| 2018 | Large-Scale Nonlinear Variable Selection via Kernel Random Features
Magda Gregorová, Jason Ramapuram, Alexandros Kalousis, Stéphane Marchand-Maillet |
ECML/PKDD (2) | 3 |
| 2018 | Structured nonlinear variable selection
Magda Gregorová, Alexandros Kalousis, Stéphane Marchand-Maillet |
UAI | 2 |
| 2017 | Learning Predictive Leading Indicators for Forecasting Time Series Systems with Unknown Clusters of Forecast TasksabstractWe present a new method for forecasting systems of multiple interrelated time series. The method learns the forecast models together with discovering leading indicators from within the system that serve as good predictors improving the forecast accuracy and a cluster structure of the predictive tasks around these. The method is based on the classical linear vector autoregressive model (VAR) and links the discovery of the leading indicators to inferring sparse graphs of Granger causality. We formulate a new constrained optimisation problem to promote the desired sparse structures across the models and the sharing of information amongst the learning tasks in a multi-task manner. We propose an algorithm for solving the problem and document on a battery of synthetic and real-data experiments the advantages of our new method over baseline VAR models as well as the state-of-the-art sparse VAR learning methods. Magda Gregorová, Alexandros Kalousis, Stéphane Marchand-Maillet |
ACML | 2 |
| 2017 | Semi Supervised Relevance Learning for Feature Selection on High Dimensional DataabstractNowadays, the advanced technologies make amounts of data growing in a fast paced way. In many application fields, this trend concerns specially dimensions of the data. It is the case where features are about thousands and tens of thousands, while the number of instances is much smaller. This phenomenon is known as the curse of dimensionality and it results in modest classification performance and feature selection instability. In order to deal with this issue, we propose a new feature selection approach that makes use of background knowledge about some dimensions known to be more relevant, as a means of directing the feature selection process. In this approach, prior knowledge about some features is used to learn new relevant features by a semi supervised approach. Experiments on three high dimensional data sets show promising results on both classification performance and stability of feature selection. Afef Ben Brahim, Alexandros Kalousis |
AICCSA | 2 |
| 2017 | Regularising Non-linear Models Using Feature Side-informationabstractVery often features come with their own vectorial descriptions which provide detailed information about their properties. We refer to these vectorial descriptions as feature side-information. In the standard learning scenario, input is represented as a vector of features and the feature side-information is most often ignored or used only for feature selection prior to model fitting. We believe that feature side-information which carries information about features intrinsic property will help improve model prediction if used in a proper way during learning process. In this paper, we propose a framework that allows for the incorporation of the feature side-information during the learning of very general model families to improve the prediction performance. We control the structures of the learned models so that they reflect features’ similarities as these are defined on the basis of the side-information. We perform experiments on a number of benchmark datasets which show significant predictive performance gains, over a number of baselines, as a result of the exploitation of the side-information. Amina Mollaysa, Pablo Strasser, Alexandros Kalousis |
ICML | 3 |
| 2017 | Forecasting and Granger Modelling with Non-linear Dynamical Dependencies
Magda Gregorová, Alexandros Kalousis, Stéphane Marchand-Maillet |
ECML/PKDD (2) | 2 |
| 2016 | Factorizing LambdaMART for cold start recommendations
Phong Nguyen 0002, Jun Wang 0017, Alexandros Kalousis |
Mach. Learn. | 3 |
| 2015 | Information Geometry and Minimum Description Length NetworksabstractWe study parametric unsupervised mixture learning. We measure the loss of intrinsic information from the observations to complex mixture models, and then to simple mixture models. We present a geometric picture, where all these representations are regarded as free points in the space of probability distributions. Based on minimum description length, we derive a simple geometric principle to learn all these models together. We present a new learning machine with theories, algorithms, and simulations. Ke Sun 0001, Jun Wang 0017, Alexandros Kalousis, Stéphane Marchand-Maillet |
ICML | 3 |
| 2015 | Space-Time Local EmbeddingsabstractSpace-time is a profound concept in physics. This concept was shown to be useful for dimensionality reduction. We present basic definitions with interesting counter-intuitions. We give theoretical propositions to show that space-time is a more powerful representation than Euclidean space. We apply this concept to manifold learning for preserving local information. Empirical results on non-metric datasets show that more information can be preserved in space-time. Ke Sun 0001, Jun Wang 0017, Alexandros Kalousis, Stéphane Marchand-Maillet |
NIPS | 3 |
| 2015 | The Data Mining OPtimization Ontology
C. Maria Keet, Agnieszka Lawrynowicz, Claudia d'Amato, Alexandros Kalousis, Phong Nguyen 0002, Raúl Palma, Robert Stevens 0001, Melanie Hilario |
J. Web Semant. | 4 |
| 2014 | Two-Stage Metric LearningabstractIn this paper, we present a novel two-stage metric learning algorithm. We first map each learning instance to a probability distribution by computing its similarities to a set of fixed anchor points. Then, we define the distance in the input data space as the Fisher information distance on the associated statistical manifold. This induces in the input data space a new family of distance metric which presents unique properties. Unlike kernelized metric learning, we do not require the similarity measure to be positive semi-definite. Moreover, it can also be interpreted as a local metric learning algorithm with well defined distance approximation. We evaluate its performance on a number of datasets. It outperforms significantly other metric learning methods and SVM. Jun Wang 0017, Ke Sun 0001, Fei Sha, Stéphane Marchand-Maillet, Alexandros Kalousis |
ICML | 5 |
| 2014 | Using Meta-mining to Support Data Mining Workflow Planning and OptimizationabstractKnowledge Discovery in Databases is a complex process that involves many different data processing and learning operators. Today's Knowledge Discovery Support Systems can contain several hundred operators. A major challenge is to assist the user in designing workflows which are not only valid but also -- ideally -- optimize some performance measure associated with the user goal. In this paper we present such a system. The system relies on a meta-mining module which analyses past data mining experiments and extracts meta-mining models which associate dataset characteristics with workflow descriptors in view of workflow performance optimization. The meta-mining model is used within a data mining workflow planner, to guide the planner during the workflow planning. We learn the meta-mining models using a similarity learning approach, and extract the workflow descriptors by mining the workflows for generalized relational patterns accounting also for domain knowledge provided by a data mining ontology. We evaluate the quality of the data mining workflows that the system produces on a collection of real world datasets coming from biology and show that it produces workflows that are significantly better than alternative methods that can only do workflow selection and not planning. P. Nguyen, Melanie Hilario, Alexandros Kalousis |
J. Artif. Intell. Res. | 3 |
| 2013 | Convex formulations of radius-margin based Support Vector MachinesabstractWe consider Support Vector Machines (SVMs) learned together with linear transformations of the feature spaces on which they are applied. Under this scenario the radius of the smallest data enclosing sphere is no longer fixed. Therefore optimizing the SVM error bound by considering both the radius and the margin has the potential to deliver a tighter error bound. In this paper we present two novel algorithms: R-SVM_μ^+—a SVM radius-margin based feature selection algorithm, and R-SVM^+ — a metric learning-based SVM. We derive our algorithms by exploiting a new tighter approximation of the radius and a metric learning interpretation of SVM. Both optimize directly the radius-margin error bound using linear transformations. Unlike almost all existing radius-margin based SVM algorithms which are either non-convex or combinatorial, our algorithms are standard quadratic convex optimization problems with linear or quadratic constraints. We perform a number of experiments on benchmark datasets. R-SVM_μ^+ exhibits excellent feature selection performance compared to the state-of-the-art feature selection methods, such as L_1-norm and elastic-net based methods. R-SVM^+ achieves a significantly better classification performance compared to SVM and its other state-of-the-art variants. From the results it is clear that the incorporation of the radius, as a means to control the data spread, in the cost function has strong beneficial effects. Huyen Do, Alexandros Kalousis |
ICML (1) | 2 |
| 2012 | Learning Heterogeneous Similarity Measures for Hybrid-Recommendations in Meta-MiningabstractThe notion of meta-mining has appeared recently and extends traditional meta-learning in two ways. First it provides support for the whole data-mining process. Second it pries open the so called algorithm black-box approach where algorithms and workflows also have descriptors. With the availability of descriptors both for datasets and data-mining workflows we are faced with a problem the nature of which is much more similar to those appearing in recommendation systems. In order to account for the meta-mining specificities we derive a novel metric-based-learning recommender approach. Our method learns two homogeneous metrics, one in the dataset and one in the workflow space, and a heterogeneous one in the dataset-workflow space. All learned metrics reflect similarities established from the dataset-workflow preference matrix. The latter is constructed from the performance results obtained by the application of workflows to datasets. We demonstrate our method on meta-mining over biological (microarray datasets) problems. The application of our method is not limited to the meta-mining problem, its formulation is general enough so that it can be applied on problems with similar requirements. Phong Nguyen 0002, Jun Wang 0017, Melanie Hilario, Alexandros Kalousis |
ICDM | 4 |
| 2012 | Model mining for robust feature selectionabstractA common problem with most of the feature selection methods is that they often produce feature sets--models--that are not stable with respect to slight variations in the training data. Different authors tried to improve the feature selection stability using ensemble methods which aggregate different feature sets into a single model. However, the existing ensemble feature selection methods suffer from two main shortcomings: (i) the aggregation treats the features independently and does not account for their interactions, and (ii) a single feature set is returned, nevertheless, in various applications there might be more than one feature sets, potentially redundant, with similar information content. In this work we address these two limitations. We present a general framework in which we mine over different feature models produced from a given dataset in order to extract patterns over the models. We use these patterns to derive more complex feature model aggregation strategies that account for feature interactions, and identify core and distinct feature models. We conduct an extensive experimental evaluation of the proposed framework where we demonstrate its effectiveness over a number of high-dimensional problems from the fields of biology and text-mining. Adam Woznica, Phong Nguyen 0002, Alexandros Kalousis |
KDD | 3 |
| 2012 | Parametric Local Metric Learning for Nearest Neighbor ClassificationabstractWe study the problem of learning local metrics for nearest neighbor classification. Most previous works on local metric learning learn a number of local unrelated metrics. While this ''independence'' approach delivers an increased flexibility its downside is the considerable risk of overfitting. We present a new parametric local metric learning method in which we learn a smooth metric matrix function over the data manifold. Using an approximation error bound of the metric matrix function we learn local metrics as linear combinations of basis metrics defined on anchor points over different regions of the instance space. We constrain the metric matrix function by imposing on the linear combinations manifold regularization which makes the learned metric matrix function vary smoothly along the geodesics of the data manifold. Our metric learning method has excellent performance both in terms of predictive power and scalability. We experimented with several large-scale classification problems, tens of thousands of instances, and compared it with several state of the art metric learning methods, both global and local, as well as to SVM with automatic kernel selection, all of which it outperforms in a significant manner. Jun Wang 0017, Alexandros Kalousis, Adam Woznica |
NIPS | 2 |
| 2012 | Learning Neighborhoods for Metric Learning
Jun Wang 0017, Adam Woznica, Alexandros Kalousis |
ECML/PKDD (1) | 3 |
| 2011 | Metric Learning with Multiple KernelsabstractMetric learning has become a very active research field. The most popular representative--Mahalanobis metric learning--can be seen as learning a linear transformation and then computing the Euclidean metric in the transformed space. Since a linear transformation might not always be appropriate for a given learning problem, kernelized versions of various metric learning algorithms exist. However, the problem then becomes finding the appropriate kernel function. Multiple kernel learning addresses this limitation by learning a linear combination of a number of predefined kernels; this approach can be also readily used in the context of multiple-source learning to fuse different data sources. Surprisingly, and despite the extensive work on multiple kernel learning for SVMs, there has been no work in the area of metric learning with multiple kernel learning. In this paper we fill this gap and present a general approach for metric learning with multiple kernel learning. Our approach can be instantiated with different metric learning algorithms provided that they satisfy some constraints. Experimental evidence suggests that our approach outperforms metric learning with an unweighted kernel combination and metric learning with cross-validation based kernel selection. Jun Wang 0017, Huyen Do, Adam Woznica, Alexandros Kalousis |
NIPS | 4 |
| 2010 | Adaptive Distances on Sets of VectorsabstractRecently, there has been a growing interest in learning distances directly from training data. While the previous works focused mainly on adapting distance measures over vectorial data, it is a well-known fact that many real-world data could not be easily represented as fixed length tuples of constants. In this paper we address this limitation and propose a novel class of distance learning techniques for learning problems in which instances are set of vectors, examples of such problems include, among others, automatic image annotation and graph classification. We investigate the behavior of the adaptive set distances on a number of artificial and real-world problems and demonstrate that they improve over the standard set distances. Adam Woznica, Alexandros Kalousis |
ICDM | 2 |
| 2010 | A New Framework for Dissimilarity and Similarity Learning
Adam Woznica, Alexandros Kalousis |
PAKDD (2) | 2 |
| 2010 | Adaptive Matching Based Kernels for Labelled Graphs
Adam Woznica, Alexandros Kalousis, Melanie Hilario |
PAKDD (2) | 2 |
| 2010 | Addressing the Challenge of Defining Valid Proteomic Biomarkers and ClassifiersabstractBACKGROUND: The purpose of this manuscript is to provide, based on an extensive analysis of a proteomic data set, suggestions for proper statistical analysis for the discovery of sets of clinically relevant biomarkers. As tractable example we define the measurable proteomic differences between apparently healthy adult males and females. We choose urine as body-fluid of interest and CE-MS, a thoroughly validated platform technology, allowing for routine analysis of a large number of samples. The second urine of the morning was collected from apparently healthy male and female volunteers (aged 21-40) in the course of the routine medical check-up before recruitment at the Hannover Medical School. RESULTS: We found that the Wilcoxon-test is best suited for the definition of potential biomarkers. Adjustment for multiple testing is necessary. Sample size estimation can be performed based on a small number of observations via resampling from pilot data. Machine learning algorithms appear ideally suited to generate classifiers. Assessment of any results in an independent test-set is essential. CONCLUSIONS: Valid proteomic biomarkers for diagnosis and prognosis only can be defined by applying proper statistical data mining procedures. In particular, a justification of the sample size should be part of the study design. Mohammed Dakna, Keith Harris, Alexandros Kalousis, Sebastien Carpentier, Walter Kolch, Joost Schanstra, Marion Haubitz, Antonia Vlahou, Harald Mischak, Mark A. Girolami |
BMC Bioinform. | 3 |
| 2009 | Feature Weighting Using Margin and Radius Based Error Bound Optimization in SVMs
Huyen Do, Alexandros Kalousis, Melanie Hilario |
ECML/PKDD (1) | 2 |
| 2009 | Margin and Radius Based Multiple Kernel Learning
Huyen Do, Alexandros Kalousis, Adam Woznica, Melanie Hilario |
ECML/PKDD (1) | 2 |
| 2008 | A general framework for estimating similarity of datasets and decision trees: exploring semantic similarity of decision treesabstractDecision trees are among the most popular pattern types in data mining due to their intuitive representation. However, little attention has been given on the definition of measures of semantic similarity between decision trees. In this work, we present a general framework for similarity estimation that includes as special cases the estimation of semantic similarity between decision trees, as well as various forms of similarity estimation on classification datasets with respect to different probability distributions defined over the attribute-class space of the datasets. The similarity estimation is based on the partitions induced by the decision trees on the attribute space of the datasets. We use the proposed framework in order to estimate the semantic similarity of decision trees induced from different subsamples of classification datasets; we evaluate its performance with respect to the empirical semantic similarity, which we estimate on the basis of independent hold-out test sets. The availability of similarity measures on decision trees opens a wide range of possibilities for meta-analysis and meta-mining of the data mining results. Eirini Ntoutsi, Alexandros Kalousis, Yannis Theodoridis |
SDM | 2 |
| 2008 | Feature Selection with the logRatio KernelabstractIn this article we present a novel kernel function, logRatio, which was designed to address two common problems in biological applications: data preprocessing and attribute interaction modelling. An extension of the SVMRFE feature selection algorithm was built around this new kernel function and compared with the original on a number of biological data and text classification problems. Experiments showed that SVMRFE based on the logRatio kernel detects relevant information and handles attribute redundancy more effectively than SVMRFE coupled with other kernels. Julien Prados, Alexandros Kalousis, Melanie Hilario |
SDM | 2 |
| 2008 | Approaches to dimensionality reduction in proteomic biomarker studiesabstractMass-spectra based proteomic profiles have received widespread attention as potential tools for biomarker discovery and early disease diagnosis. A major data-analytical problem involved is the extremely high dimensionality (i.e. number of features or variables) of proteomic data, in particular when the sample size is small. This article reviews dimensionality reduction methods that have been used in proteomic biomarker studies. It then focuses on the problem of selecting the most appropriate method for a specific task or dataset, and proposes method combination as a potential alternative to single-method selection. Finally, it points out the potential of novel dimension reduction techniques, in particular those that incorporate domain knowledge through the use of informative priors or causal inference. Melanie Hilario, Alexandros Kalousis |
Briefings Bioinform. | 2 |
| 2007 | Learning to combine distances for complex representationsabstractThe k-Nearest Neighbors algorithm can be easily adapted to classify complex objects (e.g. sets, graphs) as long as a proper dissimilarity function is given over an input space. Both the representation of the learning instances and the dissimilarity employed on that representation should be determined on the basis of domain knowledge. However, even in the presence of domain knowledge, it can be far from obvious which complex representation should be used or which dissimilarity should be applied on the chosen representation. In this paper we present a framework that allows to combine different complex representations of a given learning problem and/or different dissimilarities defined on these representations. We build on ideas developed previously on metric learning for vectorial data. We demonstrate the utility of our method in domains in which the learning instances are represented as sets of vectors by learning how to combine different set distance measures. Adam Woznica, Alexandros Kalousis, Melanie Hilario |
ICML | 2 |
| 2007 | Stability of feature selection algorithms: a study on high-dimensional spaces
Alexandros Kalousis, Julien Prados, Melanie Hilario |
Knowl. Inf. Syst. | 1 |
| 2006 | On Preprocessing of SELDI-MS Data and its EvaluationabstractMass spectrometry is becoming an important tool in proteomics. Mass spectral data are characterized by very high dimensionality and a high level of redundancy. Both issues are quite challenging when one wants to perform knowledge discovery and push existing tools to their limits. We tackle both via a preprocessing pipeline that drastically reduces dimensionality and redundancy of the initial representation in order to focus on biologically relevant information. Essentially preprocessing performs feature extraction in a manner that reflects domain knowledge. We propose a framework for the evaluation of the given pipeline and in fact of any mass spectrometry preprocessing pipeline which is based on the level of conservation of discriminatory information. The discriminatory information content of a given representation is objectively measured by the classification performance of a number of classification algorithms evaluated on the given representation. This approach also allows us to compare a number of different preprocessing possibilities, namely using peak intensities vs peak areas to represent peaks and how non observed peaks should be treated, and demonstrate which is the most informative one Julien Prados, Alexandros Kalousis, Melanie Hilario |
CBMS | 2 |
| 2006 | Distances and (Indefinite) Kernels for Sets of ObjectsabstractThe main disadvantage of most existing set kernels is that they are based on averaging, which might be inappropriate for problems where only specific elements of the two sets should determine the overall similarity. In this paper we propose a class of kernels for sets of vectors directly exploiting set distance measures and, hence, incorporating various semantics into set kernels and lending the power of regularization to learning in structural domains where natural distance functions exist. These kernels belong to two groups: (i) kernels in the proximity space induced by set distances and (ii) set distance substitution kernels (non-PSD in general). We report experimental results which show that our kernels compare favorably with kernels based on averaging and achieve results similar to other state-of-the-art methods. At the same time our kernels systematically improve over the naive way of exploiting distances. Adam Woznica, Alexandros Kalousis, Melanie Hilario |
ICDM | 2 |
| 2006 | Kernels on Lists and Sets over Relational Algebra: An Application to Classification of Protein Fingerprints
Adam Woznica, Alexandros Kalousis, Melanie Hilario |
PAKDD | 2 |
| 2005 | Stability of Feature Selection AlgorithmsabstractWith the proliferation of extremely high-dimensional data, feature selection algorithms have become indispensable components of the learning process. Strangely, despite extensive work on the stability of learning algorithms, the stability of feature selection algorithms has been relatively neglected. This study is an attempt to fill that gap by quantifying the sensitivity of feature selection algorithms to variations in the training set. We assess the stability of feature selection algorithms based on the stability of the feature preferences that they express in the form of weights-scores, ranks, or a selected feature subset. We examine a number of measures to quantify the stability of feature preferences and propose an empirical way to estimate them. We perform a series of experiments with several feature selection algorithms on a set of proteomics datasets. The experiments allow us to explore the merits of each stability measure and create stability profiles of the feature selection algorithms. Finally we show how stability profiles can support the choice of a feature selection algorithm. Alexandros Kalousis, Julien Prados, Melanie Hilario |
ICDM | 1 |
| 2005 | Kernels over Relational Algebra Structures
Adam Woznica, Alexandros Kalousis, Melanie Hilario |
PAKDD | 2 |
| 2005 | Feature Extraction from Mass Spectra for Classification of Pathological States
Alexandros Kalousis, Julien Prados, Elton Rexhepaj, Melanie Hilario |
PKDD | 1 |
| 2004 | Distilling Classification Models from Cross Validation Runs: An Application to Mass SpectrometryabstractWe present work on a proteomics application. More specifically, from the domain of mass-spectrometry and the identification of biomarkers for stroke attacks. Mass-spectrometry based biomarker identification is an application that sets a number of challenges to the knowledge discovery process. We describe how we tackle them and present a number of machine learning experiments that we performed in order to identify the most suitable learning algorithm for the given problem. However working with real world applications one of the main issues apart from good classification performance is an indication of the factors that really determine the classification decision. Usually based on the results of a resampled-based performance estimation, e.g. cross validation, an algorithm is selected that will provide the operational classification model. On a next step the operational model should be constructed, nevertheless it is not obvious how this should be done since in resampled-based procedures a number of different models are created. We propose a method for linear classifiers that examines the different models produced with cross-validation. The method examines the stability of the models produced from the different training folds and combines them to provide a single model. Alexandros Kalousis, Julien Prados, Jean-Charles Sanchez, Laure Allard, Melanie Hilario |
ICTAI | 1 |
| 2004 | On Data and Algorithms: Understanding Inductive Performance
Alexandros Kalousis, João Gama 0001, Melanie Hilario |
Mach. Learn. | 1 |
| 2003 | Representational Issues in Meta-Learning
Alexandros Kalousis, Melanie Hilario |
ICML | 1 |
| 2001 | Estimating the Predictive Accuracy of a Classifier
Hilan Bensusan, Alexandros Kalousis |
ECML | 2 |
| 2001 | Feature Selection for Meta-learning
Alexandros Kalousis, Melanie Hilario |
PAKDD | 1 |
| 2001 | Fusion of Meta-knowledge and Meta-data for Case-Based Model Selection
Melanie Hilario, Alexandros Kalousis |
PKDD | 2 |
| 2000 | Model selection via meta-learning: a comparative studyabstractThe selection of an appropriate inducer is crucial for performing effective classification. In previous work we presented a system called NOEMON which relied on a mapping between dataset characteristics and inducer performance to propose inducers for specific datasets. Instance based learning was used to create that mapping. Here we extend and refine the set of data characteristics; we also use a wider range of base-level inducers and a much larger collection of datasets to create the meta-models. We compare the performance of meta-models produced by instance based learners, decision trees and boosted decision trees. The results show that decision trees and boosted decision trees models enhance the perfomance of the system. Alexandros Kalousis, Melanie Hilario |
ICTAI | 1 |
| 2000 | Quantifying the Resilience of Inductive Classification Algorithms
Melanie Hilario, Alexandros Kalousis |
PKDD | 2 |
| 1999 | NOEMON: Design, implementation and performance results of an intelligent assistant for classifier selectionabstractThe selection of an appropriate classification model and algorithm is crucial for effective knowledge discovery on a dataset. For large databases, common in data mining, such a selection is necessary, because the cost of invoking all alternative classifiers is prohibitive. This selection task is impeded by two factors. First, there are many performance criteria, and the behaviour of a classifier varies considerably with them. Second, a classifier's performance is strongly affected by the characteristics of the dataset. Classifier selection implies mastering a lot of background information on the dataset, the models and the algorithms in question. An intelligent assistant can reduce this effort by inducing helpful suggestions from background information. In this study, we present such an assistant, NOEMON. For each registered classifier, NOEMON measures its performance for a collection of datasets. Rules are induced from those measurements and accommodated in a knowledge base. The suggestion on the most appropriate classifier(s) for a dataset is then based on those rules. Results on the performance of an initial prototype are also given. Alexandros Kalousis, Theoharis Theoharis |
Intell. Data Anal. | 1 |