VLDB 2026 Research / reviewers in the wild / expert
Pavlo Mozharovskyi
dblp:117/3535
· DBLP profile ↗
12ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-1925-3337ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Two Steps Textual Anomaly Detection through Anisotropy MitigationabstractAnomaly detection aims at distinguishing between in-distribution samples, which belong to the same distribution as the training set, and out-of-distribution samples, which lie outside of it.In textual anomaly detection, recent approaches routinely apply anomaly detection algorithms directly to embeddings extracted from pre-trained embedding models (two-stage approaches).However, the geometric properties of pre-trained embeddings can hinder the effectiveness of detection algorithms, which often rely on distance-based measures.In this work, we first highlight the relevance of similaritytrained models for textual anomaly detection.Beyond being trained to capture semantic similarities, these models also exhibit geometric properties that appear better suited to detection algorithms.We further demonstrate that, besides model choice, a simple post-processing step can significantly improve anomaly detection by adapting embeddings to the assumptions made by classical detection algorithms.The bulk of our experiments is done on a reformulation of the classification tasks from the MTEB benchmark into anomaly detection tasks 1 . Pierre Fihey, Matthieu Labeau, Pavlo Mozharovskyi |
ACL (1) | 3 |
| 2026 | Unified pipeline for generalized mental state detection using EEG signalsabstractMental states, a complex union of cognitive, emotional, and perceptual conditions, fundamentally shape how individuals perceive and interact with their surroundings. Detecting these states is vital, as it reveals the underlying processes that govern behaviour and enables targeted interventions across diverse fields such as mental health, education, and human-computer interaction. Generalisability across subjects and trials is essential to ensure that these interventions are effective and reliable in varied real-world settings, thereby enhancing their practical applicability. In this paper, we introduce an end-to-end optimised pipeline for classifying mental states from electroencephalography (EEG) signals. Through quantitative studies of data preprocessing and feature enhancement of continuous data collected under less stringent conditions, our pipeline utilises specially designed, cutting-edge, lightweight classifiers and achieves new state-of-the-art performance. Specifically addressing the challenge of generalisability in EEG signal research, our pipeline demonstrates robust performance, achieving a peak accuracy of 79.1% and an average of 71.9% in cross-subject scenarios, and a high of 89.3% with an average of 85.4% in cross-trial evaluations. Yinghao Wang, Rayan Elrawas, Anh-Dung Nguyen, Maxime Girard, Pavlo Mozharovskyi, Enzo Tartaglione |
Expert Syst. Appl. | 5 |
| 2025 | Tailoring Mixup to Data for CalibrationabstractAmong all data augmentation techniques proposed so far, linear interpolation of training samples, also called Mixup, has found to be effective for a large panel of applications.
Along with improved predictive performance, Mixup is also a good technique for improving calibration.
However, mixing data carelessly can lead to manifold mismatch, i.e., synthetic data lying outside original class manifolds, which can deteriorate calibration.
In this work, we show that the likelihood of assigning a wrong label with mixup increases with the distance between data to mix.
To this end, we propose to dynamically change the underlying distributions of interpolation coefficients
depending on the similarity between samples to mix, and define a flexible framework to do so without losing in diversity. We provide extensive experiments for classification and regression tasks, showing that our proposed method improves predictive performance
and calibration of models, while being much more efficient. Quentin Bouniot, Pavlo Mozharovskyi, Florence d'Alché-Buc |
ICLR | 2 |
| 2025 | Restyling Unsupervised Concept Based Interpretable Networks with Generative ModelsabstractDeveloping inherently interpretable models for prediction has gained prominence in recent years. A subclass of these models, wherein the interpretable network relies on learning high-level concepts, are valued because of closeness of concept representations to human communication. However, the visualization and understanding of the learnt unsupervised dictionary of concepts encounters major limitations, especially for large-scale images. We propose here a novel method that relies on mapping the concept features to the latent space of a pretrained generative model. The use of a generative model enables high quality visualization, and lays out an intuitive and interactive procedure for better interpretation of the learnt concepts by imputing concept activations and visualizing generated modifications. Furthermore, leveraging pretrained generative models has the additional advantage of making the training of the system more efficient. We quantitatively ascertain the efficacy of our method in terms of accuracy of the interpretable prediction network, fidelity of reconstruction, as well as faithfulness and consistency of learnt concepts. The experiments are conducted on multiple image recognition benchmarks for large-scale images. Project page available at https://jayneelparekh.github.io/VisCoIN_project_page/ Jayneel Parekh, Quentin Bouniot, Pavlo Mozharovskyi, Alasdair Newson, Florence d'Alché-Buc |
ICLR | 3 |
| 2025 | Self-Supervised Learning of Graph Representations for Network Intrusion DetectionabstractDetecting intrusions in network traffic is a challenging task, particularly under limited supervision and constantly evolving attack patterns. While recent works have leveraged graph neural networks for network intrusion detection, they often decouple representation learning from anomaly detection, limiting the utility of the embeddings for identifying attacks. We propose GraphIDS, a self-supervised intrusion detection model that unifies these two stages by learning local graph representations of normal communication patterns through a masked autoencoder. An inductive graph neural network embeds each flow with its local topological context to capture typical network behavior, while a Transformer‑based encoder-decoder reconstructs these embeddings, implicitly learning global co-occurrence patterns via self-attention without requiring explicit positional information. During inference, flows with unusually high reconstruction errors are flagged as potential intrusions. This end-to-end framework ensures that embeddings are directly optimized for the downstream task, facilitating the recognition of malicious traffic. On diverse NetFlow benchmarks, GraphIDS achieves up to 99.98% PR‑AUC and 99.61% macro F1-score, outperforming baselines by 5–25 percentage points. Lorenzo Guerra, Thomas Chapuis, Guillaume Duc, Pavlo Mozharovskyi |
NeurIPS | 4 |
| 2024 | Tackling Interpretability in Audio Classification Networks With Non-negative Matrix FactorizationabstractThis paper tackles two major problem settings for interpretability of audio processing networks,post-hocandby-designinterpretation. For post-hoc interpretation, we aim to interpret decisions of a network in terms of high-level audio objects that are also listenable for the end-user. This is extended to present an inherently interpretable model with high performance. To this end, we propose a novel interpreter design that incorporates non-negative matrix factorization (NMF). In particular, an interpreter is trained to generate a regularized intermediate embedding from hidden layers of a target network, learnt as time-activations of a pre-learnt NMF dictionary. Our methodology allows us to generate intuitive audio-based interpretations that explicitly enhance parts of the input signal most relevant for a network's decision. We demonstrate our method's applicability on a variety of classification tasks, including multi-label data for real-world audio and music. Jayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Gaël Richard, Florence d'Alché-Buc |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Statistical Depth Functions for Ranking Distributions: Definitions, Statistical Learning and ApplicationsabstractThe concept of median/consensus has been widely investigated in order to provide a statistical summary of ranking data, i.e. realizations of a random permutation $\Sigma$ of a finite set, $\{1,; \ldots,;{n}\}$ with $n\geq 1$ say. As it sheds light onto only one aspect of $\Sigma$’s distribution $P$, it may neglect other informative features. It is the purpose of this paper to define analogues of quantiles, ranks and statistical procedures based on such quantities for the analysis of ranking data by means of a metric-based notion of depth function on the symmetric group. Overcoming the absence of vector space structure on $\mathfrak{S}_n$, the latter defines a center-outward ordering of the permutations in the support of $P$ and extends the classic metric-based formulation of consensus ranking (medians corresponding then to the deepest permutations). The axiomatic properties that ranking depths should ideally possess are listed, while computational and generalization issues are studied at length. Beyond the theoretical analysis carried out, the relevance of the novel concepts and methods introduced for a wide variety of statistical tasks are also supported by numerous numerical experiments. Morgane Goibert, Stéphan Clémençon, Ekhine Irurozki, Pavlo Mozharovskyi |
AISTATS | 4 |
| 2022 | Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMFabstractThis paper tackles post-hoc interpretability for audio processing networks. Our goal is to interpret decisions of a trained network in terms of high-level audio objects that are also listenable for the end-user. To this end, we propose a novel interpreter design that incorporates non-negative matrix factorization (NMF). In particular, a regularized interpreter module is trained to take hidden layer representations of the targeted network as input and produce time activations of pre-learnt NMF components as intermediate outputs. Our methodology allows us to generate intuitive audio-based interpretations that explicitly enhance parts of the input signal most relevant for a network's decision. We demonstrate our method's applicability on popular benchmarks, including a real-world multi-label classification task. Jayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc, Gaël Richard |
NeurIPS | 3 |
| 2021 | When OT meets MoM: Robust estimation of Wasserstein DistanceabstractOriginated from Optimal Transport, the Wasserstein distance has gained importance in Machine Learning due to its appealing geometrical properties and the increasing availability of efficient approximations. It owes its recent ubiquity in generative modelling and variational inference to its ability to cope with distributions having non overlapping support. In this work, we consider the problem of estimating the Wasserstein distance between two probability distributions when observations are polluted by outliers. To that end, we investigate how to leverage a Medians of Means (MoM) approach to provide robust estimates. Exploiting the dual Kantorovitch formulation of the Wasserstein distance, we introduce and discuss novel MoM-based robust estimators whose consistency is studied under a data contamination model and for which convergence rates are provided. Beyond computational issues, the choice of the partition size, i.e., the unique parameter of theses robust estimators, is investigated in numerical experiments. Furthermore, these MoM estimators make Wasserstein Generative Adversarial Network (WGAN) robust to outliers, as witnessed by an empirical study on two benchmarks CIFAR10 and Fashion MNIST. Guillaume Staerman, Pierre Laforgue, Pavlo Mozharovskyi, Florence d'Alché-Buc |
AISTATS | 3 |
| 2021 | A Framework to Learn with InterpretationabstractTo tackle interpretability in deep learning, we present a novel framework to jointly learn a predictive model and its associated interpretation model. The interpreter provides both local and global interpretability about the predictive model in terms of human-understandable high level attribute functions, with minimal loss of accuracy. This is achieved by a dedicated architecture and well chosen regularization penalties. We seek for a small-size dictionary of high level attribute functions that take as inputs the outputs of selected hidden layers and whose outputs feed a linear classifier. We impose strong conciseness on the activation of attributes with an entropy-based criterion while enforcing fidelity to both inputs and outputs of the predictive model. A detailed pipeline to visualize the learnt features is also developed. Moreover, besides generating interpretable models by design, our approach can be specialized to provide post-hoc interpretations for a pre-trained neural network. We validate our approach against several state-of-the-art methods on multiple datasets and show its efficacy on both kinds of tasks. Jayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc |
NeurIPS | 2 |
| 2020 | The Area of the Convex Hull of Sampled Curves: a Robust Functional Statistical Depth measureabstractWith the ubiquity of sensors in the IoT era, statistical observations are becoming increasingly available in the form of massive (multivariate) time-series. Formulated as unsupervised anomaly detection tasks, an abundance of applications like aviation safety management, the health monitoring of complex infrastructures or fraud detection can now rely on such functional data, acquired and stored with an ever finer granularity. The concept of \textit{statistical depth}, which reflects centrality of an arbitrary observation w.r.t. a statistical population may play a crucial role in this regard, anomalies corresponding to observations with ’small’ depth. Supported by sound theoretical and computational developments in the recent decades, it has proven to be extremely useful, in particular in functional spaces. However, most approaches documented in the literature consist in evaluating independently the centrality of each point forming the time series and consequently exhibit a certain insensitivity to possible shape changes.In this paper, we propose a novel notion of functional depth based on the area of the convex hull of sampled curves, capturing gradual departures from centrality, even beyond the envelope of the data, in a natural fashion.We discuss practical relevance of commonly imposed axioms on functional depths and investigate which of them are satisfied by the notion of depth we promote here. Estimation and computational issues are also adressed and various numerical experiments provide empirical evidence of the relevance of the approach proposed. Guillaume Staerman, Pavlo Mozharovskyi, Stéphan Clémençon |
AISTATS | 2 |
| 2019 | Functional Isolation ForestabstractFor the purpose of monitoring the behavior of complex infrastructures (\textit{e.g.} aircrafts, transport or energy networks), high-rate sensors are deployed to capture multivariate data, generally unlabeled, in quasi continuous-time to detect quickly the occurrence of anomalies that may jeopardize the smooth operation of the system of interest. The statistical analysis of such massive data of functional nature raises many challenging methodological questions. The primary goal of this paper is to extend the popular {\scshape Isolation Forest} (IF) approach to Anomaly Detection, originally dedicated to finite dimensional observations, to functional data. The major difficulty lies in the wide variety of topological structures that may equip a space of functions and the great variety of patterns that may characterize abnormal curves. We address the issue of (randomly) splitting the functional space in a flexible manner in order to isolate progressively any trajectory from the others, a key ingredient to the efficiency of the algorithm. Beyond a detailed description of the algorithm, computational complexity and stability issues are investigated at length. From the scoring function measuring the degree of abnormality of an observation provided by the proposed variant of the IF algorithm, a \textit{Functional Statistical Depth} function is defined and discussed, as well as a multivariate functional extension. Numerical experiments provide strong empirical evidence of the accuracy of the extension proposed. Guillaume Staerman, Pavlo Mozharovskyi, Stéphan Clémençon, Florence d'Alché-Buc |
ACML | 2 |