Andrei Manolache

dblp:290/2275 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Graph learning · 49% Trustworthy machine learning · 18% Information extraction and text analysis · 14%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Network and information security
1 paper
Digital forensics and information hiding · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
1.522024
Probabilistic Graph Rewiring via Virtual Nodes · NeurIPS 2024
Probabilistically Rewired Message-Passing Neural Networks · ICLR 2024
Machine learning › Graph learning › graph neural network
graph rewiring
1.522024
Probabilistic Graph Rewiring via Virtual Nodes · NeurIPS 2024
Probabilistically Rewired Message-Passing Neural Networks · ICLR 2024
Machine learning › Graph learning › graph neural network
message passing
1.522024
Probabilistic Graph Rewiring via Virtual Nodes · NeurIPS 2024
Probabilistically Rewired Message-Passing Neural Networks · ICLR 2024
Machine learning › Optimization for machine learning
constrained optimization
0.912025
Learning (Approximately) Equivariant Networks via Constrained Optimization · NeurIPS 2025
Machine learning › Deep learning architectures and training
equivariant neural network
0.912025
Learning (Approximately) Equivariant Networks via Constrained Optimization · NeurIPS 2025
Machine learning › Graph learning › graph neural network
expressive power
0.812024
Probabilistic Graph Rewiring via Virtual Nodes · NeurIPS 2024
Machine learning › Graph learning
graph structure learning
0.812024
Probabilistically Rewired Message-Passing Neural Networks · ICLR 2024
Natural language and speech › Information extraction and text analysis › text mining
text anomaly detection
0.712023
AD-NLP: A Benchmark for Anomaly Detection in Natural Language Processing · EMNLP 2023
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
authorship verification
0.612022
Rethinking the Authorship Verification Experimental Setups · EMNLP 2022
Machine learning › Trustworthy machine learning › fairness
bias evaluation
0.612022
Rethinking the Authorship Verification Experimental Setups · EMNLP 2022
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.612022
AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability
explainable AI
0.612022
Rethinking the Authorship Verification Experimental Setups · EMNLP 2022
Machine learning › Trustworthy machine learning
interpretability
0.612022
Rethinking the Authorship Verification Experimental Setups · EMNLP 2022
Machine learning › Probabilistic and Bayesian machine learning
non-stationary data
0.612022
AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection · NeurIPS 2022
Natural language and speech › Information extraction and text analysis › text mining
stylometry
0.612022
VeriDark: A Large-Scale Benchmark for Authorship Verification on the Dark Web · NeurIPS 2022
Data mining
anomaly detection
0.612022
AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection · NeurIPS 2022
Data mining › anomaly detection
unsupervised anomaly detection
0.612022
AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection · NeurIPS 2022
Digital forensics and information hiding › authorship attribution
authorship verification
0.612022
VeriDark: A Large-Scale Benchmark for Authorship Verification on the Dark Web · NeurIPS 2022
Machine learning › Graph learning › graph neural network
graph transformer
0.212024
Probabilistic Graph Rewiring via Virtual Nodes · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

differentiable sampling · 1.5t-SNE · 1.1optimal transport · 1.1homotopy principles · 0.9constrained optimization · 0.9virtual nodes · 0.8probabilistic graph rewiring · 0.8k-subset sampling · 0.8explainable AI · 0.6NLP baselines · 0.6BERT · 0.6
YearPublicationVenuePosition
2025 Learning (Approximately) Equivariant Networks via Constrained Optimization
abstract
Equivariant neural networks are designed to respect symmetries through their architecture, boosting generalization and sample efficiency when those symmetries are present in the data distribution. Real-world data, however, often departs from perfect symmetry because of noise, structural variation, measurement bias, or other symmetry-breaking effects. Strictly equivariant models may struggle to fit the data, while unconstrained models lack a principled way to leverage partial symmetries. Even when the data is fully symmetric, enforcing equivariance can hurt training by limiting the model to a restricted region of the parameter space. Guided by homotopy principles, where an optimization problem is solved by gradually transforming a simpler problem into a complex one, we introduce Adaptive Constrained Equivariance (ACE), a constrained optimization approach that starts with a flexible, non-equivariant model and gradually reduces its deviation from equivariance. This gradual tightening smooths training early on and settles the model at a data-driven equilibrium, balancing between equivariance and non-equivariance. Across multiple architectures and tasks, our method consistently improves performance metrics, sample efficiency, and robustness to input perturbations compared with strictly equivariant models and heuristic equivariance relaxations.
Andrei Manolache, Luiz F. O. Chamon, Mathias Niepert
NeurIPS1
2024 Probabilistically Rewired Message-Passing Neural Networks
abstract
Message-passing graph neural networks (MPNNs) emerged as powerful tools for processing graph-structured input. However, they operate on a fixed input graph structure, ignoring potential noise and missing information. Furthermore, their local aggregation mechanism can lead to problems such as over-squashing and limited expressive power in capturing relevant graph structures. Existing solutions to these challenges have primarily relied on heuristic methods, often disregarding the underlying data distribution. Hence, devising principled approaches for learning to infer graph structures relevant to the given prediction task remains an open challenge. In this work, leveraging recent progress in exact and differentiable k-subset sampling, we devise probabilistically rewired MPNNs (PR-MPNNs), which learn to add relevant edges while omitting less beneficial ones. For the first time, our theoretical analysis explores how PR-MPNNs enhance expressive power, and we identify precise conditions under which they outperform purely randomized approaches. Empirically, we demonstrate that our approach effectively mitigates issues like over-squashing and under-reaching. In addition, on established real-world datasets, our method exhibits competitive or superior predictive performance compared to traditional MPNN models and recent graph transformer architectures.
Chendi Qian, Andrei Manolache, Kareem Ahmed, Zhe Zeng 0001, Guy Van den Broeck, Mathias Niepert, Christopher Morris 0001
ICLR2
2024 Probabilistic Graph Rewiring via Virtual Nodes
abstract
Message-passing graph neural networks (MPNNs) have emerged as a powerful paradigm for graph-based machine learning. Despite their effectiveness, MPNNs face challenges such as under-reaching and over-squashing, where limited receptive fields and structural bottlenecks hinder information flow in the graph. While graph transformers hold promise in addressing these issues, their scalability is limited due to quadratic complexity regarding the number of nodes, rendering them impractical for larger graphs. Here, we propose implicitly rewired message-passing neural networks (IPR-MPNNs), a novel approach that integrates implicit probabilistic graph rewiring into MPNNs. By introducing a small number of virtual nodes, i.e., adding additional nodes to a given graph and connecting them to existing nodes, in a differentiable, end-to-end manner, IPR-MPNNs enable long-distance message propagation, circumventing quadratic complexity. Theoretically, we demonstrate that IPR-MPNNs surpass the expressiveness of traditional MPNNs. Empirically, we validate our approach by showcasing its ability to mitigate under-reaching and over-squashing effects, achieving state-of-the-art performance across multiple graph datasets. Notably, IPR-MPNNs outperform graph transformers while maintaining significantly faster computational efficiency.
Chendi Qian, Andrei Manolache, Christopher Morris 0001, Mathias Niepert
NeurIPS2
2023 AD-NLP: A Benchmark for Anomaly Detection in Natural Language Processing
abstract
Deep learning models have reignited the interest in Anomaly Detection research in recent years.Methods for Anomaly Detection in text have shown strong empirical results on ad-hoc anomaly setups that are usually made by downsampling some classes of a labeled dataset.This can lead to reproducibility issues and models that are biased toward detecting particular anomalies while failing to recognize them in more sophisticated scenarios.In the present work, we provide a unified benchmark for detecting various types of anomalies, focusing on problems that can be naturally formulated as Anomaly Detection in text, ranging from syntax to stylistics.In this way, we are hoping to facilitate research in Text Anomaly Detection.We also evaluate and analyze two strong shallow baselines, as well as two of the current state-of-the-art neural approaches, providing insights into the knowledge the neural models are learning when performing the anomaly detection task.We provide the code for evaluation, downloading, and preprocessing the dataset at https: //github.com/mateibejan1/ad-nlp/.
Matei Bejan, Andrei Manolache, Marius Popescu
EMNLP2
2022 Rethinking the Authorship Verification Experimental Setups
abstract
One of the main drivers of the recent advances in authorship verification is the PAN large-scale authorship dataset.Despite generating significant progress in the field, inconsistent performance differences between the closed and open test sets have been reported.To this end, we improve the experimental setup by proposing five new public splits over the PAN dataset, specifically designed to isolate and identify biases related to the text topic and to the author's writing style.We evaluate several BERT-like baselines on these splits, showing that such models are competitive with authorship verification state-of-the-art methods.Furthermore, using explainable AI, we find that these baselines are biased towards named entities.We show that models trained without the named entities obtain better results and generalize better when tested on DarkReddit, our new dataset for authorship verification.Test split O2D2 * O2D2 BERT Naive † Comp.† Closed 93.5 96.4 95.6 75.6 72.2 Clopen 94.0 96.0 97.4 74.1 71.1 Open UA 92.6 92.6 90.2 78.6 68.5 Open UF 91.4 95.1 91.6 79.9 79.0 Open All 80.6 67.5 88.7 75.6 76.9 PAN Closed 93.3 93.5 -74.7 74.2 PAN Open 93.3 94.4 -75.3 74.
Florin Brad, Andrei Manolache, Elena Burceanu, Antonio Barbalau, Radu Tudor Ionescu, Marius Popescu
EMNLP2
2022 AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection
abstract
Analyzing the distribution shift of data is a growing research direction in nowadays Machine Learning (ML), leading to emerging new benchmarks that focus on providing a suitable scenario for studying the generalization properties of ML models. The existing benchmarks are focused on supervised learning, and to the best of our knowledge, there is none for unsupervised learning. Therefore, we introduce an unsupervised anomaly detection benchmark with data that shifts over time, built over Kyoto-2006+, a traffic dataset for network intrusion detection. This type of data meets the premise of shifting the input distribution: it covers a large time span (10 years), with naturally occurring changes over time (e.g. users modifying their behavior patterns, and software updates). We first highlight the non-stationary nature of the data, using a basic per-feature analysis, t-SNE, and an Optimal Transport approach for measuring the overall distribution distances between years. Next, we propose AnoShift, a protocol splitting the data in IID, NEAR, and FAR testing splits. We validate the performance degradation over time with diverse models, ranging from classical approaches to deep learning. Finally, we show that by acknowledging the distribution shift problem and properly addressing it, the performance can be improved compared to the classical training which assumes independent and identically distributed data (on average, by up to 3% for our approach). Dataset and code are available at https://github.com/bit-ml/AnoShift/.
Marius Dragoi, Elena Burceanu, Emanuela Haller, Andrei Manolache, Florin Brad
NeurIPS4
2022 VeriDark: A Large-Scale Benchmark for Authorship Verification on the Dark Web
abstract
The Dark Web represents a hotbed for illicit activity, where users communicate on different market forums in order to exchange goods and services. Law enforcement agencies benefit from forensic tools that perform authorship analysis, in order to identify and profile users based on their textual content. However, authorship analysis has been traditionally studied using corpora featuring literary texts such as fragments from novels or fan fiction, which may not be suitable in a cybercrime context. Moreover, the few works that employ authorship analysis tools for cybercrime prevention usually employ ad-hoc experimental setups and datasets. To address these issues, we release VeriDark: a benchmark comprised of three large scale authorship verification datasets and one authorship identification dataset obtained from user activity from either Dark Web related Reddit communities or popular illicit Dark Web market forums. We evaluate competitive NLP baselines on the three datasets and perform an analysis of the predictions to better understand the limitations of such approaches. We make the datasets and baselines publicly available at https://github.com/bit-ml/VeriDark .
Andrei Manolache, Florin Brad, Antonio Barbalau, Radu Tudor Ionescu, Marius Popescu
NeurIPS1
2021 DATE: Detecting Anomalies in Text via Self-Supervision of Transformers
abstract
Leveraging deep learning models for Anomaly Detection (AD) has seen widespread use in recent years due to superior performances over traditional methods.Recent deep methods for anomalies in images learn better features of normality in an end-to-end self-supervised setting.These methods train a model to discriminate between different transformations applied to visual data and then use the output to compute an anomaly score.We use this approach for AD in text, by introducing a novel pretext task on text sequences.We learn our DATE model end-to-end, enforcing two independent and complementary self-supervision signals, one at the token-level and one at the sequencelevel.Under this new task formulation, we show strong quantitative and qualitative results on the 20Newsgroups and AG News datasets.In the semi-supervised setting, we outperform state-of-the-art results by +13.5% and +6.9%, respectively (AUROC).In the unsupervised configuration, DATE surpasses all other methods even when 10% of its training data is contaminated with outliers (compared with 0% for the others).
Andrei Manolache, Florin Brad, Elena Burceanu
NAACL-HLT1