Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Vincent Vercruyssen

dblp:206/3254 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0003-3645-3135ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 85% Machine learning and data management · 15%
Artificial intelligence
3 papers
Transfer learning and domain adaptation · 44% Learning paradigms · 38% Probabilistic and Bayesian machine learning · 19%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
2.042023
Learning from Positive and Unlabeled Multi-Instance Bags in Anomaly Detection · KDD 2023
Transferring the Contamination Factor between Anomaly Detection Domains by Shape Similarity · AAAI 2022
Transfer Learning for Anomaly Detection through Localized and Unsupervised Instance Selection · AAAI 2020
Data mining › anomaly detection › label-efficient anomaly detection
semi-supervised anomaly detection
0.822020
Transfer Learning for Anomaly Detection through Localized and Unsupervised Instance Selection · AAAI 2020
Semi-Supervised Anomaly Detection with an Application to Water Analytics · ICDM 2018
Machine learning and data management › weak supervision
positive-unlabeled learning
0.712023
Learning from Positive and Unlabeled Multi-Instance Bags in Anomaly Detection · KDD 2023
Machine learning › Transfer learning and domain adaptation
cross-domain transfer
0.612022
Transferring the Contamination Factor between Anomaly Detection Domains by Shape Similarity · AAAI 2022
Data mining › anomaly detection
contamination factor estimation
0.612022
Transferring the Contamination Factor between Anomaly Detection Domains by Shape Similarity · AAAI 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
class prior estimation
0.412020
Class Prior Estimation in Active Positive and Unlabeled Learning · IJCAI 2020
Machine learning › Learning paradigms › weakly supervised learning
positive-unlabeled learning
0.412020
Class Prior Estimation in Active Positive and Unlabeled Learning · IJCAI 2020
Machine learning › Learning paradigms
semi-supervised learning
0.412020
Class Prior Estimation in Active Positive and Unlabeled Learning · IJCAI 2020
Data mining
clustering
0.112018
Semi-Supervised Anomaly Detection with an Application to Water Analytics · ICDM 2018
Data mining › clustering
constrained clustering
0.112018
Semi-Supervised Anomaly Detection with an Application to Water Analytics · ICDM 2018

Methods — techniques the papers use, named apart from their topics

shape similarity modeling · 1.1distribution modeling · 1.1instance selection · 0.9positive and unlabeled learning · 0.7multi-instance learning · 0.7autoencoder · 0.7nearest-neighbor classification · 0.4nearest neighbor classification · 0.4active learning · 0.4time series analysis · 0.3constrained clustering · 0.3
YearPublicationVenuePosition
2023 Learning from Positive and Unlabeled Multi-Instance Bags in Anomaly Detection
abstract
In the multi-instance learning (MIL) setting instances are grouped together into bags. Labels are provided only for the bags and not on the level of individual instances. A positive bag label means that at least one instance inside the bag is positive, while a negative bag label restricts all the instances in the bag to be negative. MIL data naturally arises in many contexts, such as anomaly detection, where labels are rare and costly, and one often ends up annotating the label for sets of instances. Moreover, in many real-world anomaly detection problems, only positive labels are collected because they usually represent critical events. Such a setting, where only positive labels are provided along with unlabeled data, is called Positive and Unlabeled (PU) learning. Despite being useful for several use cases, there is no work dedicated to learning from positive and unlabeled data in a multi-instance setting for anomaly detection. Therefore, we propose the first method that learns from PU bags in anomaly detection. Our method uses an autoencoder as an underlying anomaly detector. We alter the autoencoder's objective function and propose a new loss that allows it to learn from positive and unlabeled bags of instances. We theoretically analyze this method. Experimentally, we evaluate our method on 30 datasets and show that it performs better than multiple baselines adapted to work in our setting.
Lorenzo Perini, Vincent Vercruyssen, Jesse Davis
KDD2
2022 Transferring the Contamination Factor between Anomaly Detection Domains by Shape Similarity
abstract
Anomaly detection attempts to find examples in a dataset that do not conform to the expected behavior. Algorithms for this task assign an anomaly score to each example representing its degree of anomalousness. Setting a threshold on the anomaly scores enables converting these scores into a discrete prediction for each example. Setting an appropriate threshold is challenging in practice since anomaly detection is often treated as an unsupervised problem. A common approach is to set the threshold based on the dataset's contamination factor, i.e., the proportion of anomalous examples in the data. While the contamination factor may be known based on domain knowledge, it is often necessary to estimate it by labeling data. However, many anomaly detection problems involve monitoring multiple related, yet slightly different entities (e.g., a fleet of machines). Then, estimating the contamination factor for each dataset separately by labeling data would be extremely time-consuming. Therefore, this paper introduces a method for transferring the known contamination factor from one dataset (the source domain) to a related dataset where it is unknown (the target domain). Our approach does not require labeled target data and is based on modeling the shape of the distribution of the anomaly scores in both domains. We theoretically analyze how our method behaves when the (biased) target domain anomaly score distribution converges to its true one. Empirically, our method outperforms several baselines on real-world datasets.
Lorenzo Perini, Vincent Vercruyssen, Jesse Davis
AAAI2
2022 Multi-domain Active Learning for Semi-supervised Anomaly Detection
Vincent Vercruyssen, Lorenzo Perini, Wannes Meert, Jesse Davis
ECML/PKDD (4)1
2020 Transfer Learning for Anomaly Detection through Localized and Unsupervised Instance Selection
abstract
Anomaly detection attempts to identify instances that deviate from expected behavior. Constructing performant anomaly detectors on real-world problems often requires some labeled data, which can be difficult and costly to obtain. However, often one considers multiple, related anomaly detection tasks. Therefore, it may be possible to transfer labeled instances from a related anomaly detection task to the problem at hand. This paper proposes a novel transfer learning algorithm for anomaly detection that selects and transfers relevant labeled instances from a source anomaly detection task to a target one. Then, it classifies target instances using a novel semi-supervised nearest-neighbors technique that considers both unlabeled target and transferred, labeled source instances. The algorithm outperforms a multitude of state-of-the-art transfer learning methods and unsupervised anomaly detection methods on a large benchmark. Furthermore, it outperforms its rivals on a real-world task of detecting anomalous water usage in retail stores.
Vincent Vercruyssen, Wannes Meert, Jesse Davis
AAAI1
2020 Class Prior Estimation in Active Positive and Unlabeled Learning
abstract
Estimating the proportion of positive examples (i.e., the class prior) from positive and unlabeled (PU) data is an important task that facilitates learning a classifier from such data. In this paper, we explore how to tackle this problem when the observed labels were acquired via active learning. This introduces the challenge that the observed labels were not selected completely at random, which is the primary assumption underpinning existing approaches to estimating the class prior from PU data. We analyze this new setting and design an algorithm that is able to estimate the class prior for a given active learning strategy. Empirically, we show that our approach accurately recovers the true class prior on a benchmark of anomaly detection datasets and that it does so more accurately than existing methods.
Lorenzo Perini, Vincent Vercruyssen, Jesse Davis
IJCAI2
2020 Quantifying the Confidence of Anomaly Detectors in Their Example-Wise Predictions
Lorenzo Perini, Vincent Vercruyssen, Jesse Davis
ECML/PKDD (3)2
2020 "Now you see it, now you don't!" Detecting Suspicious Pattern Absences in Continuous Time Series
abstract
Given its large applicational potential, time series anomaly detection has become a crucial data mining task. Its goal is to identify periods of a time series where there is a deviation from the expected behavior. Existing approaches focus on analyzing whether the currently observed behavior differs from previously seen, normal behavior. In contrast, this paper tackles the the task where the absence of a previously observed behavior is indicative of an anomaly. In other words, a pattern that is expected to recur in the time series is absent. In real-world use cases, absent patterns can be linked to serious problems. For instance, if a scheduled, regular maintenance operation of a machine does not take place, this can be harmful to the machine at a later time. In this paper, we introduce the task of detecting when a specific pattern is absent in a real-valued time series. We propose a novel technique called FZapPa that can address this task. Empirically, FZapPa outperforms existing anomaly techniques on a benchmark of real-world datasets.
Vincent Vercruyssen, Wannes Meert, Jesse Davis
SDM1
2019 Pattern-Based Anomaly Detection in Mixed-Type Time Series
Len Feremans, Vincent Vercruyssen, Boris Cule, Wannes Meert, Bart Goethals
ECML/PKDD (1)2
2018 Semi-Supervised Anomaly Detection with an Application to Water Analytics
abstract
Nowadays, all aspects of a production process are continuously monitored and visualized in a dashboard. Equipment is monitored using a variety of sensors, natural resource usage is tracked, and interventions are recorded. In this context, a common task is to identify anomalous behavior from the time series data generated by sensors. As manually analyzing such data is laborious and expensive, automated approaches have the potential to be much more efficient as well as cost effective. While anomaly detection could be posed as a supervised learning problem, typically this is not possible as few or no labeled examples of anomalous behavior are available and it is oftentimes infeasible or undesirable to collect them. Therefore, unsupervised approaches are commonly employed which typically identify anomalies as deviations from normal (i.e., common or frequent) behavior. However, in many real-world settings several types of normal behavior exist that occur less frequently than some anomalous behaviors. In this paper, we propose a novel constrained-clustering-based approach for anomaly detection that works in both an unsupervised and semi-supervised setting. Starting from an unlabeled data set, the approach is able to gradually incorporate expert-provided feedback to improve its performance. We evaluated our approach on real-world water monitoring time series data from supermarkets in collaboration with Colruyt Group, one of Belgiums largest retail companies. Empirically, we found that our approach outperforms the current detection system as well as several other baselines. Our system is currently deployed and used by the company to analyze water usage for 20 stores on a daily basis.
Vincent Vercruyssen, Wannes Meert, Gust Verbruggen, Koen Maes, Ruben Baumer, Jesse Davis
ICDM1