EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Yang 0001
dblp:96/2976-1
· DBLP profile ↗
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-2521-2256ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMM: A Unified Weakly-Supervised Anomaly Detection Framework for Multi-Distributional DataabstractWeakly-Supervised Anomaly Detection (WSAD) has garnered increasing research interest in recent years, as it enables superior detection performance while demanding only a small fraction of labeled data. However, existing WSAD methods face two major limitations. From the data aspect, they struggle to detect anomalies between normal clusters or collective anomalies due to overlooking the multi-distribution and complex manifolds of real-world data. From the label aspect, they fall short of detecting unknown anomalies because of the label-insufficiency and anomaly contamination. To address these issues, we propose MMM, a unified WSAD framework for multi-distributional data. The framework consists of three components: a Multi-distribution data modeler captures latent representations of complex data distributions, followed by a Multiform feature extractor that extracts multiple underlying features from the modeler, highlighting the characteristics of potential anomalies. Finally, a Multi-strategy anomaly score estimator converts these features into anomaly scores, with the aid of a novel training approach with three strategies that maximize the utility of both data and labels. Experimental results showed that MMM achieved superior performance and robustness compared to state-of-the-art WSAD methods, while providing interpretable results that facilitate practical anomaly analysis. Xu Tan 0004, Junqi Chen 0001, Jiawei Yang 0001, Jie Chen 0022, Susanto Rahardja |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | Test-Time Learning for Outlier DetectionabstractIn this work, the concept of test-time learning is presented, wherein Machine-Learning (ML) models are constructed by involving unlabeled test samples. Based on this concept, we propose an unsupervised method called Local Augment (LA) designed to improve the performance of trained outlier detectors at the prediction stage without altering the trained models or accessing the training data. LA operates under the only assumption that the model should produce similar outputs for similar inputs, implying that the prediction of a given sample can be enhanced by the predictions for its similar samples. Specifically, LA boosts outlier detection performance during prediction by fusing the outlier score of a given sample with the scores of synthetically neighboring samples generated by adding random perturbations to the given sample. This simple method demonstrates an average improvement of +0.04 Area Under the Receiver Operating Characteristic curve (AUROC) across 22 real-world datasets for all 11 tested detectors. Notably, this represents the pioneering work of enhancing ML models during the prediction stage without the need to modify the trained models or access the training dataset. This work opens up new possibilities for addressing existing bottleneck problems in various ML tasks beyond outlier detection in diverse domains. Jiawei Yang 0001, Jingdong Chen, Susanto Rahardja |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Smoothing Outlier Scores is All You Need to Improve Outlier Detectors (Extended Abstract)abstractExisting outlier detectors calculate outlier scores for data objects independently, ignoring the consistency between score similarity and object similarity. As a result, these detectors may produce inconsistent scores for similar objects, leading the scores of some normal objects to exceed some of outlier objects, increasing the possibility of misclassification. To address this issue, we first assume that similar objects should have similar scores. Then, based on this assumption, we propose neighborhood averaging, an outlier score post-processing technique to improve any single outlier detector, which is the first of its kind. Jiawei Yang 0001, Susanto Rahardja, Pasi Fränti |
ICDE | 1 |
| 2025 | MSS-PAE: Saving Autoencoder-based Outlier Detection from Unexpected Reconstruction
Xu Tan 0004, Jiawei Yang 0001, Junqi Chen 0001, Sylwan Rahardja, Susanto Rahardja |
Pattern Recognit. | 2 |
| 2024 | Ensemble of Deep Variational Mixture Models for Unsupervised ClusteringabstractDeep variational mixture models (DVMMs) have demonstrated promising performance in unsupervised clustering for complicated high-dimensional data such as images. However, their prediction accuracy is often unstable and significantly influenced by randomness, particularly during the initialization of parameters. To reduce this uncertainty, we propose an ensemble approach that combines the predictions of multiple base models. Specifically, we introduce two individual ensemble strategies: voting and merging. In the voting strategy, the final label is determined by selecting the predicted class label with the most votes and lowest Shannon entropy. In the merging strategy, the class probability vectors (scaled by the temperature parameter) from different models are combined to predict the final class label. Experimental results on two image datasets demonstrate that these proposed methods yield reliable and superior clustering performance. Xu Tan 0004, Junqi Chen 0001, Jiawei Yang 0001, Sylwan Rahardja, Mou Wang, Susanto Rahardja |
ICIP | 3 |
| 2024 | FlexAE: A Self-Conditioned Detector To Prevent Model Overfitting For Unsupervised Video Anomaly DetectionabstractUnsupervised Video Anomaly Detection (VAD) has garnered significant attention for its ability to exploit unlabeled videos. However, VAD faces two primary challenges arising from the absence of labels: (i) Striking a balance between overfitting and underfitting, and (ii) Optimal parameter tuning. To tackle these challenges, we propose a novel detector named Flexible AutoEncoder (FlexAE). A fitting-parameter is introduced to regulate the model’s fitting capacity, and a novel Negative Learning (NL) mechanism is integrated to mitigate the influence of anomalies during training. For self-conditioning, a novel algorithm is devised to autonomously update the fitting-parameter and the threshold used in NL based on the reconstruction error. Comprehensive experiments on two benchmark datasets, UCF-Crime and ShanghaiTech, demonstrate that our proposed FlexAE outperforms state-of-the-art methods without the need for manual hyperparameter tuning. Junqi Chen 0001, Xu Tan 0004, Jiawei Yang 0001, Sylwan Rahardja, Susanto Rahardja |
ICIP | 3 |
| 2024 | Joint Selective State Space Model and Detrending for Robust Time Series Anomaly DetectionabstractDeep learning-based sequence models are extensively employed in Time Series Anomaly Detection (TSAD) tasks due to their effective sequential modeling capabilities. However, the ability of TSAD is limited by two key challenges: (i) the ability to model long-range dependency and (ii) the generalization issue in the presence of non-stationary data. To tackle these challenges, an anomaly detector that leverages the selective state space model known for its proficiency in capturing long-term dependencies across various domains is proposed. Additionally, a multi-stage detrending mechanism is introduced to mitigate the prominent trend component in non-stationary data to address the generalization issue. Extensive experiments conducted on real-world public datasets demonstrate that the proposed methods surpass all 12 compared baseline methods. Junqi Chen 0001, Xu Tan 0004, Sylwan Rahardja, Jiawei Yang 0001, Susanto Rahardja |
IEEE Signal Process. Lett. | 4 |
| 2024 | FOOR: Be Careful for Outlier-Score Outliers When Using Unsupervised Outlier EnsemblesabstractOutlier detection is a very important tool in analyzing patterns and detecting unexpected events in social systems. However, the process of outlier detection could be fraught with uncertainty, with difficulties in determining the veracity of an object’s outlier score. We propose a framework for outlier-score outlier removal (FOOR). FOOR is a selection method, which aims to remove inaccurate outlier scores prior to data processing by ensemble techniques, to improve the accuracy of all ensembles. FOOR has rigorously tested with 30 real-world datasets and seven state-of-the-art ensembles over 25 different base detectors. Simulated experiments showed that FOOR significantly improves the existing techniques, with an average (AVG) of +0.05 AUC (from 0.81 to 0.86 AUC). Thus, we recommend FOOR as the new standard for outlier-score preprocessing before ensembles. Jiawei Yang 0001, Sylwan Rahardja, Susanto Rahardja |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Smoothing Outlier Scores Is All You Need to Improve Outlier DetectorsabstractWe hypothesize thatsimilar objects should have similar outlier scores. To the best of our knowledge, all existing outlier detectors calculate the outlier score for each object independently regardless of the outlier scores of the other objects. Therefore, they do not guarantee that similar objects have similar outlier scores. To verify our proposed hypothesis, we propose an outlier score post-processing technique for outlier detectors, called neighborhood averaging (NA) for neighborhood smoothing in outlier score space. It pays attention to objects and their neighbors and guarantees them to have more similar outlier scores than their original scores. Given an object and its outlier score from any outlier detector, NA modifies its outlier score by combining it with its$k$nearest neighbors' scores. We demonstrate the effectivity of NA by using the well-known$k$nearest neighbors ($k$-NN). Experimental results show that NA improves all 10 tested baseline detectors by 13% on average relative to the original results (from 0.70 to 0.79 AUC) evaluated on nine real-world datasets. Moreover, deep-learning-based detectors and even outlier detectors that are already based on$k$-NN are also improved. The experiments also show that in some applications, the choice of detector is no more significant when detectors are jointly used with NA. This may pose a challenge to the generally considered idea that the data model is the most important factor. We open our code on www.outlierNet.com for reproducibility. Jiawei Yang 0001, Susanto Rahardja, Pasi Fränti |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Neighborhood representative for improving outlier detectorsabstractOver the decades, traditional outlier detectors have ignored the group-level factor when calculating outlier scores for objects in data by evaluating only the object-level factor, failing to capture the collective outliers. To mitigate this issue, we present a framework called neighborhood representative (NR), which empowers all the existing outlier detectors to efficiently detect outliers, including collective outliers, while maintaining their computational integrity. It achieves this by selecting representative objects, scoring these objects, then applies the score of the representative objects to its collective objects. Without altering existing detectors, NR is compatible with existing detectors, while improving performance on eleven real world datasets with +8% (0.72 to 0.78 AUC) on average relative to twelve state-of-the-art outlier detectors. The implementation of NR can be found via www.OutlierNet.com for reproducibility. Index Terms—Outlier detection, Preprocessing, Neighborhood representative, K nearest neighbors. Jiawei Yang 0001, Yu Chen 0019, Sylwan Rahardja |
Inf. Sci. | 1 |
| 2023 | Outlier detection: How to Select k for k-nearest-neighbors-based outlier detectors
Jiawei Yang 0001, Xu Tan 0004, Sylwan Rahardja |
Pattern Recognit. Lett. | 1 |
| 2023 | Classification of Interbeat Interval Time-Series Using Attention EntropyabstractClassification of interbeat interval time-series which fluctuates in an irregular and complex manner is very challenging. Typically, entropy methods are employed to quantify the complexity of the time-series for classifying. Traditional entropy methods focus on the frequency distribution of all the observations in a time-series. This requires a relatively long time-series with at least a couple of thousands of data points, which limits their usages in practical applications. The methods are also sensitive to the parameter settings. In this paper, we propose a conceptually new approach calledattention entropy, which pays attention only to the key observations. Instead of counting the frequency of all observations, it analyzes the frequency distribution of the intervals between the key observations in a time-series. Attention entropy does not need any parameter to tune, it is robust to the time-series length, and requires only linear time to compute. Experiments show that it outperforms fourteen state-of-the-art entropy methods evaluated by real-world datasets. It achieves average classification accuracy of AUC = 0.71 while the second-best method, multiscale entropy, achieves AUC = 0.62 when classifying four groups of people with a time-series length of 100. Jiawei Yang 0001, Gulraiz Iqbal Choudhary, Susanto Rahardja, Pasi Fränti |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Sparse random projection isolation forest for outlier detection
Xu Tan 0004, Jiawei Yang 0001, Susanto Rahardja |
Pattern Recognit. Lett. | 2 |
| 2021 | Mean-shift outlier detection and filteringabstractTraditional outlier detection methods create a model for data and then label as outliers for objects that deviate significantly from this model. However, when dat has many outliers, outliers also pollute the model. The model then becomes unreliable, thus rendering most outlier detectors to become ineffective. To solve this problem, we propose a mean-shift outlier detector. This detector employs a mean-shift technique to modify data and cancel the bias caused by the outliers. The mean-shift technique replaces every object by the mean of its k-nearest neighbors which essentially removes the effect of outliers before clustering without the need to know the outliers. In addition, it also detects outliers based on the distance shifted. Our experiments show that the proposed method works well regardless of the number of outliers in the data. This method outperforms all state-of-the-art methods tested, with both real-world numeric datasets as well as generated numeric and string datasets. Jiawei Yang 0001, Susanto Rahardja, Pasi Fränti |
Pattern Recognit. | 1 |