Anjin Liu

dblp:152/8041 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-0733-7138ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Autonomous Concept Drift Threshold Determination
abstract
Existing drift detection methods focus on designing sensitive test statistics. They treat the detection threshold as a fixed hyperparameter, set once to balance false alarms and late detections, and applied uniformly across all datasets and over time. However, maintaining model performance is the key objective from the perspective of machine learning, and we observe that model performance is highly sensitive to this threshold. This observation inspires us to investigate whether a dynamic threshold could be provably better. In this paper, we prove that a threshold that adapts over time can outperform any single fixed threshold. The main idea of the proof is that a dynamic strategy, constructed by combining the best threshold from each individual data segment, is guaranteed to outperform any single threshold that apply to all segments. Based on the theorem, we propose a Dynamic Threshold Determination algorithm. It enhances existing drift detection frameworks with a novel comparison phase to inform how the threshold should be adjusted. Extensive experiments on a wide range of synthetic and real-world datasets, including both image and tabular data, validate that our approach substantially enhances the performance of state-of-the-art drift detectors.
Pengqian Lu, Jie Lu 0001, Anjin Liu, En Yu, Guangquan Zhang 0001
AAAI3
2026 TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
abstract
Detecting hallucinations in Retrieval-Augmented Generation (RAG) remains a critical reliability challenge, as ungrounded responses can have severe consequences in high-stakes applications such as clinical decision support, legal research assistants, and autonomous agents that act on retrieved evidence.Prior approaches attribute hallucinations to a binary conflict between internal knowledge stored in FFNs and the retrieved context.However, this perspective is incomplete, failing to account for the impact of other components of the LLM, such as the user query, previously generated tokens, the self token, and the Final LayerNorm adjustment.To comprehensively capture the impact of these components on hallucination detection, we propose TPA which mathematically attributes each token's probability to seven distinct sources: Query, RAG Context, Past Token, Self Token, FFN, Final LayerNorm, and Initial Embedding.This attribution quantifies how each source contributes to the generation of the next token.Specifically, we aggregate these attribution scores by Part-of-Speech (POS) tags to quantify the contribution of each model component to the generation of specific linguistic categories within a response.By leveraging these patterns, such as detecting anomalies where Nouns rely heavily on LayerNorm, TPA effectively identifies hallucinated responses.Extensive experiments on five LLMs (Llama2-7B/13B, Llama3-8B, Mistral-7B, and Qwen3-8B) demonstrate that TPA achieves state-of-the-art performance across diverse architectures.
Pengqian Lu, Jie Lu 0001, Anjin Liu, Guangquan Zhang 0001
ACL (1)3
2025 Early Concept Drift Detection via Prediction Uncertainty
abstract
Concept drift, characterized by unpredictable changes in data distribution over time, poses significant challenges to machine learning models in streaming data scenarios. Although error rate-based concept drift detectors are widely used, they often fail to identify drift in the early stages when the data distribution changes but error rates remain constant. This paper introduces the Prediction Uncertainty Index (PU-index), derived from the prediction uncertainty of the classifier, as a superior alternative to the error rate for drift detection. Our theoretical analysis demonstrates that: (1) The PU-index can detect drift even when error rates remain stable. (2) Any change in the error rate will lead to a corresponding change in the PU-index. These properties make the PU-index a more sensitive and robust indicator for drift detection compared to existing methods. We also propose a PU-index-based Drift Detector (PUDD) that employs a novel Adaptive PU-index Bucketing algorithm for detecting drift. Empirical evaluations on both synthetic and real-world datasets demonstrate PUDD’s efficacy in detecting drift in structured and image data.
Pengqian Lu, Jie Lu 0001, Anjin Liu, Guangquan Zhang 0001
AAAI3
2025 Tracking Correlations Between Multiple Data Streams Through Evolutionary Regressor Chains
abstract
In a real-world setting, several correlational data streams are active at once. An essential question is how to use the correlations between data streams to enhance the effectiveness of machine learning models. The fact that data streams are nonstationary and the correlations across data streams might change over time presents another difficulty. We suggest an ensemble chain-structured model, Evolutionary regressor chains (RCs), to track the correlations between data streams to solve these issues. We develop a heuristic order searching approach to search for the chain's optimal order. With the ability to monitor the dynamicity of the correlations, the heuristic order searching technique can also update the chains over time. Furthermore, a way for reducing computing complexity while maintaining the ensemble's diversity is proposed. The method's theoretical foundation is established through a dynamic regret analysis proving optimal adaptation in the data streams. The outcomes of our experiments demonstrate the effectiveness of Evolutionary RCs.
Jie Lu 0001, Anjin Liu, Xin Yao 0001, Guangquan Zhang 0001
IEEE Trans. Cybern.3
2025 Adaptive Information Fusion-Based Concept Drift Learning for Evolving Multiple Data Streams
abstract
Concept drift arises from unpredictable data distribution shifts, degrading model performance. In evolving multiple data streams, these drifts pose greater challenges due to dynamic changes and uncertain inter-stream correlations, demanding robust accuracy and generalization. To address this issue, in this article, we propose a novel multiple data stream learning method, called the adaptive information fusion-based concept drift learning (AIF-CD) method, to adaptively handle multiple data streams with heterogeneous feature spaces and complex drift situations. First, a real-time learning method with a cooperation scheme is proposed to handle multiple data streams. Second, an information fusion-based augmentation process is designed to help enhance the learning efficiency of each stream. Next, a drift severity identification-based adaptation strategy and a process to selectively use the previous timestamps' data are introduced to enhance learning robustness in both synchronous and asynchronous scenarios. Moreover, a detailed runtime complexity and theoretical analysis further explains the learning efficiency of our method. Our key innovation combines real-time adaptation with theoretical guarantees for complex, evolving multi-stream learning. The experiment results in various scenarios under synchronous and asynchronous settings show that the proposed method is more efficient than other benchmark methods.
Kun Wang 0050, Jie Lu 0001, Anjin Liu
IEEE Trans. Knowl. Data Eng.3
2024 A self-adaptive ensemble for user interest drift learning
Kun Wang 0050, Li Xiong 0002, Anjin Liu, Guangquan Zhang 0001, Jie Lu 0001
Neurocomputing3
2024 TS-DM: A Time Segmentation-Based Data Stream Learning Method for Concept Drift Adaptation
abstract
Concept drift arises from the uncertainty of data distribution over time and is common in data stream. While numerous methods have been developed to assist machine learning models in adapting to such changeable data, the problem of improperly keeping or discarding data samples remains. This may results in the loss of valuable knowledge that could be utilized in subsequent time points, ultimately affecting the model's accuracy. To address this issue, a novel method called time segmentation-based data stream learning method (TS-DM) is developed to help segment and learn the streaming data for concept drift adaptation. First, a chunk-based segmentation strategy is given to segment normal and drift chunks. Building upon this, a chunk-based evolving segmentation (CES) strategy is proposed to mine and segment the data chunk when both old and new concepts coexist. Furthermore, a warning level data segmentation process (CES-W) and a high-low-drift tradeoff handling process are developed to enhance the generalization and robustness. To evaluate the performance and efficiency of our proposed method, we conduct experiments on both synthetic and real-world datasets. By comparing the results with several state-of-the-art data stream learning methods, the experimental findings demonstrate the efficiency of the proposed method.
Kun Wang 0050, Jie Lu 0001, Anjin Liu, Guangquan Zhang 0001
IEEE Trans. Cybern.3
2023 TCR-M: A Topic Change Recognition-based Method for Data Stream Learning
abstract
Data stream learning has received more and more attention in recent years, change tracking and real-time prediction of data streams under uncertainty have been highly focused. With the development of the information age, more and more different types of data streams have been generated, bringing challenges to the research in this field. Among them, text data streams, as one of the categories, also need to be mined and predicted in real-time. This paper addresses this problem by proposing a topic change recognition-based method (TCR-M) for data stream learning, thus helping support text data stream learning. We first propose a topic change recognition process that extracts the topics of the text data stream at each time point, tracks and determines the severity of the topic change, and locates the time points when significant changes occur. Next, an ensemble learning model is constructed and a separate base learner is simultaneously trained to correct the prediction results of the ensemble learning model, which is updated based on the topic change recognition results. To verify the effectiveness of the method, a number of text data streams are collected for evaluation, then outputting the topic change recognition results and prediction results. By comparing it with benchmark methods, the proposed method shows its efficiency. In future research, further improvements are needed for learning and application.
Kun Wang 0050, Jie Lu 0001, Anjin Liu, Guangquan Zhang 0001
KES3
2023 Evolving Gradient Boost: A Pruning Scheme Based on Loss Improvement Ratio for Learning Under Concept Drift
abstract
In nonstationary environments, data distributions can change over time. This phenomenon is known as concept drift, and the related models need to adapt if they are to remain accurate. With gradient boosting (GB) ensemble models, selecting which weak learners to keep/prune to maintain model accuracy under concept drift is nontrivial research. Unlike existing models such as AdaBoost, which can directly compare weak learners' performance by their accuracy (a metric between [0, 1]), in GB, weak learners' performance is measured with different scales. To address the performance measurement scaling issue, we propose a novel criterion to evaluate weak learners in GB models, called the loss improvement ratio (LIR). Based on LIR, we develop two pruning strategies: 1) naive pruning (NP), which simply deletes all learners with increasing loss and 2) statistical pruning (SP), which removes learners if their loss increase meets a significance threshold. We also devise a scheme to dynamically switch between NP and SP to achieve the best performance. We implement the scheme as a concept drift learning algorithm, called evolving gradient boost (LIR-eGB). On average, LIR-eGB delivered the best performance against state-of-the-art methods on both stationary and nonstationary data.
Kun Wang 0050, Jie Lu 0001, Anjin Liu, Guangquan Zhang 0001, Li Xiong 0002
IEEE Trans. Cybern.3
2023 Concept Drift Detection Delay Index
abstract
Data streams may encounter data distribution changes, which can significantly impair the accuracy of models. Concept drift detection tracks data distribution changes and signals when to update models. Many drift detection methods apply thresholds to distinguish between drift or non-drift streams and to claim their method outperforms others with non-aligned drift thresholds. We consider that selecting a proper drift threshold could be more important than developing a new drift detection algorithm, and different drift detection algorithms may end up with very similar performance with aligned drift thresholds. To better understand this process, we propose a novel threshold selection algorithm to align the drift thresholds of a set of algorithms so that they are all at the same sensitivity level. Based on comprehensive experiment evaluations, we observed that several state-of-the-art drift detection algorithms could achieve similar results by aligning their thresholds, providing a novel insight to explain how drift detection algorithms contribute to data stream learning. We noticed that a higher detection sensitivity improves accuracy for data streams with frequent distribution change. The evaluation results are showing that drift thresholds should not be fixed during stream learning. Rather, they should adjust dynamically based on the prevailing conditions of the data stream.
Anjin Liu, Jie Lu 0001, Yiliao Song, Junyu Xuan, Guangquan Zhang 0001
IEEE Trans. Knowl. Data Eng.1
2022 Elastic gradient boosting decision tree with adaptive iterations for concept drift adaptation
Kun Wang 0050, Jie Lu 0001, Anjin Liu, Yiliao Song, Li Xiong 0002, Guangquan Zhang 0001
Neurocomputing3
2022 Real-Time Prediction System of Train Carriage Load Based on Multi-Stream Fuzzy Learning
abstract
When a train leaves a platform, knowing the carriage load (the number of passengers in each carriage) of this train will support train managers to guide passengers at the next platform to choose carriages to avoid congestion. This capacity has become critical since the onset of the pandemic. However, with the dynamicity of passengers and the speed of trains improved (about 3 minutes travel between stations) as well as the station stop period reduced (60–90 second per station), the real-time prediction is more challenging. This paper presents an intelligent system, which is developed in collaboration with Sydney Trains, for real-time predicting carriage load across a city passenger train network. The system comprises three innovations. First, a fuzzy time-matching method significantly improves prediction accuracy in the uncertain situations and allows noisy historical data to be used for training. Second, the LightGBM model is extended with an incremental learning scheme to make forecasting in real-time possible. Third, a new multi-stream learning strategy that merges data streams with similar concept drift patterns is pioneered to increase the amount of suitable training data while reducing generalization errors. A comprehensive suite of practical tests on real-world datasets demonstrates the merit of these solutions.
Hang Yu 0006, Jie Lu 0001, Anjin Liu, Bin Wang 0045, Guangquan Zhang 0001
IEEE Trans. Intell. Transp. Syst.3
2022 A Segment-Based Drift Adaptation Method for Data Streams
abstract
In concept drift adaptation, we aim to design a blind or an informed strategy to update our best predictor for future data at each time point. However, existing informed drift adaptation methods need to wait for an entire batch of data to detect drift and then update the predictor (if drift is detected), which causes adaptation delay. To overcome the adaptation delay, we propose a sequentially updated statistic, called drift-gradient to quantify the increase of distributional discrepancy when every new instance arrives. Based on drift-gradient, a segment-based drift adaptation (SEGA) method is developed to online update our best predictor. Drift-gradient is defined on a segment in the training set. It can precisely quantify the increase of distributional discrepancy between the old segment and the newest segment when only one new instance is available at each time point. A lower value of drift-gradient on the old segment represents that the distribution of the new instance is closer to the distribution of the old segment. Based on the drift-gradient, SEGA retrains our best predictors with the segments that have the minimum drift-gradient when every new instance arrives. SEGA has been validated by extensive experiments on both synthetic and real-world, classification and regression data streams. The experimental results show that SEGA outperforms competitive blind and informed drift adaptation methods.
Yiliao Song, Jie Lu 0001, Anjin Liu, Haiyan Lu, Guangquan Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Learning Bounds for Open-Set Learning
abstract
Traditional supervised learning aims to train a classifier in the closed-set world, where training and test samples share the same label space. In this paper, we target a more challenging and re_x0002_alistic setting: open-set learning (OSL), where there exist test samples from the classes that are unseen during training. Although researchers have designed many methods from the algorith_x0002_mic perspectives, there are few methods that pro_x0002_vide generalization guarantees on their ability to achieve consistent performance on different train_x0002_ing samples drawn from the same distribution. Motivated by the transfer learning and probably approximate correct (PAC) theory, we make a bold attempt to study OSL by proving its general_x0002_ization error-given training samples with size n, the estimation error will get close to order Op(1/$\sqrt{}$n). This is the first study to provide a generalization bound for OSL, which we do by theoretically investigating the risk of the tar_x0002_get classifier on unknown classes. According to our theory, a novel algorithm, called auxiliary open-set risk (AOSR) is proposed to address the OSL problem. Experiments verify the efficacy of AOSR. The code is available at github.com/AnjinLiu/Openset_Learning_AOSR.
Zhen Fang 0001, Jie Lu 0001, Anjin Liu, Feng Liu 0003, Guangquan Zhang 0001
ICML3
2021 Confident Anchor-Induced Multi-Source Free Domain Adaptation
abstract
Unsupervised domain adaptation has attracted appealing academic attentions by transferring knowledge from labeled source domain to unlabeled target domain. However, most existing methods assume the source data are drawn from a single domain, which cannot be successfully applied to explore complementarily transferable knowledge from multiple source domains with large distribution discrepancies. Moreover, they require access to source data during training, which are inefficient and unpractical due to privacy preservation and memory storage. To address these challenges, we develop a novel Confident-Anchor-induced multi-source-free Domain Adaptation (CAiDA) model, which is a pioneer exploration of knowledge adaptation from multiple source domains to the unlabeled target domain without any source data, but with only pre-trained source models. Specifically, a source-specific transferable perception module is proposed to automatically quantify the contributions of the complementary knowledge transferred from multi-source domains to the target domain. To generate pseudo labels for the target domain without access to the source data, we develop a confident-anchor-induced pseudo label generator by constructing a confident anchor group and assigning each unconfident target sample with a semantic-nearest confident anchor. Furthermore, a class-relationship-aware consistency loss is proposed to preserve consistent inter-class relationships by aligning soft confusion matrices across domains. Theoretical analysis answers why multi-source domains are better than a single source domain, and establishes a novel learning bound to show the effectiveness of exploiting multi-source domains. Experiments on several representative datasets illustrate the superiority of our proposed CAiDA model. The code is available at https://github.com/Learning-group123/CAiDA.
Jiahua Dong 0001, Zhen Fang 0001, Anjin Liu, Gan Sun, Tongliang Liu
NeurIPS3
2021 Concept Drift Detection via Equal Intensity k-Means Space Partitioning
abstract
The data stream poses additional challenges to statistical classification tasks because distributions of the training and target samples may differ as time passes. Such a distribution change in streaming data is called concept drift. Numerous histogram-based distribution change detection methods have been proposed to detect drift. Most histograms are developed on the grid-based or tree-based space partitioning algorithms which makes the space partitions arbitrary, unexplainable, and may cause drift blind spots. There is a need to improve the drift detection accuracy for the histogram-based methods with the unsupervised setting. To address this problem, we propose a cluster-based histogram, called equal intensity k -means space partitioning (EI-kMeans). In addition, a heuristic method to improve the sensitivity of drift detection is introduced. The fundamental idea of improving the sensitivity is to minimize the risk of creating partitions in distribution offset regions. Pearson's chi-square test is used as the statistical hypothesis test so that the test statistics remain independent of the sample distribution. The number of bins and their shapes, which strongly influence the ability to detect drift, are determined dynamically from the sample based on an asymptotic constraint in the chi-square test. Accordingly, three algorithms are developed to implement concept drift detection, including a greedy centroids initialization algorithm, a cluster amplify-shrink algorithm, and a drift detection algorithm. For drift adaptation, we recommend retraining the learner if a drift is detected. The results of experiments on the synthetic and real-world datasets demonstrate the advantages of EI-kMeans and show its efficacy in detecting concept drift.
Anjin Liu, Jie Lu 0001, Guangquan Zhang 0001
IEEE Trans. Cybern.1
2021 Concept Drift Detection: Dealing With Missing Values via Fuzzy Distance Estimations
abstract
In data streams, the data distribution of arriving observations at different time points may change—a phenomenon called concept drift. While detecting concept drift is a relatively mature area of study, solutions to the uncertainty introduced by observations with missing values have only been studied in isolation. No one has yet explored whether or how these solutions might impact drift detection performance. We, however, believe that data imputation methods may actually increase uncertainty in the data rather than reducing it. We also conjecture that imputation can introduce bias into the process of estimating distribution changes during drift detection, which can make it more difficult to train a learning model. Our idea is to focus on estimating the distance between observations rather than estimating the missing values, and to define membership functions that allocate observations to histogram bins according to the estimation errors. Our solution comprises a novel masked distance learning (MDL) algorithm to reduce the cumulative errors caused by iteratively estimating each missing value in an observation and a fuzzy-weighted frequency (FWF) method for identifying discrepancies in the data distribution. The concept drift detection algorithm proposed in this article is a singular and unified algorithm that can handle missing values, but not an imputation algorithm combined with a concept drift detection algorithm. Experiments on both synthetic and real-world datasets demonstrate the advantages of this method and show its robustness in detecting drift in data with missing values. The results show that compared to the best-performing algorithm that handles imputation and drift detection separately, MDL-FWF reduced the average drift detection difference from 10.75% to 5.83%. This is a nearly 46% improvement. These findings reveal that missing values exert a profound impact on concept drift detection, but using fuzzy set theory to model observations can produce more reliable results than imputation.
Anjin Liu, Jie Lu 0001, Guangquan Zhang 0001
IEEE Trans. Fuzzy Syst.1
2021 Diverse Instance-Weighting Ensemble Based on Region Drift Disagreement for Concept Drift Adaptation
abstract
Concept drift refers to changes in the distribution of underlying data and is an inherent property of evolving data streams. Ensemble learning, with dynamic classifiers, has proved to be an efficient method of handling concept drift. However, the best way to create and maintain ensemble diversity with evolving streams is still a challenging problem. In contrast to estimating diversity via inputs, outputs, or classifier parameters, we propose a diversity measurement based on whether the ensemble members agree on the probability of a regional distribution change. In our method, estimations over regional distribution changes are used as instance weights. Constructing different region sets through different schemes will lead to different drift estimation results, thereby creating diversity. The classifiers that disagree the most are selected to maximize diversity. Accordingly, an instance-based ensemble learning algorithm, called the diverse instance-weighting ensemble (DiwE), is developed to address concept drift for data stream classification problems. Evaluations of various synthetic and real-world data stream benchmarks show the effectiveness and advantages of the proposed algorithm.
Anjin Liu, Jie Lu 0001, Guangquan Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Fast Switch Naïve Bayes to Avoid Redundant Update for Concept Drift Learning
abstract
In data stream mining, concept drift may cause the predictions given by machine learning models become less accurate as time passes. Existing concept drift detection and adaptation methods are built based on a framework that is buffering new samples if a drift-warming level is triggered and retraining a new model if a drift-alarm level is triggered. However, these methods neglected the problem that the performance of a learning model could be more sensitive to the amount of training data rather than the concept drift. In other words, a retrained model built on very few data instances could be even worse than the old model trained before the drift. To elaborate and address this problem, we propose a fast switch Naïve Bayes model (fsNB) for concept drift detection and adaptation. The intuition is to apply the idea of following the leader in online learning. We manipulate a sliding and an incremental Naïve Bayes classifier, if the sliding one overwhelms the incremental one, the model reports a drift. The experimental evaluation shows the advantages of fsNB and demonstrates that retraining may not be the best options for a marginal drift.
Anjin Liu, Guangquan Zhang 0001, Kun Wang 0050, Jie Lu 0001
IJCNN1
2019 Knowledge graph-based entity importance learning for multi-stream regression on Australian fuel price forecasting
abstract
A knowledge graph (KG) represents a collection of interlinked descriptions of entities. It has become a key focus for organising and utilising this type of data for applications. Many graph embedding techniques have been proposed to simplify the manipulation while preserving the inherent structure of the KG. However, scant attention has been given to the investigation of the importance of the entities (the nodes of KGs). In this paper, we propose a novel entities importance learning framework that investigates how to weight the entities and use them as a prior knowledge for solving multi-stream regression problems. The framework consists of KG feature extraction, multi-stream correlation analysis, and entity importance learning. To evaluate the proposed method, we implemented the framework based on Wikidata and applied it to Australian retail fuel price forecasting. The experiment results indicate that the proposed method reduces prediction error, which supports the weighted knowledge graph information as a means for improving machine learning model accuracy.
Dennis Chow, Anjin Liu, Guangquan Zhang 0001, Jie Lu 0001
IJCNN2
2019 Learning under Concept Drift: A Review
abstract
Concept drift describes unforeseeable changes in the underlying distribution of streaming data overtime. Concept drift research involves the development of methodologies and techniques for drift detection, understanding, and adaptation. Data analysis has revealed that machine learning in a concept drift environment will result in poor learning results if the drift is not addressed. To help researchers identify which research topics are significant and how to apply related techniques in data analysis tasks, it is necessary that a high quality, instructive review of current research developments and trends in the concept drift field is conducted. In addition, due to the rapid development of concept drift in recent years, the methodologies of learning under concept drift have become noticeably systematic, unveiling a framework which has not been mentioned in literature. This paper reviews over 130 high quality publications in concept drift related research areas, analyzes up-to-date developments in methodologies and techniques, and establishes a framework of learning under concept drift including three main components: concept drift detection, concept drift understanding, and concept drift adaptation. This paper lists and discusses 10 popular synthetic datasets and 14 publicly available benchmark datasets used for evaluating the performance of learning algorithms aiming at handling concept drift. Also, concept drift related research directions are covered and discussed. By providing state-of-the-art knowledge, this survey will directly support researchers in their understanding of research developments in the field of learning under concept drift.
Jie Lu 0001, Anjin Liu, João Gama 0001, Guangquan Zhang 0001
IEEE Trans. Knowl. Data Eng.2
2018 Accumulating regional density dissimilarity for concept drift detection in data streams
Anjin Liu, Jie Lu 0001, Feng Liu 0003, Guangquan Zhang 0001
Pattern Recognit.1
2017 Fuzzy time windowing for gradual concept drift adaptation
abstract
The aim of machine learning is to find hidden insights into historical data, and then apply them to forecast the future data or trends. Machine learning algorithms optimize learning models for lowest error rate based on the assumption that the historical data and the data to be predicted conform to the same knowledge pattern (data distribution). However, if the historical data is not enough, or the knowledge pattern keeps changing (data uncertainty), this assumption will become invalid. In data stream mining, this phenomenon of knowledge pattern changing is called concept drift. To address this issue, we propose a novel fuzzy windowing concept drift adaptation (FW-DA) method. Compared to conventional windowing-based drift adaptation algorithms, FW-DA achieves higher accuracy by allowing the sliding windows to keep an overlapping period so that the data instances belonging to different concepts can be determined more precisely. In addition, FW-DA statistically guarantees that the upcoming data conforms to the inferred knowledge pattern with a certain confidence level. To evaluate FW-DA, four experiments were conducted using both synthetic and real-world data sets. The experiment results show that FW-DA outperforms the other windowing-based methods including state-of-the-art drift adaptation methods.
Anjin Liu, Guangquan Zhang 0001, Jie Lu 0001
FUZZ-IEEE1
2017 Regional Concept Drift Detection and Density Synchronized Drift Adaptation
abstract
In data stream mining, the emergence of new patterns or a pattern ceasing to exist is called concept drift. Concept drift makes the learning process complicated because of the inconsistency between existing data and upcoming data. Since concept drift was first proposed, numerous articles have been published to address this issue in terms of distribution analysis. However, most distribution-based drift detection methods assume that a drift happens at an exact time point, and the data arrived before that time point is considered not important. Thus, if a drift only occurs in a small region of the entire feature space, the other non-drifted regions may also be suspended, thereby reducing the learning efficiency of models. To retrieve non-drifted information from suspended historical data, we propose a local drift degree (LDD) measurement that can continuously monitor regional density changes. Instead of suspending all historical data after a drift, we synchronize the regional density discrepancies according to LDD. Experimental evaluations on three public data sets show that our concept drift adaptation algorithm improves accuracy compared to other methods.
Anjin Liu, Yiliao Song, Guangquan Zhang 0001, Jie Lu 0001
IJCAI1
2014 Concept Drift Detection Based on Anomaly Analysis
Anjin Liu, Guangquan Zhang 0001, Jie Lu 0001
ICONIP (1)1