Yiliao Song

dblp:186/7620 · also Lia Song · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
19since 2021 · last 2025
0000-0002-6633-2695ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal Metaphors
abstract
Senqi Yang, Dongyu Zhang, Jing Ren, Ziqi Xu, Xiuzhen Zhang, Yiliao Song, Hongfei Lin, Feng Xia. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Senqi Yang, Dongyu Zhang 0001, Jing Ren 0001, Ziqi Xu 0001, Xiuzhen Zhang 0001, Yiliao Song, Hongfei Lin, Feng Xia 0001
ACL (1)6
2025 Flow: Modularized Agentic Workflow Automation
abstract
Multi-agent frameworks powered by large language models (LLMs) have demonstrated great success in automated planning and task execution. However, the effective adjustment of agentic workflows during execution has not been well studied. An effective workflow adjustment is crucial in real-world scenarios, as the initial plan must adjust to unforeseen challenges and changing conditions in real time to ensure the efficient execution of complex tasks. In this paper, we define workflows as an activity-on-vertex (AOV) graph, which allows continuous workflow refinement by LLM agents through dynamic subtask allocation adjustment based on historical performance and previous AOVs. To further enhance framework performance, we emphasize modularity in workflow design based on evaluating parallelism and dependency complexity. With this design, our proposed multi-agent framework achieves efficient concurrent execution of subtasks, effective goal achievement, and enhanced error tolerance. Empirical results across various practical tasks demonstrate significant improvements in the efficiency of multi-agent frameworks through dynamic workflow refinement and modularization.
Boye Niu, Yiliao Song, Kai Lian, Yifan Shen 0004, Yu Yao 0005, Kun Zhang 0001, Tongliang Liu
ICLR2
2025 Deep Kernel Relative Test for Machine-generated Text Detection
abstract
Recent studies demonstrate that two-sample test can effectively detect machine-generated texts (MGTs) with excellent adaptation ability to texts generated by newer LLMs. However, two-sample test-based detection relies on the assumption that human-written texts (HWTs) must follow the distribution of seen HWTs. As a result, it tends to make mistakes in identifying HWTs that deviate from the seen HWT distribution, limiting their use in sensitive areas like academic integrity verification. To address this issue, we propose to employ non-parametric kernel relative test to detect MGTs by testing whether it is statistically significant that the distribution of a text to be tested is closer to the distribution of HWTs than to the MGTs' distribution. We further develop a kernel optimisation algorithm in relative test to select the best kernel that can enhance the testing capability for MGT detection. As relative test does not assume that a text to be tested must belong exclusively to either MGTs or HWTs, relative test can largely reduce the false positive error compared to two-sample test, offering significant advantages in practice. Extensive experiments demonstrate the superior performance of our method, compared to state-of-the-art non-parametric and parametric detectors. The code and demo are available: https://github.com/xLearn-AU/R-Detect.
Yiliao Song, Zhenqiao Yuan, Shuhai Zhang, Zhen Fang 0001, Feng Liu 0003
ICLR1
2025 A Unified Solution to Diverse Heterogeneities in One-Shot Federated Learning
abstract
One-Shot Federated Learning (OSFL) restricts communication between the server and clients to a single round, significantly reducing communication costs and minimizing privacy leakage risks compared to traditional Federated Learning (FL), which requires multiple rounds of communication. However, existing OSFL frameworks remain vulnerable to distributional heterogeneity, as they primarily focus on model heterogeneity while neglecting data heterogeneity. To bridge this gap, we propose FedHydra, a unified, data-free, OSFL framework designed to effectively address both model and data heterogeneity. Unlike existing OSFL approaches, FedHydra introduces a novel two-stage learning mechanism. Specifically, it incorporates model stratification and heterogeneity-aware stratified aggregation to mitigate the challenges posed by both model and data heterogeneity. By this design, the data and model heterogeneity issues are simultaneously monitored from different aspects during learning. Consequently, FedHydra can effectively mitigate both issues by minimizing their inherent conflicts. We compared FedHydra with five SOTA baselines on four benchmark datasets. Experimental results show that our method outperforms the previous OSFL methods in both homogeneous and heterogeneous settings. The code is available at https://github.com/Jun-B0518/FedHydra.
Yiliao Song, Di Wu 0050, Atul Sajjanhar, Yong Xiang 0001, Wei Zhou 0044, Xiaohui Tao 0001, Yan Li 0002, Yue Li 0017
KDD (2)2
2025 Can Dependencies Induced by LLM-Agent Workflows Be Trusted?
abstract
LLM-agent systems often decompose high-level objectives into subtask dependency graphs, assuming that each subtask’s output is reliable and conditionally independent of others given its parent responses. However, this assumption frequently breaks during execution, as ground-truth responses are inaccessible, leading to inter-agent misalignment—failures caused by inconsistencies and coordination breakdowns among agents. To address this, we propose SeqCV, a dynamic framework for reliable execution under violated conditional independence. SeqCV executes subtasks sequentially, each conditioned on all prior verified responses, and performs consistency checks immediately after agents generate short token sequences. At each checkpoint, a token sequence is accepted only if it represents shared knowledge consistently supported across diverse LLM models; otherwise, it is discarded, triggering recursive subtask decomposition for finer-grained reasoning. Despite its sequential nature, SeqCV avoids repeated corrections on the same misalignment and achieves higher effective throughput than parallel pipelines. Across multiple reasoning and coordination tasks, SeqCV improves accuracy by up to 30\% over existing LLM-agent systems. Code is available at https://github.com/tmllab/2025_NeurIPS_SeqCV.
Yu Yao 0005, Yiliao Song, Yian Xie, Mengdan Fan, Mingyu Guo 0001, Tongliang Liu
NeurIPS2
2025 A Multistream Concept Drift Handling Framework via Data Sharing
abstract
A frequent problem in data stream mining is concept drift, meaning the data distribution changes over time. A common issue when dealing with concept drift is insufficient data. Real-world applications of data stream mining often involve multiple data streams. However, most concept drift methods handle these data streams separately. This study uses data from other data streams to handle the problem of insufficient data. We propose a novel Multistream Concept Drift Handling Framework via data sharing, containing a fuzzy membership-based drift detection (FMDD) component and a fuzzy membership-based drift adaptation (FMDA) component, to train the new learning model for drifting streams by sharing weighted data from other nondrifting streams. A stream fuzzy set is defined with membership functions that measure the degree to which samples belong to a data stream. Our Concept Drift Handling Framework can detect when and in which stream concept drift occurs, and therefore the insufficient data issue can be solved by adding the weighted data from nondrifting streams to train new learning models. Synthetic and real-world experimental results show that our method can help avoid the insufficient data issue and thereby significantly improve the prediction performance.
Jie Lu 0001, Yiliao Song, Guangquan Zhang 0001
IEEE Trans. Cybern.3
2025 TrapNet: Model Inversion Defense via Trapdoor
abstract
Model inversion (MI) attacks, for which effective defense strategies are still lacking, pose significant risks to privacy by reconstructing private training data through access to well-trained classifiers. Addressing this concern, this study introduces TrapNet, designed to defend against advanced MI attacks while maintaining good model utility. TrapNet intentionally injects trapdoors into the classification manifold of the protected target model. In this way, TrapNet can effectively mislead MI attack optimization. Specifically, TrapNet leverages a conditional GAN (cGAN) trained on the private dataset to generate diverse and realistic trapdoor samples. In addition, we propose a graph-matching self-obfuscation strategy and an entropy regularization technique to optimize trapdoor injection while preserving model utility. Compared to the existing defense, TrapNet can provide universal protection to all target classes without access to any auxiliary public data. Extensive experiments on CelebA, VGG-Face, and VGG-Face2 datasets demonstrate TrapNet’s superior performance over existing defenses, including the most advanced NetGuard and BiDO, against state-of-the-art model inversion attacks, i.e., PLG-MI, LOMMA, and Plug&Play.
Wanlun Ma, Derui Wang, Yiliao Song, Minhui Xue 0001, Sheng Wen, Zhengdao Li, Yang Xiang 0001
IEEE Trans. Inf. Forensics Secur.3
2024 FedInverse: Evaluating Privacy Leakage in Federated Learning
abstract
Federated Learning (FL) is a distributed machine learning technique where multiple devices (such as smartphones or IoT devices) train a shared global model by using their local data. FL claims that the data privacy of local participants is preserved well because local data will not be shared with either the server-side or other training participants. However, this paper discovers a pioneering finding that a model inversion (MI) attacker, who acts as a benign participant, can invert the shared global model and obtain the data belonging to other participants. This will lead to severe data-leakage risk in FL because it is difficult to identify attackers from benign participants. In addition, we found even the most advanced defense approaches could not effectively address this issue. Therefore, it is important to evaluate such data-leakage risks of an FL system before using it. To alleviate this issue, we propose FedInverse to evaluate whether the FL global model can be inverted by MI attackers. In particular, FedInverse can be optimized by leveraging the Hilbert-Schmidt independence criterion (HSIC) as a regularizer to adjust the diversity of the MI attack generator. We test FedInverse with three typical MI attackers, GMI, KED-MI, and VMI, and the experiments show our FedInverse method can successfully obtain the data belonging to other participants. The code of this work is available at https://github.com/Jun-B0518/FedInverse
Di Wu 0050, Yiliao Song, Wei Zhou 0044, Yong Xiang 0001, Atul Sajjanhar
ICLR3
2024 Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy
abstract
Large language models (LLMs) such as ChatGPT have exhibited remarkable performance in generating human-like texts. However, machine-generated texts (MGTs) may carry critical risks, such as plagiarism issues and hallucination information. Therefore, it is very urgent and important to detect MGTs in many situations. Unfortunately, it is challenging to distinguish MGTs and human-written texts because the distributional discrepancy between them is often very subtle due to the remarkable performance of LLMS. In this paper, we seek to exploit \textit{maximum mean discrepancy} (MMD) to address this issue in the sense that MMD can well identify distributional discrepancies. However, directly training a detector with MMD using diverse MGTs will incur a significantly increased variance of MMD since MGTs may contain \textit{multiple text populations} due to various LLMs. This will severely impair MMD's ability to measure the difference between two samples. To tackle this, we propose a novel \textit{multi-population} aware optimization method for MMD called MMD-MP, which can \textit{avoid variance increases} and thus improve the stability to measure the distributional discrepancy. Relying on MMD-MP, we develop two methods for paragraph-based and sentence-based detection, respectively. Extensive experiments on various LLMs, \eg, GPT2 and ChatGPT, show superior detection performance of our MMD-MP.
Shuhai Zhang, Yiliao Song, Yuanqing Li 0001, Bo Han 0003, Mingkui Tan
ICLR2
2024 The "Code" of Ethics: A Holistic Audit of AI Code Generators
abstract
AI-powered programming language generation (PLG) models have gained increasing attention due to their ability to generate source code of programs in a few seconds with a plain program description. Despite their remarkable performance, many concerns are raised over the potential risks of their development and deployment, such as legal issues of copyright infringement induced by training usage of licensed code, and malicious consequences due to the unregulated use of these models. In this paper, we present the first-of-its-kind study to systematically investigate the accountability of PLG models from the perspectives of both model development and deployment. In particular, we develop a holistic framework not only to audit the training data usage of PLG models, but also to identify neural code generated by PLG models as well as determine its attribution to a source model. To this end, we propose using membership inference to audit whether a code snippet used is in the PLG model's training data. In addition, we propose a learning-based method to distinguish between human-written code and neural code. In neural code attribution, through both empirical and theoretical analysis, we show that it is impossible to reliably attribute the generation of one code snippet to one model. We then propose two feasible alternative methods: one is to attribute one neural code snippet to one of the candidate PLG models, and the other is to verify whether a set of neural code snippets can be attributed to a given PLG model. The proposed framework thoroughly examines the accountability of PLG models which are verified by extensive experiments. The implementations of our proposed framework are also encapsulated into a new artifact, named CODEFORENSIC, to foster further research.
Wanlun Ma, Yiliao Song, Minhui Xue 0001, Sheng Wen, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.2
2024 Type-LDD: A Type-Driven Lite Concept Drift Detector for Data Streams
abstract
Concept drift is a phenomenon that the distribution of data streams changes with time. When this happens, model predictions become less accurate. Hence, concept drift needs to be detected and adapted. Existing drift detection methods are good at determining when drift has occurred, but few retrieve information about how the drift came to be present in the stream, i.e., what type of drift has occurred. Hence, discussing the impact of the type of drift on adaptation is a difficult thing. To fill this gap, we propose a pre-trained framework for training a drift detector called a type-driven lite concept drift detector (Type-LDD) that retrieves information about both when and how a drift has occurred. In our proposed pre-trained framework, the Type-LDD including a drift-type identifier and a drift-point locator was based on a synthetic dataset containing a range of drift types. When repurposing the pre-trained model for detecting new data streams, a knowledge distillation module fine-tunes the proposed Type-LDD to speed up inference and keep detection accuracy. The proposed Type-LDD is validated on both synthetic data and real-world data, and demonstrated that accurately identifying the type of drift that has occurred can improve adaptation accuracy.
Hang Yu 0006, Jie Lu 0001, Yiliao Song, Shaorong Xie, Guangquan Zhang 0001
IEEE Trans. Knowl. Data Eng.4
2023 Concept Drift Detection Delay Index
abstract
Data streams may encounter data distribution changes, which can significantly impair the accuracy of models. Concept drift detection tracks data distribution changes and signals when to update models. Many drift detection methods apply thresholds to distinguish between drift or non-drift streams and to claim their method outperforms others with non-aligned drift thresholds. We consider that selecting a proper drift threshold could be more important than developing a new drift detection algorithm, and different drift detection algorithms may end up with very similar performance with aligned drift thresholds. To better understand this process, we propose a novel threshold selection algorithm to align the drift thresholds of a set of algorithms so that they are all at the same sensitivity level. Based on comprehensive experiment evaluations, we observed that several state-of-the-art drift detection algorithms could achieve similar results by aligning their thresholds, providing a novel insight to explain how drift detection algorithms contribute to data stream learning. We noticed that a higher detection sensitivity improves accuracy for data streams with frequent distribution change. The evaluation results are showing that drift thresholds should not be fixed during stream learning. Rather, they should adjust dynamically based on the prevailing conditions of the data stream.
Anjin Liu, Jie Lu 0001, Yiliao Song, Junyu Xuan, Guangquan Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2023 Multi-Stream Concept Drift Self-Adaptation Using Graph Neural Network
abstract
Concept drift is the phenomenon where the data distribution in a data stream changes over time. It is a ubiquitous problem in the real-world, for example, a traffic accident would cause a jam in a certain period, leading to a distribution change in traffic speed. Most research in the concept drift field focuses on single data stream, however, few of them consider multi-stream environments which are more in line with the application needs. To fill this gap, we propose a multi-stream prediction setting and a multi-stream concept drift self-adaptation framework using graph neural network, named SAGN. In SAGN, we reconsider the learning procedure of GNN-based predictors from an aspect of concept drift adaptation for multi-stream. By this design, the prediction task is converted into online streaming data tasks in sub-graphs. Each sub-graph corresponds to an adaptation target and will be updated over time. In this way, locally we can overcome drift in each sub-graph by a designed adaptation technique, and globally the correlation between different data streams is well-preserved as a graph structure. Therefore, whether drift occurs or not, in one or several streams, SAGN can provide consistently accurate prediction results. We comprehensively tested SAGN on both synthetic and real-world, drift and non-drift data in the multi-step prediction task. The experiment results show that SAGN is able to achieve state-of-the-art performance in most cases.
Jie Lu 0001, Yiliao Song, Guangquan Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2023 Learning Data Streams With Changing Distributions and Temporal Dependency
abstract
In a data stream, concept drift refers to unpredictable distribution changes over time, which violates the identical-distribution assumption required by conventional machine learning methods. Current concept drift adaptation techniques mostly focus on a data stream with changing distributions. However, since each variable of a data stream is a time series, these variables normally have temporal dependency problems in the real world. How to solve concept drift and temporal dependency problems at the same time is rarely discussed in the concept-drift literature. To solve this situation, this article proves and validates that the testing error decreases faster if a predictor is trained on a temporally reconstructed space when drift occurs. Based on this theory, a novel drift adaptation regression (DAR) framework is designed to predict the label variable for data streams with concept drift and temporal dependency. A new statistic called local drift degree (LDD+) is proposed and used as a drift adaptation technique in the DAR framework to discard outdated instances in a timely way, thereby guaranteeing that the most relevant instances will be selected during the training process. The performance of DAR is demonstrated by a set of experimental evaluations on both synthetic data and real-world data streams.
Yiliao Song, Jie Lu 0001, Haiyan Lu, Guangquan Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Elastic gradient boosting decision tree with adaptive iterations for concept drift adaptation
Kun Wang 0050, Jie Lu 0001, Anjin Liu, Yiliao Song, Li Xiong 0002, Guangquan Zhang 0001
Neurocomputing4
2022 Learn-to-adapt: Concept drift adaptation for hybrid multiple streams
En Yu, Yiliao Song, Guangquan Zhang 0001, Jie Lu 0001
Neurocomputing2
2022 A Drift Region-Based Data Sample Filtering Method
abstract
Concept drift refers to changes in the underlying data distribution of data streams over time. A well-trained model will be outdated if concept drift occurs. Once concept drift is detected, it is necessary to understand where the drift occurs to support the drift adaptation strategy and effectively update the outdated models. This process, called drift understanding, has rarely been studied in this area. To fill this gap, this article develops a drift region-based data sample filtering method to update the obsolete model and track the new data pattern accurately. The proposed method can effectively identify the drift region and utilize information on the drift region to filter the data sample for training models. The theoretical proof guarantees the identified drift region converges uniformly to the real drift region as the sample size increases. Experimental evaluations based on four synthetic datasets and two real-world datasets demonstrate our method improves the learning accuracy when dealing with data streams involving concept drift.
Jie Lu 0001, Yiliao Song, Feng Liu 0003, Guangquan Zhang 0001
IEEE Trans. Cybern.3
2022 A Segment-Based Drift Adaptation Method for Data Streams
abstract
In concept drift adaptation, we aim to design a blind or an informed strategy to update our best predictor for future data at each time point. However, existing informed drift adaptation methods need to wait for an entire batch of data to detect drift and then update the predictor (if drift is detected), which causes adaptation delay. To overcome the adaptation delay, we propose a sequentially updated statistic, called drift-gradient to quantify the increase of distributional discrepancy when every new instance arrives. Based on drift-gradient, a segment-based drift adaptation (SEGA) method is developed to online update our best predictor. Drift-gradient is defined on a segment in the training set. It can precisely quantify the increase of distributional discrepancy between the old segment and the newest segment when only one new instance is available at each time point. A lower value of drift-gradient on the old segment represents that the distribution of the new instance is closer to the distribution of the old segment. Based on the drift-gradient, SEGA retrains our best predictors with the segments that have the minimum drift-gradient when every new instance arrives. SEGA has been validated by extensive experiments on both synthetic and real-world, classification and regression data streams. The experimental results show that SEGA outperforms competitive blind and informed drift adaptation methods.
Yiliao Song, Jie Lu 0001, Anjin Liu, Haiyan Lu, Guangquan Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 An Efficient Bayesian Neural Network for Multiple Data Streams
abstract
Spatial and temporal data such as multiple data streams often have concept drift problems, which refers to changes of the data distributions over time. Once concept drift occurs, a stationary machine learning predictor will probably be invalid because the testing data have a different data distribution from that of the training data. Recent studies in data streams aim to address this issue by concept drift adaptation techniques. Concept drift adaptation methods update the predictor by time and have been validated to provide accurate real-time prediction for a single data stream. However, handling multiple relevant data streams is still an unsolved challenge in this area, considering many real-world applications generate multiple data streams that are simultaneously evolving. To fill this gap, we here present a predicting network for multiple data streams, named MuNet. MuNet leverages the dependency between data streams to efficiently lower down the computational cost of real-time predictions for multiple data streams. In MuNet, an online-learned Bayesian neural network (BNN) is designed as a connector between streams. The BNN-connector can continuously use the real-time information from only one base stream to correct the stationary predictor for the other streams, which avoids the high cost caused by repeated adaptation.
Yiliao Song, Guangquan Zhang 0001, Jie Lu 0001
IJCNN2
2020 A Fuzzy Drift Correlation Matrix for Multiple Data Stream Regression
abstract
How to handle concept drift problem is a big challenge for algorithms designed for the data streams. Currently, techniques related to the concept drift problem focus on single data stream. However, it normally needs to handle multiple relevant data streams in the real-world application. Current concept drift methods can not be directly used in the multistream setting. They can only be limitedly applied on each stream separately, which omits the drift correlation between streams. In the multi-stream scenario, when drift occurs in a stream, other streams may face or have faced a similar drift problem as well. This pattern of simultaneous or delayed occurrence of drift is critical to analyze and predict multiple streams as a whole dynamic system. To fill the gap in the multi-stream scenario, this paper proposes a fuzzy drift variance (FDV) to measure the correlated drift patterns among streams. FDA is able to present how the pattern of drift occurrence for any two streams correlates and how delayed this correlation is. Seven synthetic streams are designed to validate FDA. The experimental results show a good presentation ability of FDA for drift-correlated multiple streams.
Yiliao Song, Guangquan Zhang 0001, Haiyan Lu, Jie Lu 0001
FUZZ-IEEE1
2020 Fuzzy Clustering-Based Adaptive Regression for Drifting Data Streams
abstract
Current models and algorithms have been increasingly required to learn in a nonstationary environment because the phenomenon of concept drift (or pattern shift) may occur, that is, the assumption that data are identically distributed may be invalid in data streams. Once the data pattern changes, a well-trained model built on the previous, now obsolete data cannot provide an accurate prediction for future data. To obtain reliable prediction, it is important to understand the existing patterns in the data stream and to know which pattern the current examples belong to during the modeling process. However, it is ambiguous to classify an example to a certain pattern in many real-world cases. In this paper, we propose a novel adaptive regression approach, called FUZZ-CARE, to dynamically recognize, train, and store patterns, and assign the membership degree of the upcoming examples belonging to these patterns. Membership degrees are presented by the membership matrix obtained from a kernel fuzzy c-means clustering, which is synchronously trained and adapted with regression parameters. Rather than designing a complicated procedure to continuously chase the newest pattern, which is a common approach in the literature, FUZZ-CARE abstracts useful past information to help predict newly arrived examples. It thus effectively avoids the risk of insufficient training due to the lack of new data and improves prediction accuracy. Experiments on six synthetic datasets and 21 real-world datasets validate the high accuracy and robustness of our approach.
Yiliao Song, Jie Lu 0001, Haiyan Lu, Guangquan Zhang 0001
IEEE Trans. Fuzzy Syst.1
2019 A Noise-tolerant Fuzzy c-Means based Drift Adaptation Method for Data Stream Regression
abstract
Concept drift referring to the changes of data distributions has been one critical challenge typically associated with mining data streams. Current drift detection and adaptation methods focus on how to immediately detect the distribution changes once the concept drift occurs and swiftly update the model to be applicable to the newly arrived data instances. Most of those methods assume the data does not have noise or the noise is too weak to affect the modeling procedure. However, realworld data are normally contaminated, and denoise techniques are highly preferred as a necessary preprocess. This issue is more complex for a data stream with concept drift because the noise is very likely to be confused with drift. Motivated by that, this paper proposes a Noise-tolerant Fuzzy c-means based drift Adaptation method (NFA) which can adapt to the changing distributions and is suitable for noisy data streams. The concept drift problem is solved by using a fuzzy c-means based regression model to continuously include the most relevant data instances to the latest pattern in the training set. In addition, a denoise technique is designed in NFA to remove noise, and the ability of incremental updating enables it to be embedded in the incremental drift adaptation process, and therefore NFA can solve concept drift and noise problems at the same time. Experimental evaluation results also show good performance of our method on handling data streams with concept drift and noise.
Yiliao Song, Guangquan Zhang 0001, Haiyan Lu, Jie Lu 0001
FUZZ-IEEE1
2018 A Self-adaptive Fuzzy Network for Prediction in Non-stationary Environments
abstract
Prediction in non-stationary environments, where data streams are ever-changing at very high speeds, has become more and more important in real-world applications. The uncertainty in data streams caused by changes in data distribution is described as concept drift. The appearance of concept drift in a data stream results in inconsistencies between the existing data and incoming data. Such inconsistencies pose a great challenge to conventional machine learning methods, given they are built on the assumption of independent and identically distributed data and cannot adapt to unpredictable changes in knowledge patterns. To solve such data stream uncertainty problem, this paper presents a window-based self-adaptive fuzzy network called adaptive fuzzy network (AFN), which can continuously modify the network through identifying new knowledge from the previous data samples. Three components are embedded in ANF: a drift detection module to identify whether the current window of data samples presents different pattern from the previous; a drift adaption module to retain useful knowledge in previous samples; and a fuzzy inference system, which integrates the detection and adaption modules for prediction. ANF has been evaluated through a set of experiments on non-stationary data streams. The experimental results show a good effectiveness of our method.
Yiliao Song, Guangquan Zhang 0001, Haiyan Lu, Jie Lu 0001
FUZZ-IEEE1
2017 A fuzzy kernel c-means clustering model for handling concept drift in regression
abstract
Concept drift, given the huge volume of high-speed data streams, requires traditional machine learning models to be self-adaptive. Techniques to handle drift are especially needed in regression cases for a wide range of applications in the real world. There is, however, a shortage of research on drift adaptation for regression cases in the literature. One of the main obstacles to further research is the resulting model complexity when regression methods and drift handling techniques are combined. This paper proposes a self-adaptive algorithm, based on a fuzzy kernel c-means clustering approach and a lazy learning algorithm, called FKLL, to handle drift in regression learning. Using FKLL, drift adaptation first updates the learning set using lazy learning, then fuzzy kernel c-means clustering is used to determine the most relevant learning set. Experiments show that the FKLL algorithm is better able to respond to drift as soon as the learning sets are updated, and is also suitable for dealing with reoccurring drift, when compared to the original lazy learning algorithm and other state-of-the-art regression methods.
Yiliao Song, Guangquan Zhang 0001, Jie Lu 0001, Haiyan Lu
FUZZ-IEEE1
2017 Regional Concept Drift Detection and Density Synchronized Drift Adaptation
abstract
In data stream mining, the emergence of new patterns or a pattern ceasing to exist is called concept drift. Concept drift makes the learning process complicated because of the inconsistency between existing data and upcoming data. Since concept drift was first proposed, numerous articles have been published to address this issue in terms of distribution analysis. However, most distribution-based drift detection methods assume that a drift happens at an exact time point, and the data arrived before that time point is considered not important. Thus, if a drift only occurs in a small region of the entire feature space, the other non-drifted regions may also be suspended, thereby reducing the learning efficiency of models. To retrieve non-drifted information from suspended historical data, we propose a local drift degree (LDD) measurement that can continuously monitor regional density changes. Instead of suspending all historical data after a drift, we synchronize the regional density discrepancies according to LDD. Experimental evaluations on three public data sets show that our concept drift adaptation algorithm improves accuracy compared to other methods.
Anjin Liu, Yiliao Song, Guangquan Zhang 0001, Jie Lu 0001
IJCAI2