EDBT 2026 Demo / reviewers in the wild / expert
Fangzhou Zhu
dblp:74/8725
· DBLP profile ↗
17ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
7 papers |
Mathematical optimization · 94% Automated reasoning and model checking · 6% | |
| Artificial intelligence
2 papers |
Planning, search and constraint satisfaction · 44% Optimization for machine learning · 44% Language models and text generation · 13% | |
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 50% Spatial and temporal data management · 50% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization › discrete optimization
mixed integer linear programming |
2.5 | 3 | 2025 | Dynamic Configuration for Cutting Plane Separators via Reinforcement Learning on Incremental Graph · NeurIPS 2025 Apollo-MILP: An Alternating Prediction-Correction Neural Solving Framework for Mixed-Integer Linear Programming · ICLR 2025 Learning to Cut via Hierarchical Sequence/Set Model for Efficient Mixed-Integer Programming · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Mathematical optimization › black-box optimization
solver configuration |
1.7 | 2 | 2025 | Accelerate Presolve in Large-Scale Linear Programming via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Dynamic Configuration for Cutting Plane Separators via Reinforcement Learning on Incremental Graph · NeurIPS 2025 |
Mathematical optimization
combinatorial optimization |
1.5 | 2 | 2024 | Towards General Algorithm Discovery for Combinatorial Optimization: Learning Symbolic Branching Policy from Bipartite Graph · ICML 2024 Rethinking Branching on Exact Combinatorial Optimization Solver: The First Deep Symbolic Discovery Framework · ICLR 2024 |
Machine learning › Optimization for machine learning
learned optimizer |
0.9 | 1 | 2025 | Apollo-MILP: An Alternating Prediction-Correction Neural Solving Framework for Mixed-Integer Linear Programming · ICLR 2025 |
Mathematical optimization › optimization for machine learning
differentiable optimization |
0.9 | 1 | 2025 | Differentiable Integer Linear Programming · ICLR 2025 |
Mathematical optimization
integer programming |
0.9 | 1 | 2025 | Differentiable Integer Linear Programming · ICLR 2025 |
Mathematical optimization
learning to optimize |
0.9 | 1 | 2025 | Differentiable Integer Linear Programming · ICLR 2025 |
Mathematical optimization
linear programming |
0.9 | 1 | 2025 | Accelerate Presolve in Large-Scale Linear Programming via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Mathematical optimization
presolving |
0.9 | 1 | 2025 | Accelerate Presolve in Large-Scale Linear Programming via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Mathematical optimization › discrete optimization › mixed integer linear programming
branching |
0.8 | 1 | 2024 | Rethinking Branching on Exact Combinatorial Optimization Solver: The First Deep Symbolic Discovery Framework · ICLR 2024 |
Automated reasoning and model checking › satisfiability › SAT solving
branching heuristic |
0.8 | 1 | 2024 | Towards General Algorithm Discovery for Combinatorial Optimization: Learning Symbolic Branching Policy from Bipartite Graph · ICML 2024 |
Mathematical optimization
discrete optimization |
0.8 | 1 | 2024 | Rethinking Branching on Exact Combinatorial Optimization Solver: The First Deep Symbolic Discovery Framework · ICLR 2024 |
Cellular and mobile networks
mobility management |
0.5 | 1 | 2021 | A Data-Driven Sequential Localization Framework for Big Telco Data · IEEE Trans. Knowl. Data Eng. 2021 |
Natural language and speech › Language models and text generation › large language model reasoning
tree of thought reasoning |
0.3 | 1 | 2025 | BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem Solving · ACL (1) 2025 |
Data mining › predictive analytics
churn prediction |
0.2 | 1 | 2015 | Telco Churn Prediction with Big Data · SIGMOD Conference 2015 |
Data mining
predictive analytics |
0.2 | 1 | 2015 | Telco Churn Prediction with Big Data · SIGMOD Conference 2015 |
Data mining
big data analytics |
0.1 | 1 | 2015 | Telco Churn Prediction with Big Data · SIGMOD Conference 2015 |
Methods — techniques the papers use, named apart from their topics
transformer · 2.4graph neural network · 2.4uncertainty-based error bound · 1.7trust-region search · 1.7neural prediction-correction · 1.7reinforcement learning · 1.7materialized views · 1.0indexing · 1.0unsupervised learning · 0.9tree-of-thoughts · 0.9search · 0.9probabilistic modeling · 0.9incremental graph · 0.9gradient descent · 0.9neural network search · 0.8knowledge distillation · 0.8deep symbolic discovery · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem SolvingabstractTeng Wang, Wing Yin Yu, Zhenqi He, Zehua Liu, HaileiGong HaileiGong, Han Wu, Xiongwei Han, Wei Shi, Ruifeng She, Fangzhou Zhu, Tao Zhong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wing Yin Yu, Zhenqi He, Zehua Liu, HaileiGong HaileiGong, Han Wu 0004, Xiongwei Han, Ruifeng She, Fangzhou Zhu, Tao Zhong 0004 |
ACL (1) | 10 |
| 2025 | Apollo-MILP: An Alternating Prediction-Correction Neural Solving Framework for Mixed-Integer Linear ProgrammingabstractLeveraging machine learning (ML) to predict an initial solution for mixed-integer linear programming (MILP) has gained considerable popularity in recent years. These methods predict a solution and fix a subset of variables to reduce the problem dimension. Then, they solve the reduced problem to obtain the final solutions. However, directly fixing variable values can lead to low-quality solutions or even infeasible reduced problems if the predicted solution is not accurate enough. To address this challenge, we propose an Alternating prediction-correction neural solving framework (Apollo-MILP) that can identify and select accurate and reliable predicted values to fix. In each iteration, Apollo-MILP conducts a prediction step for the unfixed variables, followed by a correction step to obtain an improved solution (called reference solution) through a trust-region search. By incorporating the predicted and reference solutions, we introduce a novel Uncertainty-based Error upper BOund (UEBO) to evaluate the uncertainty of the predicted values and fix those with high confidence. A notable feature of Apollo-MILP is the superior ability for problem reduction while preserving optimality, leading to high-quality final solutions. Experiments on commonly used benchmarks demonstrate that our proposed Apollo-MILP significantly outperforms other ML-based approaches in terms of solution quality, achieving over a 50% reduction in the solution gap. Haoyang Liu 0002, Jie Wang 0005, Zijie Geng, Xijun Li, Yuxuan Zong, Fangzhou Zhu, Jianye Hao, Feng Wu 0001 |
ICLR | 6 |
| 2025 | Differentiable Integer Linear ProgrammingabstractMachine learning (ML) techniques have shown great potential in generating high-quality solutions for integer linear programs (ILPs).
However, existing methods typically rely on a *supervised learning* paradigm, leading to (1) *expensive training cost* due to repeated invocations of traditional solvers to generate training labels, and (2) *plausible yet infeasible solutions* due to the misalignment between the training objective (minimizing prediction loss) and the inference objective (generating high-quality solutions).
To tackle this challenge, we propose **DiffILO** (**Diff**erentiable **I**nteger **L**inear Programming **O**ptimization), an *unsupervised learning paradigm for learning to solve ILPs*.
Specifically, through a novel probabilistic modeling, DiffILO reformulates ILPs---discrete and constrained optimization problems---into continuous, differentiable (almost everywhere), and unconstrained optimization problems.
This reformulation enables DiffILO to simultaneously solve ILPs and train the model via straightforward gradient descent, providing two major advantages.
First, it significantly reduces the training cost, as the training process does not need the aid of traditional solvers at all.
Second, it facilitates the generation of feasible and high-quality solutions, as the model *learns to solve ILPs* in an end-to-end manner, thus aligning the training and inference objectives.
Experiments on commonly used ILP datasets demonstrate that DiffILO not only achieves an average training speedup of $13.2$ times compared to supervised methods, but also outperforms them by generating heuristic solutions with significantly higher feasibility ratios and much better solution qualities. Zijie Geng, Jie Wang 0005, Xijun Li, Fangzhou Zhu, Jianye Hao, Bin Li 0025, Feng Wu 0001 |
ICLR | 4 |
| 2025 | Dynamic Configuration for Cutting Plane Separators via Reinforcement Learning on Incremental GraphabstractCutting planes (cuts) are essential for solving mixed-integer linear programming (MILP) problems, as they tighten the feasible solution space and accelerate the solving process. Modern MILP solvers offer diverse cutting plane separators to generate cuts, enabling users to leverage their potential complementary strengths to tackle problems with different structures. Recent machine learning approaches learn to configure separators based on problem-specific features, selecting effective separators and deactivating ineffective ones to save unnecessary computing time. However, they ignore the dynamics of separator efficacy at different stages of cut generation and struggle to adapt the configurations for the evolving problems after multiple rounds of cut generation. To address this challenge, we propose a novel **dyn**amic **sep**arator configuration (**DynSep**) method that models separator configuration in different rounds as a reinforcement learning task, making decisions based on an incremental triplet graph updated by iteratively added cuts. Specifically, we tokenize the incremental subgraphs and utilize a decoder-only Transformer as our policy to autoregressively predict when to halt separation and which separators to activate at each round. Evaluated on synthetic and large-scale real-world MILP problems, DynSep speeds up average solving time by 64% on easy and medium datasets, and reduces primal-dual gap integral within the given time limit by 16% on hard datasets. Moreover, experiments demonstrate that DynSep well generalizes to MILP instances of significantly larger sizes than those seen during training. Mingxuan Ye, Jie Wang 0005, Fangzhou Zhu, Yufei Kuang, Xijun Li, Weilin Luo, Jianye Hao, Feng Wu 0001 |
NeurIPS | 3 |
| 2025 | Accelerate Presolve in Large-Scale Linear Programming via Reinforcement LearningabstractAs one of the most critical components in modern LP solvers, presolve in linear programming (LP) employs a rich set of presolvers to remove different types of redundancy in input problems by equivalent transformations. We found from extensive experiments that the presolve routine-that is, the method determining (P1) which presolvers to select, (P2) in what order to execute, and (P3) when to stop-significantly impacts the efficiency of solving LPs. However, designing high-quality presolve routines is highly challenging due to the enormous search space, and further optimizing the routines on different tasks for high performance demands extensive domain knowledge and manual tuning. To tackle this problem, we propose the first learning based framework-that is, reinforcement learning for presolve (RL4Presolve)-to learn high-quality presolve routines. An appealing feature is that we employ a novel adaptive action sequence that learns complex routines efficiently by generating combinations of presolvers automatically at each step. Extensive experiments demonstrate that RL4Presolve achieves significant improvement (up to roughly 90% ) in the efficiency of solving LPs. Furthermore, we extract routines from learned policies for simple and efficient deployment without GPU resources to Huawei's supply chain, where extensive manual tuning for each separate task was required previously due to the high economic value. Yufei Kuang, Xijun Li, Jie Wang 0005, Fangzhou Zhu, Houqiang Li, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Rethinking Branching on Exact Combinatorial Optimization Solver: The First Deep Symbolic Discovery FrameworkabstractMachine learning (ML) has been shown to successfully accelerate solving NP-hard combinatorial optimization (CO) problems under the branch and bound framework.
However, the high training and inference cost and limited interpretability of ML approaches severely limit their wide application to modern exact CO solvers. In contrast, human-designed policies---though widely integrated in modern CO solvers due to their compactness and reliability---can not capture data-driven patterns for higher performance. To combine the advantages of the two paradigms, we propose the first symbolic discovery framework---namely, deep symbolic discovery for exact combinatorial optimization solver (Symb4CO)---to learn high-performance symbolic policies on the branching task. Specifically, we show the potential existence of small symbolic policies empirically, employ a large neural network to search in the high-dimensional discrete space, and compile the learned symbolic policies directly for fast deployment. Experiments show that the Symb4CO learned purely CPU-based policies consistently achieve *comparable* performance to previous GPU-based state-of-the-art approaches.
Furthermore, the appealing features of Symb4CO include its high training (*ten training instances*) and inference (*one CPU core*) efficiency and good interpretability (*one-line expressions*), making it simple and reliable for deployment. The results show encouraging potential for the *wide* deployment of ML to modern CO solvers. Yufei Kuang, Jie Wang 0005, Haoyang Liu 0002, Fangzhou Zhu, Xijun Li, Jianye Hao, Bin Li 0025, Feng Wu 0001 |
ICLR | 4 |
| 2024 | Towards General Algorithm Discovery for Combinatorial Optimization: Learning Symbolic Branching Policy from Bipartite GraphabstractMachine learning (ML) approaches have been successfully applied to accelerating exact combinatorial optimization (CO) solvers. However, many of them fail to explain what patterns they have learned that accelerate the CO algorithms due to the black-box nature of ML models like neural networks, and thus they prevent researchers from further understanding the tasks they are interested in. To tackle this problem, we propose the first graph-based algorithm discovery framework—namely, graph symbolic discovery for exact combinatorial optimization solver (GS4CO)—that learns interpretable branching policies directly from the general bipartite graph representation of CO problems. Specifically, we design a unified representation for symbolic policies with graph inputs, and then we employ a Transformer with multiple tree-structural encodings to generate symbolic trees end-to-end, which effectively reduces the cumulative error from iteratively distilling graph neural networks. Experiments show that GS4CO learned interpretable and lightweight policies outperform all the baselines on CPU machines, including both the human-designed and the learning-based. GS4CO shows an encouraging step towards general algorithm discovery on modern CO solvers. Yufei Kuang, Jie Wang 0005, Yuyan Zhou, Xijun Li, Fangzhou Zhu, Jianye Hao, Feng Wu 0001 |
ICML | 5 |
| 2024 | Multidimensional Fingerprints-Based Multiattacker Detection for 6G SystemsabstractThe future 6G systems are expected to achieve intelligent connection and interaction between various heterogeneous terminals, increasing the fragility for spoofing attacks. Due to the high security and energy efficiency, Physical Layer Authentication (PLA) has been regarded as a powerful method to verify the identity of devices. Nevertheless, due to the inaccurate identifying fingerprints caused by the imperfect estimation and variations of the limited fingerprints, most of the state-of-the-art PLA schemes have low reliability and robustness in low Signal-Noise Ratio (SNR) environments. Besides, most PLA schemes rely on the prior knowledge of attackers to establish authentication models, thus reducing the feasibility of actual communications. To address the first challenge, we propose a multi-attacker detection architecture based on multi-dimensional fingerprints, which can provide more robust identifiable spatial attributes for devices by using fingerprints observed by receivers in multi-locations. Upon the designed detection architecture, to tackle the second issue, we propose four clustering-based PLA schemes without requiring their training fingerprint sets. Considering that the aforementioned schemes can divide fingerprints from different transmitters into several disjoint clusters but can not precisely identify forged fingerprints, we further propose the graph learning-based PLA approaches with only a few labeled fingerprints. The simulation results on real industrial outdoor and indoor datasets demonstrate the superiority of the designed detection system in Adjusted Mutual Information (AMI) and authentication accurate rate (AucRate) over the single observation-based PLA schemes. Xiaodong Xu 0001, Gangyi Li, Bingxuan Xu, Fangzhou Zhu, Bizhu Wang, Ping Zhang 0003 |
IEEE Internet Things J. | 5 |
| 2024 | Learning to Cut via Hierarchical Sequence/Set Model for Efficient Mixed-Integer ProgrammingabstractCutting planes (cuts) play an important role in solving mixed-integer linear programs (MILPs), which formulate many important real-world applications. Cut selection heavily depends on (P1) which cuts to prefer and (P2) how many cuts to select. Although modern MILP solvers tackle (P1)-(P2) by human-designed heuristics, machine learning carries the potential to learn more effective heuristics. However, many existing learning-based methods learn which cuts to prefer, neglecting the importance of learning how many cuts to select. Moreover, we observe that (P3) what order of selected cuts to prefer significantly impacts the efficiency of MILP solvers as well. To address these challenges, we propose a novel hierarchical sequence/set model (HEM) to learn cut selection policies. Specifically, HEM is a bi-level model: (1) a higher-level module that learns how many cuts to select, (2) and a lower-level module-that formulates the cut selection as a sequence/set to sequence learning problem-to learn policies selecting an ordered subset with the cardinality determined by the higher-level module. To the best of our knowledge, HEM is the first data-driven methodology that well tackles (P1)-(P3) simultaneously. Experiments demonstrate that HEM significantly improves the efficiency of solving MILPs on eleven challenging MILP benchmarks, including two Huawei's real problems. Jie Wang 0005, Xijun Li, Yufei Kuang, Zhihao Shi, Fangzhou Zhu, Mingxuan Yuan, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | A Data-Driven Sequential Localization Framework for Big Telco DataabstractThe proliferation of telco networks and mobile terminals brings the accumulation of tremendous amounts of measure report(MR) data at a rapid pace. The MR data is generated by mobile objects while connecting to data services and is stored in backend data centers. To geo-tag or localize such MR data is believed to have a profound effect on the analytics and optimizations of telco and traffic networks. However, MR records are of noisy and partial observations regarding to mobile objects' geo-locations and hence pose challenges to accurate telco data localization. There have been quite a few attempts. Single-point localization methods map a MR record to a location, but come out with limited accuracies due to the ignorance of spatiotemporal coherence of successive MR records. Recent efforts on sequential localization techniques alleviate this by mapping a sequence of MR records to a trajectory. However, existing solutions are often with assumptions on specific models, e.g., mobility and signal strength distributions, or priori knowledge on topology space, e.g., road networks, limiting the deployment in practice. To this end, we propose a data-driven framework to tackle the challenges in sequential telco localization. We solely use raw MR records and a public third-party GPS dataset for the learning of the correlations between mobile objects' locations and MR records, requiring no model assumptions and priori knowledge. To handle the data-intensive workloads during the learning process, we use materialized views for efficient online localization and light-weighted indexing techniques for periodical parameters tuning, in order to improve the efficiency and scalability. Results on real data show that our solution achieves 58.8 percent improvement in median localization errors compared with state-of-art sequential localization techniques that require hypothesis models and priori knowledge, making our solution superior in terms of effectiveness, efficiency, and employability. Fangzhou Zhu, Mingxuan Yuan, Xike Xie, Shenglin Zhao, Weixiong Rao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | SCAFFISD: A Scalable Framework for Fine-grained Identification and Security Detection of Wireless RoutersabstractThe security of wireless network devices has received widespread attention, but most existing schemes cannot achieve fine-grained device identification. In practice, the security vulnerabilities of a device are heavily depending on its model and firmware version. Motivated by this issue, we propose a universal, extensible and device-independent framework called SCAFFISD, which can provide fine-grained identification of wireless routers. It can generate access rules to extract effective information from the router admin page automatically and perform quick scans for known device vulnerabilities. Meanwhile, SCAFFISD can identify rogue access points (APs) in combination with existing detection methods, with the purpose of performing a comprehensive security assessment of wireless networks. We implement the prototype of SCAFFISD and verify its effectiveness through security scans of actual products. Fangzhou Zhu, Liang Liu 0006, Weizhi Meng 0001, Ting Lv, Renjun Ye |
TrustCom | 1 |
| 2017 | Experimental Study of Telco Localization MethodsabstractTelecommunication (Telco) localization is a technique to accurately locate mobile devices (MDs) using measurement report (MR) data, and has been widely used in Telco industry. Many techniques have been proposed, including measurement-based statistical algorithms, fingerprinting algorithms and different machine learning-based algorithms. However, it has not been well studied yet on how these algorithms perform on various Telco MR data sets. In this paper, we conduct a comprehensive experimental study of five state-of-art algorithms for Telco localization. Based on real data sets from two Telco networks, we study the localization performance of such algorithms. We find that a Random Forest-based machine learning algorithm performs best in most experiments due to high localization accuracy and insensitivity to data volume. The experimental result and observation in this paper may inspire and enhance future research in Telco localization. Weixiong Rao, Fangzhou Zhu, Mingxuan Yuan |
MDM | 3 |
| 2016 | City-Scale Localization with Telco Big DataabstractIt is still challenging in telecommunication (telco) industry to accurately locate mobile devices (MDs) at city-scale using the measurement report (MR) data, which measure parameters of radio signal strengths when MDs connect with base stations (BSs) in telco networks for making/receiving calls or mobile broadband (MBB) services. In this paper, we find that the widely-used location based services (LBSs) have accumulated lots of over-the-top (OTT) global positioning system (GPS) data in telco networks, which can be automatically used as training labels for learning accurate MR-based positioning systems. Benefiting from these telco big data, we deploy a context-aware coarse-to-fine regression (CCR) model in Spark/Hadoop-based telco big data platform for city-scale localization of MDs with two novel contributions. First, we design map-matching and interpolation algorithms to encode contextual information of road networks. Second, we build a two-layer regression model to capture coarse-to-fine contextual features in a short time window for improved localization performance. In our experiments, we collect 108 GPS-associated MR records in the centroid of Shanghai city with 12 x 11 square kilometers for 30 days, and measure four important properties of real-world MR data related to localization errors: stability, sensitivity, uncertainty and missing values. The proposed CCR works well under different properties of MR data and achieves a mean error of 110m and a median error of $80m$, outperforming the state-of-art range-based and fingerprinting localization methods. Fangzhou Zhu, Chen Luo 0003, Mingxuan Yuan, Yijian Zhu, Zhengqing Zhang, Tao Gu 0001, Weixiong Rao |
CIKM | 1 |
| 2015 | Telco Churn Prediction with Big DataabstractWe show that telco big data can make churn prediction much more easier from the $3$V's perspectives: Volume, Variety, Velocity. Experimental results confirm that the prediction performance has been significantly improved by using a large volume of training data, a large variety of features from both business support systems (BSS) and operations support systems (OSS), and a high velocity of processing new coming data. We have deployed this churn prediction system in one of the biggest mobile operators in China. From millions of active customers, this system can provide a list of prepaid customers who are most likely to churn in the next month, having $0.96$ precision for the top $50000$ predicted churners in the list. Automatic matching retention campaigns with the targeted potential churners significantly boost their recharge rates, leading to a big business value. Fangzhou Zhu, Mingxuan Yuan, Bing Ni, Wenyuan Dai, Qiang Yang 0001 |
SIGMOD Conference | 2 |
| 2013 | Probabilistic nearest neighbor queries of uncertain data via wireless data broadcast
Fangzhou Zhu, Guohui Li 0001, Xiaosong Zhao, Cong Zhang 0007 |
Peer-to-Peer Netw. Appl. | 1 |
| 2011 | Active Learning to Defend Poisoning Attack against Semi-Supervised Intrusion Detection ClassifierabstractIntrusion detection systems play an important role in computer security. To make intrusion detection systems adaptive to changing environments, supervised learning techniques had been applied in intrusion detection. However, supervised learning needs a large amount of training instances to obtain classifiers with high accuracy. Limited to lack of high quality labeled instances, some researchers focused on semi-supervised learning to utilize unlabeled instances enhancing classification. But involving the unlabeled instances into the learning process also introduces vulnerability: attackers can generate fake unlabeled instances to mislead the final classifier so that a few intrusions can not be detected. In this paper we show that the attacker could mislead the semi-supervised intrusion detection classifier by poisoning the unlabeled instances. And we propose a defend method based on active learning to defeat the poisoning attack. Experiments show that the poisoning attack can reduce the accuracy of the semi-supervised learning classifier and the proposed defending method based on active learning can obtain higher accuracy than the original semi-supervised learner under the presented poisoning attack. Fangzhou Zhu, Zhiping Cai |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 3 |
| 2010 | A Misleading Attack against Semi-supervised Learning for Intrusion Detection
Fangzhou Zhu, Zhiping Cai |
MDAI | 1 |