VLDB 2026 Research / reviewers in the wild / expert
Haizhou Du
dblp:48/10264
· DBLP profile ↗
42ranked-venue papers
27as first author
38since 2021 · last 2026
0000-0002-1875-5159ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 15 first-author · 17 since 2021Computer networks · 13 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 6 since 2021Systems, architecture and hardware · 5 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mnemosyne: Accelerating Multi-Hop Question Answering via Cache Hit Order FittingabstractMulti-Hop Question Answering (MHQA) requires step-by-step reasoning across multiple pieces of information to answer complex questions. The cache-aided Retrieval-Augmented Generation (RAG) can accelerate the process of external knowledge retrieval at each reasoning step for MHQA. However, existing methods focus on the internal structure and ignore the misalignment between the queries’ arrival order and cache hit order. To tackle this, we propose Mnemosyne, a cache hit order fitting method designed to accelerate the RAG progress for MHQA. Specifically, our cache-aware order fitting strategy adjusts the order of queries arrival via graph reordering to better align with the cache hit order, thereby reducing the likelihood of failed or unproductive retrieval attempts. The multi-granularity caching storage mechanism is designed to loosen the strict hit condition to multiple similar semantic matching modes, facilitating that relevant documents can still be retrieved. Experiments conducted on four multi-hop QA datasets demonstrate that Mnemosyne effectively reduces retrieval latency while enhancing task answer F1 score, achieving a superior trade-off between efficiency and effectiveness. Haizhou Du, Jiujiu Li, Luobin Huang, Lisheng Wang |
AAAI | 1 |
| 2026 | FedQCab: A Common Plug-and-Play Layer-wise Quantization Tool for Communication-Efficient Federated Learning
Haizhou Du, Tonghao Chen |
IWQoS | 1 |
| 2026 | DNSMedic: Automated Precision Diagnosis and Repair Guidance for DNS Misconfigurations
Kaiqiang Hu, Mochun Long, Haizhou Du, Qiao Xiang, Mengrui Zhang, Letian Zhu |
IWQoS | 5 |
| 2026 | M3RAG: Orchestrating Multi-agent Reasoning for Multi-hop, Multi-modal Understanding
Haizhou Du |
MMM (1) | 1 |
| 2026 | SemDNS: A Declarative Semantics for DNS Resolution
Kaiqiang Hu, Haizhou Du, Xiaoliang Wang 0001 |
SIGCOMM | 2 |
| 2026 | FedRGL: Federated Riemannian Graph Learning in Mixed-Curvature Spaces with Ricci-Gated ConvolutionabstractFederated Graph Learning (FGL) has emerged as an efficient paradigm to address the pronounced data heterogeneity common in real-world decentralized graph datasets, attracting significant interest from both academia and industry. Existing FGL methods often degrade the global model's performance due to a fundamental mismatch between the model's fixed-geometry embedding space and the diverse geometric structures of client data. To address this issue, we propose a novel framework for Personalized Federated Riemannian Graph Learning, namely FedRGL. FedRGL introduces a personalized mixed-curvature product space for each heterogeneous client, mapping local graph data into a tailored geometric space composed of Euclidean, hyperbolic, and spherical manifolds. Furthermore, it leverages a Ricci-Gated Graph Convolutional Network to dynamically adapt its message-passing mechanism to the local topology of each graph. Extensive experiments across diverse geometric heterogeneity settings demonstrate that FedRGL significantly outperforms state-of-the-art FGL methods in terms of model accuracy and generalization. Haizhou Du, Zicheng Shi |
WWW | 1 |
| 2026 | VeriDNS: incremental distributed verification of DNS configurations
Haizhou Du, Kaiqiang Hu, Yao Wang 0022 |
Comput. Networks | 1 |
| 2026 | FedHyperGraph: A layer-wise personalized federated learning with correlation graphs in hyperbolic space
Haizhou Du, Chongyi Qiu, Huan Huo |
Neural Networks | 1 |
| 2025 | A Communication-Efficient Paradigm for Decentralized Machine Learning with Model Logits
Yiwen Cai, Haizhou Du |
APNet | 2 |
| 2025 | Fedcab: A Novel Communication Compression Method for Efficient Federated Learning under Dynamic Bandwidth
Tonghao Chen, Haizhou Du |
APNet | 2 |
| 2025 | DNS-GraphSentry: A Robust Graph Neural Network Approach for Automated Misconfiguration Diagnosis
Kaiqiang Hu, Haizhou Du |
APNet | 2 |
| 2025 | Decoupling Feature Entanglement for Personalized Federated Learning via Neural Collapse
Haizhou Du |
CIKM | 1 |
| 2025 | FedHigh: A Graph-Guided Aggregation Framework for Layer-Wised Personalized Federated LearningabstractThe goal of Personalized Federated Learning (PFL) achieves model personalization on statistical heterogeneous scenarios. Existing PFL approaches roughly treat the entire model parameters as a whole unit for generating personalized weights, and neglect the latent correlations among parameters of different client models, which is crucial for model personalization by the same learning tasks on statistical heterogeneity. To solve this issue, we propose a graph-guided aggregation framework for layer-wised PFL, namely FedHigh. The framework leverages layered knowledge sharing of client models to generate latent correlation graphs in the hyperbolic space. The latent correlation graphs can efficiently guide the personalized aggregation. Experimental results demonstrate that FedHigh achieves significant accuracy improvements up to 42.6%, 33.1%, and 7.5% on graph, computer vision (CV), and natural language processing (NLP) datasets, respectively. Meanwhile, FedHigh accelerates convergence speed up to 52.6% compared to state-of-the-art baselines on CV datasets. Additionally, the scalability of FedHigh also attractively outperforms other baselines with different world sizes. Haizhou Du, Chongyi Qiu, Huan Huo |
ECAI | 1 |
| 2025 | FedHyperClass: Boosting Cross-Modal Consensus for Federated Learning with Unimodal Clients
Haizhou Du, Chongyi Qiu |
ICA3PP (2) | 1 |
| 2025 | DNSHolmes: A Scalable Framework for DNS Repair Using Large Language ModelsabstractWith the widespread application of large language models (LLMs), their advantages in context analysis and semantic inferencing are becoming increasingly prominent, and we are attempting to introduce them into the realm of DNS repair. As is well known, DNS, as the core of Internet infrastructure, has complex policies and a fragile system, where even a small misconfiguration can lead to catastrophic service failures. Especially in large-scale networks, analyzing detected errors and generating repair solutions often requires operators to invest a significant amount of time and effort. This paper proposes a scalable framework, DNSHolmes, designed to leverage LLMs for generating DNS configuration repair solutions. Specifically, this method first addresses the numerous potential root causes in large-scale errors by abstracting them into a small number of State Equivalence Classes (SECs). It then adopts a deterministic finite automaton (DFA) to compute a Critical Path Graph (CPG) for each class, precisely isolating the minimal set of records responsible for the failure. Crucially, the CPG serves as a focused, verifiable context for a Large Language Model (LLM), guiding it to generate accurate patches while overcoming the fundamental context-length limitations that help LLM with effective repair reasoning. Our evaluation on large-scale public datasets and a real-world campus network demonstrates that DNSHolmes can reduce operational effort by 78.4% and achieve efficient repair in large-scale DNS configuration. Kaiqiang Hu, Haizhou Du |
IPCCC | 2 |
| 2025 | FedFree: Breaking Knowledge-sharing Barriers through Layer-wise Alignment in Heterogeneous Federated LearningabstractHeterogeneous Federated Learning (HtFL) enables collaborative learning across clients with diverse model architectures and non-IID data distributions, which are prevalent in real-world edge computing applications. Existing HtFL approaches typically employ proxy datasets to facilitate knowledge sharing or implement coarse-grained model-level knowledge transfer. However, such approaches not only elevate risks of user privacy leakage but also lead to the loss of fine-grained model-specific knowledge, ultimately creating barriers to effective knowledge sharing. To address these challenges, we propose FedFree, a novel data-free and model-free HtFL framework featuring two key innovations. First, FedFree introduces a reverse layer-wise knowledge transfer mechanism that aggregates heterogeneous client models into a global model solely using Gaussian-based pseudo data, eliminating reliance on proxy datasets. Second, it leverages Knowledge Gain Entropy (KGE) to guide targeted layer-wise knowledge alignment, ensuring that each client receives the most relevant global updates tailored to its specific architecture. We provide rigorous theoretical convergence guarantees for FedFree and conduct extensive experiments on CIFAR-10 and CIFAR-100. Results demonstrate that FedFree achieves substantial performance gains, with relative accuracy improving up to 46.3% over state-of-the-art baselines. The framework consistently excels under highly heterogeneous model/data distributions and in large scale settings. Haizhou Du, Yiran Xiang, Yiwen Cai, Xiufeng Liu 0001, Zonghan Wu, Huan Huo, Guodong Long |
NeurIPS | 1 |
| 2025 | Toward Verifying and Interpreting Learning-Based Networking Systems With SMTabstractThere has been a growing interest in applying machine learning to real-world tasks. However, due to the black-box nature of machine learning models, it is crucial to 1) verify important properties of a model and 2) understand the reasons behind a model’s prediction before deploying them in a production environment. Existing approaches typically handle them as two separate and sometimes orthogonal topics. In this paper, we show that the verification and interpretability of machine learning models are tightly related and can be unified by satisfiability modulo theories (SMT). Our key insight is: not only a wide range of properties of machine learning models can be formulated as SMT problems and verified accordingly, but many commonly studied interpretability questions can also be answered by iteratively checking the satisfiability and related properties of multiple SMT problems. Leveraging this insight, we design UINT, a general verification and interpretability framework for learning-based networking systems. UINT 1) allows operators to specify verification and interpretability problems as SMT formulas, 2) encodes the target machine learning models into SMT constraints, and 3) automatically simplifies and solves the corresponding verification and interpretability problems using commodity SMT solvers. We implement a prototype of UINT and evaluate it on real-world learning-based networking systems. Results demonstrate the efficiency and efficacy of UINT in verifying and interpreting key questions for these systems. Yuling Lin, Yangfan Huang, Haizhou Du, Qiao Xiang, Yijian Chen, Linghe Kong, Qiang Li 0045, Franck Le, Jiwu Shu |
IEEE Trans. Netw. | 4 |
| 2024 | Rethinking DNS Configuration Verification with a Distributed ArchitectureabstractDNS misconfiguration can result in severe social and financial consequences. Existing DNS configuration verification tools employ a centralized architecture, where all zone files are collected for verification. This architecture faces significant scalability issues (e.g., the verifier becoming the performance bottleneck and not supporting incremental verification). Inspired by the recent proposal of distributed data plane verification and the resemblance between the network data plane and DNS configuration, we propose to rearchitect DNS configuration verification with a distributed design. Our key insight is that by analyzing the query processing behavior of each DNS zone file in parallel and stitching the results in a symbolic way, we can substantially scale up the verification of DNS configuration. Evaluation shows that an up to 9.51× speed up on a dataset with over 410,000 resource records while having small overhead. Yao Wang 0022, Kaiqiang Hu, Haizhou Du, Qiao Xiang, Ruiting Zhou, Linghe Kong, Jiwu Shu |
APNet | 5 |
| 2024 | FlowSeer: A Novel Framework for Generalized Network Performance Estimation at Flow LevelabstractNetwork performance estimation is a key enabler in achieving efficient network operation. It can assist in important network tasks such as capacity planning, topology design, and parameter tuning. The traditional methods often rely on discrete event simulation and modeling by queuing theory. However, they consider some ideal scenarios with assumptions at a packet level and hardly generalize with ideal assumptions to the real world. In this paper, we propose FlowSeer, a novel framework for generalized network performance estimation at a flow level. FlowSeer employs a graph neural network to catch topology features and a pioneering transformer to learn and model flow features. Furthermore, the FlowSeer framework adopts a two-stage feature extraction mechanism, which can improve accuracy in large-scale networks. Compared with the state-of-the-art models, our extensive experiments demonstrate FlowSeer can improve the MAPE by up to 65.18% in diverse scenarios. Haizhou Du |
CSCWD | 1 |
| 2024 | HyperPrism: An Adaptive Non-linear Aggregation Framework for Distributed Machine Learning over Non-IID Data and Time-varying Communication LinksabstractWhile Distributed Machine Learning (DML) has been widely used to achieve decent performance, it is still challenging to take full advantage of data and devices distributed at multiple vantage points to adapt and learn, especially it is non-trivial to address dynamic and divergence challenges based on the linear aggregation framework as follows: (1) heterogeneous learning data at different devices (i.e., non-IID data) resulting in model divergence and (2) in the case of time-varying communication links, the limited ability for devices to reconcile model divergence. In this paper, we contribute a non-linear class aggregation framework HyperPrism that leverages distributed mirror descent with averaging done in the mirror descent dual space and adapts the degree of Weighted Power Mean (WPM) used in each round. Moreover, HyperPrism could adaptively choose different mapping for different layers of the local model with a dedicated hypernetwork per device, achieving automatic optimization of DML in high divergence settings. We perform rigorous analysis and experimental evaluations to demonstrate the effectiveness of adaptive, mirror-mapping DML. In particular, we extend the generalizability of existing related works and position them as special cases within HyperPrism. Our experimental results show that HyperPrism can improve the convergence speed up to 98.63% and scale well to more devices compared with the state-of-the-art, all with little additional computation overhead compared to traditional linear aggregation. Haizhou Du, Yijian Chen, Ryan Yang, Yuchen Li 0006, Linghe Kong |
NeurIPS | 1 |
| 2024 | FedZipper: A Layer-wise Quantization Compression Framework for Federated Learning with Statistical Heterogeneity
Haizhou Du |
NPC (2) | 1 |
| 2024 | FedPrime: An Adaptive Critical Learning Periods Control Framework for Efficient Federated Learning in Heterogeneity Scenarios
Haizhou Du |
ECML/PKDD (5) | 1 |
| 2024 | A unified momentum-based paradigm of decentralized SGD for non-convex models and heterogeneous data
Haizhou Du, Chaoqian Cheng, Chengdong Ni |
Artif. Intell. | 1 |
| 2024 | FedSwarm: An Adaptive Federated Learning Framework for Scalable AIoTabstractFederated learning (FL) is a key solution for datadriven the Artificial Intelligence of Things (AIoT). Although much progress has been made, scalability remains a core challenge for real-world FL deployments. Existing solutions either suffer from accuracy loss or do not fully address the connectivity dynamicity of FL systems. In this article, we tackle the scalability issue with a novel, adaptive FL framework called FedSwarm, which improves system scalability for AIoT by deploying multiple collaborative edge servers. FedSwarm has two novel features: 1) adaptiveness on the number of local updates and 2) dynamicity of the synchronization between edge devices and edge servers. We formulate FedSwarm as a local update adaptation and perdevice dynamic server selection problem and prove FedSwarm‘s convergence bound. We further design a control mechanism consisting of a learning-based algorithm for collaboratively providing local update adaptation on the servers’ side and a bonus-based strategy for spurring dynamic per-device server selection on the devices’ side. Our extensive evaluation shows that FedSwarm significantly outperforms other studies with better scalability, lower energy consumption, and higher model accuracy. Haizhou Du, Chengdong Ni, Chaoqian Cheng, Qiao Xiang, Xi Chen 0009, Xue (Steve) Liu |
IEEE Internet Things J. | 1 |
| 2024 | An efficient federated learning framework for graph learning in hyperbolic space
Haizhou Du, Conghao Liu, Huan Huo |
Knowl. Based Syst. | 1 |
| 2023 | What Appears Suboptimal May Surprise You: A Fixed-Rate Scheduling Policy for Geo-Distributed CoFlowsabstractAll existing coflow scheduling algorithms compute dynamic-rate schedules that change the rates of flows during transmission. In this paper, we make a crucial finding: although dynamically adjusting the rates of flows could lead to a better coflow completion time (CCT) in theory, it would introduce additional pressures on the congestion control mechanism in the underlying network, which result in poor CCT in practice. This difference between theoretical CCT and practical CCT is further exacerbated in wide-area networks, where the topology does not provide any bisection guarantee as data center networks do. To this end, we designed a fixed-rate coflow scheduling policy called FSCO. Although in theory, the best fixed-rate schedule is usually suboptimal, it keeps the in-flight traffic relatively steady, reducing the risk of triggering congestion control. The core of FSCO is an efficient scheduling algorithm based on the classic network utilization maximization (NUM) framework. We implement a prototype of FSCO and evaluate its performance extensively using real-world topologies and coflow traces. Experimental results show that the total CCT reduces up to 30% compared to baselines while yielding up to 12× speedups compared to the solver. Feiyan Ding, Yao Wang 0022, Qiao Xiang, Jiwu Shu, Haizhou Du, Linghe Kong, Xue (Steve) Liu |
ICPADS | 6 |
| 2023 | Toward a Unified Framework for Verifying and Interpreting Learning-Based Networking SystemsabstractThere has been a growing interest in applying machine learning to real-world tasks. However, due to the blackbox nature of machine learning models, it is crucial to (1) verify important properties of a model and (2) understand the reasons behind a model's prediction before deploying them in a production environment. Existing approaches typically handle them as two separate and sometimes orthogonal topics. In this paper, we show that the verification and interpretability of machine learning models are tightly related and can be unified by satisfiability modulo theories (SMT). Our key insight is: not only a wide range of properties of machine learning models can be formulated as SMT problems and verified accordingly, but many commonly studied interpretability questions can also be answered by iteratively checking the satisfiability and related properties of multiple SMT problems. Leveraging this insight, we design UINT, a general verification and interpretability framework for learning-based networking systems. UINT (1) allows operators to specify verification and interpretability problems as SMT formulas, (2) encodes the target machine learning models into SMT constraints, and (3) automatically solves the corresponding verification and interpretability problems using commodity SMT solvers. We implement a prototype of UINT and evaluate it on real-world learning-based networking systems. Results demonstrate the efficiency and efficacy of UINT in verifying and interpreting key questions for these systems. Yangfan Huang, Yuling Lin, Haizhou Du, Yijian Chen, Linghe Kong, Qiao Xiang, Qiang Li 0045, Franck Le, Jiwu Shu |
IWQoS | 3 |
| 2023 | An efficient federated learning framework for multi-channeled mobile edge network with layered gradient compression
Haizhou Du, Yijian Chen, Xiaojie Feng, Qiao Xiang |
Comput. Networks | 1 |
| 2023 | An efficient joint framework for interacting knowledge graph and item recommendation
Haizhou Du, Zebang Cheng |
Knowl. Inf. Syst. | 1 |
| 2023 | Heterogeneous Reinforcement Learning Network for Aspect-Based Sentiment Classification With External KnowledgeabstractAspect-based sentiment classification aims to automatically predict the sentiment polarity of the specific aspect in a text. However, it is challenging to confirm the mapping between the aspect and the core context since a number of existing methods concentrate on building the global relations of the full context rather than the partial connections based on the aspects. Motivated by the fundamental insights of reinforcement learning, we propose a novelHeterogeneousReinforcementLearningNetwork for aspect-based sentiment analysis (HRLN) to alleviate these issues, which contains two primary components, a heterogeneous network module, and a knowledge graph-based reinforcement learning module consistent with common-sense knowledge and emotional knowledge. To evaluate the effectiveness of HRLN, we conduct extensive experiments on five benchmark datasets, which indicate that HRLN achieves competitive performance and yields state-of-the-art results on all datasets. Additionally, we present an intuitive comprehension of why our HRLN model is more robust for aspect-based sentiment classification via case studies. Yukun Cao, Yijia Tang, Haizhou Du, Ziyue Wei, ChengKun Jin |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Orchestra: adaptively accelerating distributed deep learning in heterogeneous environmentsabstractThe synchronized Local-SGD(Stochastic gradient descent) strategy becomes a more popular in distributed deep learning (DML) since it can effectively reduce the frequency of model communication and ensure global model convergence. However, it works not well and leads to excessive training time in heterogeneous environments due to the difference in workers' performance. Especially, in some data unbalanced scenarios, these differences between workers may aggravate low utilization of resources and eventually lead to stragglers, which seriously hurt the whole training procedure. Existing solutions either suffer from a heterogeneity of computing resources or do not fully address the environment dynamics. Haizhou Du, Qiao Xiang |
CF | 1 |
| 2022 | Sailfish: A Fast Bayesian Change Point Detection Framework with Gaussian Process for Time Series
Haizhou Du |
ICANN (3) | 1 |
| 2022 | Finder: A novel approach of change point detection for multivariate time series
Haizhou Du, Ziyi Duan |
Appl. Intell. | 1 |
| 2022 | Isomer: Transfer enhanced dual-channel heterogeneous dependency attention network for aspect-based sentiment classification
Yukun Cao, Yijia Tang, Haizhou Du, Ziyue Wei, ChengKun Jin |
Knowl. Based Syst. | 3 |
| 2022 | Nostradamus: A novel event propagation prediction approach with spatio-temporal characteristics in non-Euclidean space
Haizhou Du |
Neural Networks | 1 |
| 2021 | Trident: Change Point Detection for Multivariate Time Series via Dual-Level Attention Learning
Ziyi Duan, Haizhou Du |
ACIIDS | 2 |
| 2021 | Symbiosis: A Novel Framework for Integrating Hierarchies from Knowledge Graph into Recommendation System
Haizhou Du |
KSEM | 1 |
| 2021 | Astrologer: Exploiting graph neural Hawkes process for event propagation prediction with spatio-temporal characteristics
Haizhou Du, Yunpu Ma |
Knowl. Based Syst. | 1 |
| 2020 | Controllable Multi-Character Psychology-Oriented Story GenerationabstractStory generation, which aims to generate a long and coherent story automatically based on the title or an input sentence, is an important research area in the field of natural language generation. There is relatively little work on story generation with appointed emotions. Most existing works focus on using only one specific emotion to control the generation of a whole story and ignore the emotional changes in the characters in the course of the story. In our work, we aim to design an emotional line for each character that considers multiple emotions common in psychological theories, with the goal of generating stories with richer emotional changes in the characters. To the best of our knowledge, this work is first to focuses on characters' emotional lines in story generation. We present a novel model-based attention mechanism that we call SoCP (Storytelling of multi-Character Psychology). We show that the proposed model can generate stories considering the changes in the psychological state of different characters. To take into account the particularity of the model, in addition to commonly used evaluation indicators(BLEU, ROUGE, etc.), we introduce the accuracy rate of psychological state control as a novel evaluation metric. The new indicator reflects the effect of the model on the psychological state control of story characters. Experiments show that with SoCP, the generated stories follow the psychological state for each character according to both automatic and human evaluations. Xinpeng Wang 0001, Yunpu Ma, Volker Tresp, Yuyi Wang 0001, Shanlin Zhou, Haizhou Du |
CIKM | 7 |
| 2019 | OctopusKing: A TCT-Aware Task Scheduling on Spark PlatformabstractThe majority of large-scale data intensive applications executed by data centers are based on Map-Reduce or its open-source implementation, Spark. However, the native Spark's data locality has significant side effects. Spark prioritizes high-priority tasks. But tasks at the next level of data locality are never started because Spark's delay scheduling mechanism. Additionally, the current scheduling approach on Spark only considers allocating reasonably computing resources for each task, but ignores the balance between computing resources and network resources. Network-intensive jobs will generate a large number of cross-rack tasks on heterogeneous clusters. More extensively, geographically distributed big data analytics clusters will produce greater network latency due to network transfer, which will decrease the performance of the whole cluster. Therefore, we propose OctopusKing a scheduling optimization approach based on the task completion time awareness. OctopusKing employs a deep learning algorithm to predict task completion time(TCT), which comprehensively considers four influence factors including the data size of each task, data nodes heterogeneity, the computation complexity of each task and network transfers. Based on the prediction of task completion time, our approach balances non-local task scheduling and network transfer time, which decreases the overall computation time of a job and improves the utilization of cluster resources effectively. Finally, we have implemented OctopusKing on Spark cluster with a 10 nodes. We show the convincing evidence that OctopusKing reduces average job completion time by 21% in comparison to Spark's default scheduler with little overhead. Haizhou Du |
ICPADS | 1 |
| 2019 | Hephaistos: A fast and distributed outlier detection approach for big mixed attribute dataabstractThis paper tackles a new problem in outlier detection: how to promptly detect the local outlier of a large-scale mixed attribute data in the big data era. This poses significant challenges due to a lack of access to the entire mixed attribute dataset at any individual compute machine. Proposed appr oaches firstly form a mechanism that deletes the massive clear non-noise and extracts cluster-based pre-noise set. Furthermore, we analyze pre-noise set using multi-step distributed LOF computing method on the Spark platform. Finally, the ordered LOF list is the output result. Comprehensive experiments are implemented by large-scale Benchmark datasets and the Spark platform. Extensive results show that the performance of our approaches are superior to the previous ones (4X faster than baseline LOF/2X faster than DLOF) when compared to state-of-the-art techniques, and therefore is believed to be able to give better guidance to local outlier detection of mixed attribute data. Haizhou Du |
Intell. Data Anal. | 1 |
| 2016 | Robust K-means algorithm with automatically splitting and merging clusters and its applications for surveillance data
Jingsheng Lei, Teng Jiang, Kui Wu 0001, Haizhou Du, Guokang Zhu, Zhaoqing Wang |
Multim. Tools Appl. | 4 |