Miao Xie

dblp:00/7808 · DBLP profile ↗
← Back
28ranked-venue papers
15as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Security and privacy · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Structure Guided Retrieval-Augmented Generation for Factual Queries
abstract
Retrieval-Augmented Generation (RAG) has been proposed to mitigate hallucinations in large language models (LLMs), where generated outputs may be factually incorrect.However, existing RAG approaches predominantly rely on vector similarity for retrieval, which is prone to semantic noise and fails to ensure that generated responses fully satisfy the complex conditions specified by factual queries.To address this challenge, we introduce a novel research problem, named EXACT RETRIEVAL PROBLEM (ERP).To the best of our knowledge, this is the first problem formulation that explicitly incorporates structural information into RAG for factual questions to satisfy all query conditions.For this novel problem, we propose STRUCTURE GUIDED RETRIEVAL-AUGMENTED GENERATION (SG-RAG) 1 , which models the retrieval process as an embedding-based subgraph matching task, and uses the retrieved topological structures to guide the LLM to generate answers that meet all specified query conditions.To facilitate evaluation of ERP, we construct and publicly release EXACT RETRIEVAL QUESTION ANSWERING (ERQA) 2 , a large-scale dataset comprising 120,000 fact-oriented QA pairs, each involving complex conditions, spanning 20 diverse domains.The experimental results demonstrate that SG-RAG significantly outperforms strong baselines on ERQA, delivering absolute gains of 20.68-50.88percentage points, corresponding to 34%-450% relative improvements across metrics, while maintaining reasonable computational overhead.
Miao Xie, Chunli Lv
ACL (1)1
2026 Beyond Single Slot: Joint Optimization for Multi-Slot Guaranteed Display Advertising
abstract
Guaranteed display advertising is crucial for platform monetization, yet existing methods often operate under a single-slot assumption, limiting their ability to optimize allocation across multi-slot page views. In this paper, we propose a novel joint optimization framework for multi-slot GD allocation, addressing key challenges such as slot-level redundancy, contract imbalance, and exposure concentration. Our approach formulates the allocation as an offline bipartite matching problem with a contract roulette mechanism for slot exclusivity and Page View constraints for impression control, and incorporates a scalable allocation optimization algorithm for efficient large-scale deployment. Extensive online tests on the Meituan advertising platform demonstrate that our method significantly improves merchant ROI, platform revenue efficiency, and contract fulfillment robustness. Specifically, online A/B tests show a 28.99% increase in Average Revenue Per User under 70% traffic, and DID analysis further indicates improved contract stability, demonstrating the strong applicability and effectiveness of our framework in real-world advertising deployments.
Jiaming Deng, Miao Xie, Linyou Cai, Qianlong Xie, Siqiang Luo, Gao Cong
SIGIR3
2026 T2Net: Tongue Image-Based T2DM Detection via Simulated Clinical Diagnostic Reasoning
abstract
Clinical studies indicate that the progression of Type 2 Diabetes Mellitus (T2DM) is associated with characteristic alterations in tongue features, which may facilitate non-invasive early detection. However, current deep learning-based tongue imaging approaches for diabetes diagnosis remain constrained by limited datasets, subtle feature variations, dependence on clinical expertise, and the lack of quantitative evaluation. To address these issues, we developed an open-source dataset for T2DM tongue diagnosis (DMT) and benchmarked it using multiple baseline models. Building on DMT, we propose T2Net, a tongue image recognition model for T2DM that simulates the clinical diagnostic process. T2Net comprises four core components: local inspection, pathological clue integration, syndrome identification, and diagnostic confidence estimation. First, T2Net automatically extracts key ROIs by combining large-kernel decomposition with multi-scale learning. Then, a multi-order feature interaction module enables effective fusion of tongue image features across scales to capture pathological clues. Meanwhile, we design a context-aware dynamic aggregation convolution to model long-range dependencies, and propose a flexible focal loss to mimic the diagnostic reasoning process of clinicians, enabling brain-inspired inference. Finally, we propose a clustering-based confidence estimation approach to quantitatively evaluate the reliability of model predictions. Experimental results demonstrate that T2Net achieves highly competitive performance on the DMT dataset, outperforming the second-best baseline by 2.7% in accuracy and 2.0% in F1 score. Moreover, the quantitative evaluation scores are largely consistent with clinical assessments by physicians.
Yanyi Huang, Liyun Li, Xiaojie Feng, Miao Xie, Jiayu Ye, An Zeng, Jianlu Bi
IEEE J. Biomed. Health Informatics6
2026 UrbanMFM: Spatial Graph-Based Multiscale Foundation Models for Learning Generalized Urban Representation
abstract
As geospatial data from web platforms becomes increasingly accessible and regularly updated, urban representation learning has emerged as a critical research area for advancing urban planning. Recent studies have developed foundation model-based algorithms to leverage this data for various urban-related downstream tasks. However, current research has inadequately explored deep integration strategies for multiscale, multimodal urban data in the context of urban foundation models. This gap arises primarily because the relationships between micro-scale (e.g., individual points of interest and street view imagery) and macro-scale (e.g., region-wide satellite imagery) urban features are inherently implicit and highly complex, making traditional interaction modeling insufficient. This paper introduces a novel research problem – how to learn multiscale urban representations by integrating diverse geographic data modalities and modeling complex multimodal relationships across different spatial scales. To address this significant challenge, we propose UrbanMFM, a spatial graph-based multiscale foundation model framework explicitly designed to capture and leverage these intricate relationships. UrbanMFM utilizes a self-supervised learning paradigm that integrates diverse geographic data modalities, including POI data and urban imagery, through novel contrastive learning objectives and advanced sampling techniques. By explicitly modeling spatial graphs to represent complex multiscale urban relationships, UrbanMFM effectively facilitates deep interactions between multimodal data sources. Extensive experiments on datasets from Singapore, New York, and Beijing demonstrate that UrbanMFM outperforms the strongest baselines significantly in four representative downstream tasks. By effectively modelling spatial hierarchies with diverse data, UrbanMFM provides a more comprehensive and adaptable representation of urban environments.
Miao Xie, Pasquale Balsebre, Weiming Huang 0001, Siqiang Luo, Gao Cong
IEEE Trans. Knowl. Data Eng.2
2022 Interconnected Neural Linear Contextual Bandits with UCB Exploration
Yang Chen 0028, Miao Xie, Jiamou Liu, Kaiqi Zhao 0001
PAKDD (1)2
2021 AutoBandit: A Meta Bandit Online Learning System
abstract
Recently online multi-armed bandit (MAB) is growing rapidly, as novel problem settings and algorithms motivated by various practical applications are being studied, building on the top of the classic bandit problem. However, identifying the best bandit algorithm from lots of potential candidates for a given application is not only time-consuming but also relying on human expertise, which hinders the practicality of MAB. To alleviate this problem, this paper outlines an intelligent system called AutoBandit, equipped with many out-of-the-box MAB algorithms, for automatically and adaptively choosing the best with suitable hyper-parameters online. It is effective to help a growing application for continuously maximizing cumulative rewards of its whole life-cycle. With a flexible architecture and user-friendly web-based interfaces, it is very convenient for the user to integrate and monitor online bandits in a business system. At the time of publication, AutoBandit has been deployed for various industrial applications.
Miao Xie, Wotao Yin
IJCAI1
2021 Geographical Hidden Markov Tree
abstract
Given a spatial raster framework with explanatory feature layers, a spatial contextual layer (e.g., a potential field), as well as a set of training samples with class labels, the spatial prediction problem aims to learn a model that can predict a class layer. The problem is important in societal applications such as flood extent mapping for disaster response and national water forecasting, but is challenging due to the noise, obstacles, and heterogeneity in feature maps, implicit spatial dependency between locations based on the contextual layer (e.g., gradient directions on a potential field), and the large number of sample locations. Existing work often assumes undirected spatial dependency, or directed dependency with a total order, and thus cannot reflect complex directed dependency with a partial order. In contrast, we recently proposed geographical hidden Markov tree, a probabilistic graphical model that generalizes the common hidden Markov model from a one-dimensional sequence to a two-dimensional map. Partial order class dependency is incorporated in the hidden class layer with a reverse tree structure. We also investigated computational algorithms for reverse tree construction, model parameter learning and class inference. This paper extends our recent model with overlaying class nodes between observation nodes and underlying hidden class nodes. The additional overlaying class layer makes the model more robust to large scale feature obstacles. We also proposed corresponding learning and inference methods. Extensive evaluations on real world datasets show that our models outperform multiple baselines in flood mapping applications, our algorithms are scalable on large data sizes, and the proposed extension enhances classification performance.
Zhe Jiang 0001, Miao Xie, Arpan Man Sainju
IEEE Trans. Knowl. Data Eng.2
2021 Characterizing Crowds to Better Optimize Worker Recommendation in Crowdsourced Testing
abstract
Crowdsourced testing is an emerging trend, in which test tasks are entrusted to the online crowd workers. Typically, a crowdsourced test task aims to detect as many bugs as possible within a limited budget. However not all crowd workers are equally skilled at finding bugs; Inappropriate workers may miss bugs, or report duplicate bugs, while hiring them requires nontrivial budget. Therefore, it is of great value to recommend a set of appropriate crowd workers for a test task so that more software bugs can be detected with fewer workers. This paper first presents a new characterization of crowd workers and characterizes them with testing context, capability, and domain knowledge. Based on the characterization, we then propose Multi-Objective Crowd wOrker recoMmendation approach (MOCOM), which aims at recommending a minimum number of crowd workers who could detect the maximum number of bugs for a crowdsourced testing task. Specifically, MOCOM recommends crowd workers by maximizing the bug detection probability of workers, the relevance with the test task, the diversity of workers, and minimizing the test cost. We experimentally evaluate MOCOM on 532 test tasks, and results show that MOCOM significantly outperforms five commonly-used and state-of-the-art baselines. Furthermore, MOCOM can reduce duplicate reports and recommend workers with high relevance and larger bug detection probability; because of this it can find more bugs with fewer workers.
Junjie Wang 0001, Song Wang 0009, Tim Menzies, Qiang Cui 0001, Miao Xie, Qing Wang 0001
IEEE Trans. Software Eng.6
2020 Dual attention convolutional network for action recognition
abstract
Action recognition has been an active research area for many years. Extracting discriminative spatial and temporal features of different actions plays a key role in accomplishing this task. Current popular methods of action recognition are mainly based on two‐stream Convolutional Networks (ConvNets) or 3D ConvNets. However, the computational cost of two‐stream ConvNets is high for the requirement of optical flow while 3D ConvNets takes too much memory because they have a large amount of parameters. To alleviate such problems, the authors propose a Dual Attention ConvNet (DANet) based on dual attention mechanism which consists of spatial attention and temporal attention. The former concentrates on main motion objects in a video frame by using ConvNet structure and the latter captures related information of multiple video frames by adopting self‐attention. Their network is entirely based on 2D ConvNet and takes in only RGB frames. Experimental results on UCF‐101 and HMDB‐51 benchmarks demonstrate that DANet gets comparable results among leading methods, which proves the effectiveness of the dual attention mechanism.
Xiaoqiang Li 0002, Miao Xie, Guangtai Ding, Weiqin Tong
IET Image Process.2
2020 FERRARI: an efficient framework for visual exploratory subgraph search in graph databases
Chaohui Wang, Miao Xie, Sourav S. Bhowmick, Byron Choi, Xiaokui Xiao, Shuigeng Zhou
VLDB J.2
2019 An Indexing Framework for Efficient Visual Exploratory Subgraph Search in Graph Databases
abstract
Although exploratory search has received significant attention recently in the context of structured data, scant attention has been paid for graph-structured data. In this paper, we present two novel index structures called VACCINE and ADVISE to efficiently support exploratory subgraph search in a visual environment (VESS). VACCINE is an offline, feature-based index that stores rich information related to frequent and infrequent subgraphs in the underlying graph database and how they can be transformed from one subgraph to another. ADVISE, on the other hand, is an adaptive, compact, on-the-fly index instantiated during iterative visual formulation/reformulation of a subgraph query for exploratory search and records relevant information to efficiently support its repeated evaluation. These indexes engender more efficient and scalable visual exploratory subgraph search framework compared to a state-of-the-art technique.
Chaohui Wang, Miao Xie, Sourav S. Bhowmick, Byron Choi, Xiaokui Xiao, Shuigeng Zhou
ICDE2
2019 A Practical Semi-Parametric Contextual Bandit
abstract
Classic multi-armed bandit algorithms are inefficient for a large number of arms. On the other hand, contextual bandit algorithms are more efficient, but they suffer from a large regret due to the bias of reward estimation with finite dimensional features. Although recent studies proposed semi-parametric bandits to overcome these defects, they assume arms' features are constant over time. However, this assumption rarely holds in practice, since real-world problems often involve underlying processes that are dynamically evolving over time especially for the special promotions like Singles' Day sales. In this paper, we formulate a novel Semi-Parametric Contextual Bandit Problem to relax this assumption. For this problem, a novel Two-Steps Upper-Confidence Bound framework, called Semi-Parametric UCB (SPUCB), is presented. It can be flexibly applied to linear parametric function problem with a satisfied gap-free bound on the n-step regret. Moreover, to make our method more practical in online system, an optimization is proposed for dealing with high dimensional features of a linear function. Extensive experiments on synthetic data as well as a real dataset from one of the largest e-commercial platforms demonstrate the superior performance of our algorithm.
Miao Xie, Xuying Meng, Nan Li 0019, Rong Jin 0001
IJCAI2
2018 Geographical Hidden Markov Tree for Flood Extent Mapping
abstract
Flood extent mapping plays a crucial role in disaster management and national water forecasting. Unfortunately, traditional classification methods are often hampered by the existence of noise, obstacles and heterogeneity in spectral features as well as implicit anisotropic spatial dependency across class labels. In this paper, we propose geographical hidden Markov tree, a probabilistic graphical model that generalizes the common hidden Markov model from a one dimensional sequence to a two dimensional map. Partial order class dependency is incorporated in the hidden class layer with a reverse tree structure. We also investigate computational algorithms for reverse tree construction, model parameter learning and class inference. Extensive evaluations on both synthetic and real world datasets show that proposed model outperforms multiple baselines in flood mapping, and our algorithms are scalable on large data sizes.
Miao Xie, Zhe Jiang 0001, Arpan Man Sainju
KDD1
2018 PANDA: A System for Partial Topology-based Search on Large Networks
abstract
A large body of research on subgraph query processing on large networks assumes that a query is posed in the form of a connected graph. Unfortunately, end users in practice may not always have precise knowledge about the topological relationships between nodes in a query graph to formulate a connected query. In this demonstration, we present a novel graph querying paradigm called partial topology-based network search and a query processing system called panda to efficiently find top-k matches of a partial topology query ( ptq ) in a single machine. A ptq is a disconnected query graph containing multiple connected query components . ptq s allow an end user to formulate queries without demanding precise information about the complete topology of a query graph. We demonstrate various innovative features of panda and its promising performance.
Miao Xie, Sourav S. Bhowmick, Gao Cong, Wook-Shin Han
Proc. VLDB Endow.1
2017 Who Should Be Selected to Perform a Task in Crowdsourced Testing?
abstract
Crowdsourced testing is an emerging trend in software testing, which relies on crowd workers to accomplish test tasks. Due to the cost constraint, a test task usually involves a limited number of crowd workers. Furthermore, more workers does not necessarily result in detecting more bugs. Different workers, who may have different testing experience and expertise, may make much differences in the test outcomes. For example, some inappropriate workers may miss true bug, introduce false bugs or report duplicated bugs, which decreases the test quality. In current practice, a test task is usually dispatched in a random manner, and the quality of testing cannot be guaranteed. Therefore, it is important to select an appropriate subset of workers to perform a test task to ensure high bug detection rate. This paper introduces ExReDiv, a novel hybrid approach to select a set of workers for a test task. It consists of three key strategies: the experience strategy selects experienced workers, the relevance strategy selects workers with expertise relevant to the given test task, the diversity strategy selects diverse workers to avoid detecting duplicated bugs. We evaluate ExReDiv based on 42 test tasks from one of the largest crowdsourced testing platforms in China, and the experimental results show its effectiveness.
Qiang Cui 0001, Junjie Wang 0001, Guowei Yang 0001, Miao Xie, Qing Wang 0001, Mingshu Li 0001
COMPSAC (1)4
2017 COCOON: Crowdsourced Testing Quality Maximization Under Context Coverage Constraint
abstract
Mobile app testing is challenging since each test needs to be executed in a variety of operating contexts including heterogeneous devices, various wireless networks and different locations. Crowdsourcing enables a mobile app test to be distributed as a crowdsourced task to leverage crowd workers to accomplish the test. However, high test quality and expected test context coverage are difficult to achieve in crowdsourced testing. Upon distributing a test task, mobile app providers neither know who to participate nor predict whether all the expected test contexts can be covered in the task. To address this problem, we put forward a novel research problem called Crowdsourced Testing Quality Maximization Under Context Coverage Constraint (Cocoon). Given a mobile app test task, our objective is to recommend a set of workers, from available crowd workers, such that the expected test context coverage and a high test quality can be achieved. We prove that the Cocoon problem is NP-Complete and then introduce two greedy approaches. Based on a real dataset from the largest Chinese crowdsourced testing platform, our evaluation shows the effectiveness and efficiency of the two approaches, which can be potentially used as online services in practice.
Miao Xie, Qing Wang 0001, Guowei Yang 0001, Mingshu Li 0001
ISSRE1
2017 A multi-objective biclustering algorithm based on fuzzy mathematics
Xiaoshu Zhu, Miao Xie, Jianxin Wang 0001
Neurocomputing3
2017 Distributed Segment-Based Anomaly Detection With Kullback-Leibler Divergence in Wireless Sensor Networks
abstract
In this paper, we focus on detecting a special type of anomaly in wireless sensor network (WSN), which appears simultaneously in a collection of neighboring nodes and lasts for a significant period of time. Existing point-based techniques, in this context, are not very effective and efficient. With the proposed distributed segment-based recursive kernel density estimation, a global probability density function can be tracked and its difference between every two periods of time is continuously measured for decision making. Kullback-Leibler (KL) divergence is employed as the measure and, in order to implement distributed in-network estimation at a lower communication cost, several types of approximated KL divergence are proposed. In the meantime, an entropic graph-based algorithm that operates in the manner of centralized computing is realized, in comparison with the proposed KL divergence-based algorithms. Finally, the algorithms are evaluated using a real-world data set, which demonstrates that they are able to achieve a comparable performance at a much lower communication cost.
Miao Xie, Jiankun Hu, Song Guo 0001, Albert Y. Zomaya
IEEE Trans. Inf. Forensics Secur.1
2017 PANDA: toward partial topology-based search on large networks in a single machine
Miao Xie, Sourav S. Bhowmick, Gao Cong, Qing Wang 0001
VLDB J.1
2015 DynaDiffuse: A Dynamic Diffusion Model for Continuous Time Constrained Influence Maximization
abstract
Studying the spread of phenomena in social networks is critical but still not fully solved. Existing influence maximization models assume a static network, disregarding its evolution over time. We introduce the continuous time constrained influence maximization problem for dynamic diffusion networks, based on a novel diffusion model called DynaDiffuse. Although the problem is NP-hard, the influence spread functions are monotonic and submodular, enabling fast approximations on top of an innovative stochastic model checking approach. Experiments on real social network data show that our model finds higher quality solutions and our algorithm outperforms state-of-art alternatives.
Miao Xie, Qiusong Yang, Qing Wang 0001, Gao Cong, Gerard de Melo
AAAI1
2015 Segment-Based Anomaly Detection with Approximated Sample Covariance Matrix in Wireless Sensor Networks
abstract
In wireless sensor networks (WSNs), it has been observed that most abnormal events persist over a considerable period of time instead of being transient. As existing anomaly detection techniques usually operate in a point-based manner that handles each observation individually, they are unable to reliably and efficiently report such long-term anomalies appeared in an individual sensor node. Therefore, in this paper, we focus on a new technique for handling data in a segment-based manner. Considering a collection of neighbouring data segments as random variables, we determine those behaving abnormally by exploiting their spatial predictabilities and, motivated by spatial analysis, specifically investigate how to implement a prediction variance detector in a WSN. As the communication cost incurred in aggregating a covariance matrix is finally optimised using the Spearman’s rank correlation coefficient and differential compression, the proposed scheme is able to efficiently detect a wide range of long-term anomalies. In theory, comparing to the regular centralised approach, it can reduce the communication cost by approximately 80 percent. Moreover, its effectiveness is demonstrated by the numerical experiments, with a real world data set collected by the Intel Berkeley ResearchLab (IBRL).
Miao Xie, Jiankun Hu, Song Guo 0001
IEEE Trans. Parallel Distributed Syst.1
2014 Evaluating Host-Based Anomaly Detection Systems: Application of the Frequency-Based Algorithms to ADFA-LD
Miao Xie, Jiankun Hu, Xinghuo Yu 0001, Elizabeth Chang 0001
NSS1
2014 A vertex centric parallel algorithm for linear temporal logic model checking in Pregel
Miao Xie, Qiusong Yang, Jian Zhai, Qing Wang 0001
J. Parallel Distributed Comput.1
2013 Scalable Hypergrid k-NN-Based Online Anomaly Detection in Wireless Sensor Networks
abstract
Online anomaly detection (AD) is an important technique for monitoring wireless sensor networks (WSNs), which protects WSNs from cyberattacks and random faults. As a scalable and parameter-free unsupervised AD technique, $(k)$-nearest neighbor (kNN) algorithm has attracted a lot of attention for its applications in computer networks and WSNs. However, the nature of lazy-learning makes the kNN-based AD schemes difficult to be used in an online manner, especially when communication cost is constrained. In this paper, a new kNN-based AD scheme based on hypergrid intuition is proposed for WSN applications to overcome the lazy-learning problem. Through redefining anomaly from a hypersphere detection region (DR) to a hypercube DR, the computational complexity is reduced significantly. At the same time, an attached coefficient is used to convert a hypergrid structure into a positive coordinate space in order to retain the redundancy for online update and tailor for bit operation. In addition, distributed computing is taken into account, and position of the hypercube is encoded by a few bits only using the bit operation. As a result, the new scheme is able to work successfully in any environment without human interventions. Finally, the experiments with a real WSN data set demonstrate that the proposed scheme is effective and robust.
Miao Xie, Jiankun Hu, Song Han 0006, Hsiao-Hwa Chen
IEEE Trans. Parallel Distributed Syst.1
2012 Histogram-Based Online Anomaly Detection in Hierarchical Wireless Sensor Networks
abstract
Online anomaly detection is critical for protecting wireless sensor networks (WSNs) from cyber-attacks and random faults, which handles the streaming data in real-time. Comparing to other techniques, histogram-based anomaly detection is cheaper in computation, which should be suitable for WSNs. However, performing histogram-based anomaly detection with an online manner in WSNs is not a straightforward issue. Most of the existing histogram-based schemes have to depend on a verification procedure, which costs a great amount of computational overhead as well as communication overhead. Thus, it almost wipes out the advantage of low complexity of histogram-based anomaly detection. This paper introduces a simple estimating approach to detect anomalies with the histogram, which takes account into the distributed manner and online manner at the same time. It also proves the error caused by the new estimate is very small, through a theoretical analysis. Moreover, the optimal parameter will be suggested by minimizing the error. Finally, a set of experiments are implemented with a real WSN dataset, which prove the new scheme is effective and efficient.
Miao Xie, Jiankun Hu, Biming Tian
TrustCom1
2011 Generalized hash-binary-tree based self-healing key distribution with implicit authentication
abstract
We propose a self-healing key distribution scheme with implicit authentication following a hash-binary-tree based key distribution scheme. The scheme reduces storage overhead without increase of communication and computation overhead. Implication authentication is subtly introduced to detect tamper attack against broadcast messages during transmission. The scheme achieves the same security level with Dutta et al’s approach which is an improved version of one of our schemes. The security of the proposed scheme is analyzed under an security model.
Biming Tian, Song Han 0004, Sazia Parvin, Miao Xie
IWCMC4
2011 Highly Efficient Distance-Based Anomaly Detection through Univariate with PCA in Wireless Sensor Networks
abstract
Unsupervised anomaly detection (UAD) techniques have received increasing attention in wireless sensor networks (WSNs). However, the high dimensional training data often make sensor nodes unable to sustain in computation, and result in quite expensive communication overhead. The feature reduction techniques make great sense through the reduction of the dimensionality when the features are strongly interrelated. Among these UAD techniques, distance-based anomaly detection (DB-AD) is a special one that allows to be described by a probability model. Based on this observation, DB-AD is explored deeply with a feature reduction technique, principal component analysis (PCA). Through examining the proportion of the variance explained by the first principal component (PC), a new feature reduction approach is proposed for DB- AD in WSNs, which enables to reduce the dimensionality to one in any situation. Specifically, the first PC is alone used for representing the original data as long as it retains most of the variance; otherwise, the information loss is geometrically reverted to neutralize the error. By obtaining a tradeoff between the detection error and performance overload, this approach is significant for resource-constrained WSNs, as the computational complexity and communication overhead will be reduced to a fraction of the original magnitude. Finally, this approach is evaluated with a real WSN dataset.
Miao Xie, Song Han 0004, Biming Tian
TrustCom1
2011 Anomaly detection in wireless sensor networks: A survey
Miao Xie, Song Han 0004, Biming Tian, Sazia Parvin
J. Netw. Comput. Appl.1