EDBT 2026 Demo / reviewers in the wild / expert
Yijie Wang 0001
dblp:91/1726-1 · also Yi-Jie Wang 0001
· DBLP profile ↗
82ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0002-2913-4016ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 17 · 4 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Computer networks · 4Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modeling heterogeneous normality in time series anomaly detection
Yijie Wang 0001, Hongzuo Xu |
Inf. Process. Manag. | 2 |
| 2026 | RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D ReconstructionabstractThe introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it substantially improves the mapping completeness and memory efficiency. However, the lack of reconstruction details and the time-consuming learning of neural representations hinder the widespread application of neural-based methods to large-scale online reconstruction. We introduce RemixFusion, a novel residual-based mixed representation for scene reconstruction and camera pose estimation dedicated to high-quality and large-scale online RGB-D reconstruction. In particular, we propose a residual-based map representation comprised of an explicit coarse TSDF grid and an implicit neural module that produces residuals representing fine-grained details to be added to the coarse grid. Such mixed representation allows for detail-rich reconstruction with bounded time and memory budget, contrasting with the overly-smoothed results by the purely implicit representations, thus paving the way for high-quality camera tracking. Furthermore, we extend the residual-based representation to handle multi-frame joint pose optimization via bundle adjustment (BA). In contrast to the existing methods, which optimize poses directly, we opt to optimize pose changes. Combined with a novel technique for adaptive gradient amplification, our method attains better optimization convergence and global optimality. Furthermore, we adopt a local moving volume to factorize the whole mixed scene representation with a divide-and-conquer design to facilitate efficient online learning in our residual-based framework. Extensive experiments demonstrate that our method surpasses all state-of-the-art ones, including those based either on explicit or implicit representations, in terms of the accuracy of both mapping and tracking on large-scale scenes. Project page can be found at https://lanlan96.github.io/RemixFusion/ . Yuqing Lan, Chenyang Zhu 0002, Shuaifeng Zhi, Jiazhao Zhang, Zhoufeng Wang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004 |
ACM Trans. Graph. | 7 |
| 2025 | LLM-based Rumor Detection via Influence Guided Sample Selection and Game-based Perspective AnalysisabstractZhiliang Tian, Jingyuan Huang, Zejiang He, Zhen Huang, Menglong Lu, Linbo Qiao, Songzhu Mei, Yijie Wang, Dongsheng Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhiliang Tian, Zejiang He, Zhen Huang 0006, Menglong Lu, Linbo Qiao, Songzhu Mei, Yijie Wang 0001 |
ACL (1) | 8 |
| 2025 | Graph Structure Learning via Transfer Entropy for Multivariate Time Series Anomaly DetectionabstractMultivariate time series anomaly detection (MTAD) poses a challenge due to temporal and feature dependencies. The critical aspects of enhancing the detection performance lie in accurately capturing the dependencies between variables within the sliding window and effectively leveraging them. Existing studies rely on domain knowledge to pre-set the window size, and overlook the strength of dependencies while calculating direction based on variable similarity. This paper proposes GSLTE, a graph structure learning method for MTAD. GSLTE employs Fast Fourier Transform to conduct iterative segmentation of the whole series, selecting the dominant Fourier frequency as the window size for each subsequence within the minimum interval. GSLTE quantifies the direction and strength of the dependencies based on variable-lag transfer entropy which is achieved through Dynamic Time Warping method to learn asymmetric links between variables. Extensive experiments show that GNN-based MTAD methods applying GSLTE can further improve anomaly detection performance while outperforming state-of-the-art competitors. Yijie Wang 0001 |
ICASSP | 2 |
| 2025 | BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box FusionabstractAbstract Open‐vocabulary 3D object detection has gained significant interest due to its critical applications in autonomous driving and embodied AI. Existing detection methods, whether offline or online, typically rely on dense point cloud reconstruction, which imposes substantial computational overhead and memory constraints, hindering real‐time deployment in downstream tasks. To address this, we propose a novel reconstruction‐free online framework tailored for memory‐efficient and real‐time 3D detection. Specifically, given streaming posed RGB‐D video input, we leverage Cubify Anything as a pre‐trained visual foundation model (VFM) for single‐view 3D object detection, coupled with CLIP to capture open‐vocabulary semantics of detected objects. To fuse all detected bounding boxes across different views into a unified one, we employ an association module for correspondences of multi‐views and an optimization module to fuse the 3D bounding boxes of the same instance. The association module utilizes 3D Non‐Maximum Suppression (NMS) and a box correspondence matching module. The optimization module uses an IoU‐guided efficient random optimization technique based on particle filtering to enforce multi‐view consistency of the 3D bounding boxes while minimizing computational complexity. Extensive experiments on CA‐1M and ScanNetV2 datasets demonstrate that our method achieves state‐of‐the‐art performance among online methods. Benefiting from this novel reconstruction‐free paradigm for 3D object detection, our method exhibits great generalization abilities in various scenarios, enabling real‐time perception even in environments exceeding 1000 square meters. Yuqing Lan, Chenyang Zhu 0002, Zhirui Gao, Jiazhao Zhang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004 |
Comput. Graph. Forum | 7 |
| 2025 | Correction: Deep anomaly detection with partition contrastive learning for tabular data
Yijie Wang 0001, Hongzuo Xu, Bin Li 0030 |
Data Min. Knowl. Discov. | 2 |
| 2024 | Boundary-Driven Active Learning for Anomaly Detection in Time Series Data StreamsabstractThe key to anomaly detection in time series data streams (TSDS) lies in the ability to adapt to evolving data. Active learning for anomaly detection has shown such ability by leveraging expert feedback. However, many studies in this research line strive to optimize performance by exhausting the query budget, lacking consideration of query necessity, which means some unnecessary queries may wrongly lead the model to overfit the trivial information and incur additional consumption in both human labeling and model execution. This paper proposes Boundary-driven Active Learning for Anomaly Detection (BALAD). BALAD utilizes deep one-class classification to construct a hypersphere boundary to sense data abnormality and filters out unnecessary queries by dividing the boundary region. We further harness the hypersphere boundary to quantitatively measure data difficulty, and a focal loss is introduced to prioritize hard samples. The boundary is flexibly adapted during each feedback iteration to accommodate changes in TSDS. Extensive experiments on six datasets demonstrate that BALAD significantly outperforms the state-of-the-art anomaly detection methods. Yijie Wang 0001, Hongzuo Xu |
ICASSP | 2 |
| 2024 | Adaptive and augmented active anomaly detection on dynamic network traffic streamsabstractActive anomaly detection queries labels of sampled instances and uses them to incrementally update the detection model, and has been widely adopted in detecting network attacks. However, existing methods cannot achieve desirable performance on dynamic network traffic streams because (1) their query strategies cannot sample informative instances to make the detection model adapt to the evolving stream and (2) their model updating relies on limited query instances only and fails to leverage the enormous unlabeled instances on streams. To address these issues, we propose an active tree based model, adaptive and augmented active prior-knowledge forest (A 3 PF), for anomaly detection on network traffic streams. A prior-knowledge forest is constructed using prior knowledge of network attacks to find feature subspaces that better distinguish network anomalies from normal traffic. On one hand, to make the model adapt to the evolving stream, a novel adaptive query strategy is designed to sample informative instances from two aspects: the changes in dynamic data distribution and the uncertainty of anomalies. On the other hand, based on the similarity of instances in the neighborhood, we devise an augmented update method to generate pseudo labels for the unlabeled neighbors of query instances, which enables usage of the enormous unlabeled instances during model updating. Extensive experiments on two benchmarks, CIC-IDS2017 and UNSW-NB15, demonstrate that A 3 PF achieves significant improvements over previous active methods in terms of the area under the receiver operating characteristic curve (AUC-ROC) (20.9% and 21.5%) and the area under the precision-recall curve (AUC-PR) (44.6% and 64.1%). Bin Li 0030, Yijie Wang 0001, Li Cheng 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2024 | Calibrated One-Class Classification for Unsupervised Time Series Anomaly DetectionabstractTime series anomaly detection is instrumental in maintaining system availability in various domains. Current work in this research line mainly focuses on learning data normality deeply and comprehensively by devising advanced neural network structures and new reconstruction/prediction learning objectives. However, their one-class learning process can be misled by latent anomalies in training data (i.e., anomaly contamination) under the unsupervised paradigm. Their learning process also lacks knowledge about the anomalies. Consequently, they often learn a biased, inaccurate normality boundary. To tackle these problems, this paper proposes calibrated one-class classification for anomaly detection, realizing contamination-tolerant, anomaly-informed learning of data normality via uncertainty modeling-based calibration and native anomaly-based calibration. Specifically, our approach adaptively penalizes uncertain predictions to restrain irregular samples in anomaly contamination during optimization, while simultaneously encouraging confident predictions on regular samples to ensure effective normality learning. This largely alleviates the negative impact of anomaly contamination. Our approach also creates native anomaly examples via perturbation to simulate time series abnormal behaviors. Through discriminating these dummy anomalies, our one-class learning is further calibrated to form a more precise normality boundary. Extensive experiments on ten real-world datasets show that our model achieves substantial improvement over sixteen state-of-the-art contenders. Hongzuo Xu, Yijie Wang 0001, Songlei Jian, Qing Liao 0001, Guansong Pang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Smoothing Point Adjustment-Based Evaluation of Time Series Anomaly DetectionabstractAnomalies in time series appear consecutively, forming anomaly segments. Applying the classical point-based evaluation metrics to evaluate the detection performance of segments leads to considerable underestimation, so most related studies resort to point adjustment. This operation treats all points as true positives within a segment equally when only one individual point alarms, resulting in significant overestimation and creating an illusion of superior performance. This paper proposes smoothing point adjustment, a novel range-based evaluation protocol for time series anomaly detection. Our protocol reflects detection performance impartially by carefully considering the specific location and frequency of alarms in the raw results. It is achieved by smoothly determining the adjustment range and rewarding early detection via a ranging function and a rewarding function. Compared with other evaluation metrics, experiments on different datasets show that our protocol can yield a performance ranking of various methods more consistent with the desired situation. Yijie Wang 0001, Hongzuo Xu, Bin Li 0030 |
ICASSP | 2 |
| 2023 | Fascinating Supervisory Signals and Where to Find Them: Deep Anomaly Detection with Scale LearningabstractDue to the unsupervised nature of anomaly detection, the key to fueling deep models is finding supervisory signals. Different from current reconstruction-guided generative models and transformation-based contrastive models, we devise novel data-driven supervision for tabular data by introducing a characteristic -- scale -- as data labels. By representing varied sub-vectors of data instances, we define scale as the relationship between the dimensionality of original sub-vectors and that of representations. Scales serve as labels attached to transformed representations, thus offering ample labeled data for neural network training. This paper further proposes a scale learning-based anomaly detection method. Supervised by the learning objective of scale distribution alignment, our approach learns the ranking of representations converted from varied subspaces of each data instance. Through this proxy task, our approach models inherent regularities and patterns within data, which well describes data "normality". Abnormal degrees of testing instances are obtained by measuring whether they fit these learned patterns. Extensive experiments show that our approach leads to significant improvement over state-of-the-art generative/contrastive anomaly detection methods. Hongzuo Xu, Yijie Wang 0001, Juhui Wei, Songlei Jian, Ning Liu 0015 |
ICML | 2 |
| 2023 | Local-Adaptive Transformer for Multivariate Time Series Anomaly Detection and DiagnosisabstractTime series data are pervasive in varied real-world applications, and accurately identifying anomalies in time series is of great importance. Many current methods are insufficient to model long-term dependence, whereas some anomalies can be only identified through long temporal contextual information. This may finally lead to disastrous outcomes due to false negatives of these anomalies. Prior arts employ Transformers (i.e., a neural network architecture that has powerful capability in modeling long-term dependence and global association) to alleviate this problem; however, Transformers are insensitive in sensing local context, which may neglect subtle anomalies. Therefore, in this paper, we propose a local-adaptive Transformer based on cross-correlation for time series anomaly detection, which unifies both global and local information to capture comprehensive time series patterns. Specifically, we devise a cross-correlation mechanism by employing causal convolution to adaptively capture local pattern variation, offering diverse local information into the long-term temporal learning process. Furthermore, a novel optimization objective is utilized to jointly optimize reconstruction of the entire time series and matrix derived from cross-correlation mechanism, which prevents the cross-correlation from becoming trivial in the training phase. The generated cross-correlation matrix reveals underlying interactions between dimensions of multivariate time series, which provides valuable insights into anomaly diagnosis. Extensive experiments on six real-world datasets demonstrate that our model outperforms state-of-the-art competing methods and achieves 6.8%-27.5%$F_{1}$score improvement. Our method also has good anomaly interpretability and is effective for anomaly diagnosis. Yijie Wang 0001, Hongzuo Xu, Ruyi Zhang 0002 |
SMC | 2 |
| 2023 | RoSAS: Deep semi-supervised anomaly detection with contamination-resilient continuous supervision
Hongzuo Xu, Yijie Wang 0001, Guansong Pang, Songlei Jian, Ning Liu 0015 |
Inf. Process. Manag. | 2 |
| 2023 | Deep Isolation Forest for Anomaly DetectionabstractIsolation forest (iForest) has been emerging as arguably the most popular anomaly detector in recent years due to its general effectiveness across different benchmarks and strong scalability. Nevertheless, its linear axis-parallel isolation method often leads to (i) failure in detecting hard anomalies that are difficult to isolate in high-dimensional/non-linear-separable data space, and (ii) notorious algorithmic bias that assigns unexpectedly lower anomaly scores to artefact regions. These issues contribute to high false negative errors. Several iForest extensions are introduced, but they essentially still employ shallow, linear data partition, restricting their power in isolating true anomalies. Therefore, this paper proposes deep isolation forest. We introduce a new representation scheme that utilises casually initialised neural networks to map original data into random representation ensembles, where random axis-parallel cuts are subsequently applied to perform the data partition. This representation scheme facilitates high freedom of the partition in the original data space (equivalent to non-linear partition on subspaces of varying sizes), encouraging a unique synergy between random representations and random partition-based isolation. Extensive experiments show that our model achieves significant improvement over state-of-the-art isolation-based methods and deep detectors on tabular, graph and time series datasets; our model also inherits desired scalability from iForest. Hongzuo Xu, Guansong Pang, Yijie Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | DPSS: Dynamic Parameter Selection for Outlier Detection on Data StreamsabstractOutlier detection on data streams identifies unusual states to sense and alarm potential risks and faults of the target systems in both the cyber and physical world. As different parameter settings of machine learning algorithms can result in dramatically different performance, automatic parameter selection is also of great importance in deploying outlier detection algorithms in data streams. However, current canonical parameter selection methods suffer from two key challenges: (i) Data streams generally evolve over time, but these existing methods use a fixed training set, which fails to handle this evolving environment and often results in suboptimal parameter recommendations; (ii) The stream is infinite, and thus any parameter selection method taking the entire stream as input is infeasible. In light of these limitations, this paper introduces a Dynamic Parameter Selection method for outlier detection on data Streams (DPSS for short). DPSS uses Gaussian process regression to model the relationship between parameters and detecting performance and uses Bayesian optimization to explore the optimal parameter setting. For each new subsequence, DPSS updates the recommended parameter setting to suit the evolving characteristics. Besides, DPSS only uses historical calculations to guide the parameter setting sampling and adjust the Gaussian process regression results. DPSS can be employed as an auxiliary plug-in tool to improve the detection performance of outlier detection methods. Extensive experiments show that our method can significantly improve the F-score of outlier detectors in data streams compared to its counterparts and obtains more superior parameter selection performance than other state-of the-art parameter selection approaches. DPSS also achieves better time and memory efficiency compared to competitors. Ruyi Zhang 0002, Yijie Wang 0001, Haifang Zhou, Bin Li 0030, Hongzuo Xu |
ICPADS | 2 |
| 2022 | Factorization Machine-based Unsupervised Model Selection Method*abstractMachine learning is broadly used in many intelligent cybernetic systems. With the burgeoning of the communities of AI, the number of machine learning-based models is rapidly increasing, but picking a suitable and optimal (or relatively good) model from overwhelming options has become a conundrum when deploying a new system. Therefore, we are motivated by an intriguing question: Can we automatically select a proper model for new data? However, unsupervised model selection poses two main challenges: (i) Evaluation and comparison of candidate models on the new data are infeasible due to the lack of labels; and (ii) It is non-trivial to build relationships between model performance and data characteristics when the interaction between these characteristics should be considered. In light of these limitations, this paper proposes a factorization machine-based unsupervised model selection method. Following mainstream model selection protocols, we also leverage model performance on prior known datasets. Differently, we learn higher-order complex relationships between model performance and dataset characteristics. Specifically, our method transfers the historical performance into a second-order function of meta-features and embedding weights by harnessing the power of factorization machine. This function can be subsequently used to select a proper model when given a new dataset. Extensive experiments show that our method obtains more superior model selection performance than five state-of-the-art approaches, and our method executes faster than its competitors by approximate three magnitudes. Ruyi Zhang 0002, Yijie Wang 0001, Hongzuo Xu, Haifang Zhou |
SMC | 2 |
| 2022 | DFAID: Density-aware and feature-deviated active intrusion detection over network traffic streams
Bin Li 0030, Yijie Wang 0001, Kele Xu, Li Cheng 0001, Zhiquan Qin |
Comput. Secur. | 2 |
| 2022 | ESDU: An elastic stripe-based delta update method for erasure-coded cross-data center storage systems
Han Bao 0006, Yijie Wang 0001 |
J. Parallel Distributed Comput. | 2 |
| 2021 | Neighborhood Consensus Networks for Unsupervised Multi-view Outlier DetectionabstractMulti-view outlier detection recently attracted rapidly growing attention with the development of multi-view learning. Although promising performance demonstrated, we observe that identifying outliers in multi-view data is still a challenging task due to the complicated characteristics of multi-view data. Specifically, an effective multi-view outlier detection method should be able to handle (1) different types of outliers; (2) two or more views; (3) samples without clusters; (4) high dimensional data. Unfortunately, little is known about how these four issues can be handled simultaneously. In this paper, we propose an unsupervised multi-view outlier detection method to address these issues. Our method is based on the proposed novel neighborhood consensus networks termed NC-Nets, which automatically encodes intrinsic information into a comprehensive latent space for each view (for issue (4)) and uniforms the neighborhood structures among different views (for issue (2)). Accordingly, we propose an outlier score measurement which consists of two parts: the within-view reconstruction score and the cross-view neighborhood consensus score. The measurement is designed based on the characteristics of the different outlier types (for issue (1)) and no cluster assumption is needed (for issue (3)). Experimental results show that our method significantly outperforms state-of-the-art methods. On average, our method achieves 11.2% ~ 96.2% improvement in term of AUC and 33.5% ~ 352.7% improvement in term of F1-Score. Li Cheng 0001, Yijie Wang 0001 |
AAAI | 2 |
| 2021 | NH-CIL: A Nested Hierarchy Algorithm for Class Incremental LearningabstractClass incremental learning is widely applied in the classification scenarios as the number of classes is usually dynamically changing. However, the existing algorithms increase computational cost to implement class incremental learning in order to increase classification quality. In this paper, we propose a nested hierarchy algorithm based on OCSVM for class incremental learning, called NH-CIL. We reuse support vectors to eliminate redundant instances and catch the key ones to replace the whole model because of the generalization ability of OCSVM. When a new class arrives, NH-CIL adopts OCSVM on the new class and the old classes respectively to get the corresponding sketching supports vectors. Then NH-CIL reuses these two kinds of sketching support vectors to build a binary sub-classifier. These two steps are repeatedly nested to form a hierarchy classification model in a bottom-up manner while the number of classes increases. On the contrary, the testing phase is in a top-down manner. NH-CIL can be used as a flexible approach in the classification scenarios during the collaborative information processing. We conduct the experiments on 8 real-world benchmark datasets to compare NH-CIL with some other class incremental learning algorithms, e.g. SD-CIL, HS-CIL and OP-CIL. The experiment results show that NH-CIL averagely achieves more than 5.1%, 8.6% and 11.6% accuracy improvement and 39.8%, 24.7% and 12.6% efficiency improvement over SD-CIL, HS-CIL and OP-CIL, respectively. Yijie Wang 0001 |
CSCWD | 2 |
| 2021 | Fden: Mining Effective Information of Features in Detecting Network AnomaliesabstractNetwork anomaly detection is important for detecting and reacting to the presence of network attacks. In this paper, we propose a novel method to effectively leverage the features in detecting network anomalies, named FDEn, consisting of flow-based Feature Derivation (FD) and prior knowledge incorporated Ensemble models (Enpk). To mine the effective information in features, 149 features are derived to enrich the feature set of the original data with covering more characteristics of network traffic. To leverage these features effectively, an ensemble model Enpk, including CatBoost and XGBoost, based on the bagging strategy is proposed to first detect anomalies by combining numerical features and categorical features. And then, Enpkadjusts the predicted label of specific data by incorporating the prior knowledge of network security. We conduct empirically experiments on the data set provided by the Network Anomaly Detection Challenge (NADC), in which we obtain average improvement up to 61.6%, 31.7%, 50.2%, and 45.0%, in terms of the cost score, precision, recall and F1-score, respectively. Bin Li 0030, Yijie Wang 0001, Kele Xu, Li Cheng 0001 |
ICASSP | 2 |
| 2021 | OADA: An Online Data Augmentation Method for Raw Histopathology Images
Zhiyue Wu, Yijie Wang 0001, Haibo Mi, Hongzuo Xu, Lanlan Feng |
ICONIP (6) | 2 |
| 2021 | Effective Anomaly Detection Based on Reinforcement Learning in Network Traffic DataabstractMixed-type data with both categorical and numerical features are ubiquitous in network security, but the existing methods are minimal to deal with them. Existing methods usually process mixed-type data through feature conversion, whereas their performance is downgraded by information loss and noise caused by the transformation. Meanwhile, existing methods usually superimpose domain knowledge and machine learning in which fixed thresholds are used. It cannot dynamically adjust the anomaly threshold to the actual scenario, resulting in inaccurate anomalies obtained, which results in poor performance. To address these issues, this paper proposes a novel Anomaly Detection method based on Reinforcement Learning, termed ADRL, which uses reinforcement learning to dynamically search for thresholds and accurately obtain anomaly candidate sets, fusing domain knowledge and machine learning fully and promoting each other. Specifically, ADRL uses prior domain knowledge to label known anomalies and uses entropy and deep autoencoder in the categorical and numerical feature spaces, respectively, to obtain anomaly scores combining with known anomaly information, which are integrated to get the overall anomaly scores via a dynamic integration strategy. To obtain accurate anomaly candidate sets, ADRL uses reinforcement learning to search for the best threshold. Detailedly, it initializes the anomaly threshold to get the initial anomaly candidate set and carries on the frequent rule mining to the anomaly candidate set to form the new knowledge. Then, ADRL uses the obtained knowledge to adjust the anomaly score and get the score modification rate. According to the modification rate, different threshold modification strategies are executed, and the best threshold, that is, the threshold under the maximum modification rate, is finally obtained, and the modified anomaly scores are obtained. The scores are used to re-carry out machine learning to improve the algorithm's accuracy for anomalous data. Repeat the above process until the method is stable. We experiment on ten real network traffic datasets. Experiments show ADRL averagely improves ROC-AUC and PR-AUC than eight state-of-the-art competitors by 89.6% and 286.0%, respectively. Yijie Wang 0001, Hongzuo Xu |
ICPADS | 2 |
| 2021 | DRAM Failure Prediction in Large-Scale Data CentersabstractCloud computing is developing rapidly. Data centers are important infrastructures of cloud service and JointCloud structure. DRAM failure is one of the main causes which can lead to node outage in data centers. This paper proposes a decision-tree-based DRAM failure prediction method for large-scale data centers of cloud service. We utilize the first public-available DRAM failure prediction dataset released in PAKDD 2021 AIOps competition. We construct a suite of handcrafted features based on the system kernel log data and MCA log data. Feature engineering is detailedly introduced in this paper, which can inspire and foster future research in this field. Harnessing the power of a state-of-the-art classifier (i.e., XGBoost), our method can effectively and timely predict DRAM failures. Our solution has good performance on the PAKDD 2021 dataset, it can generally achieve more than 60% precision in the validation phase. Extensive experiments investigate the performance of variants of our method to validate the significance of different strategies in the proposed solution. Hongzuo Xu, Songlei Jian, Chenlin Huang, Yijie Wang 0001, Zhiyue Wu |
JCC | 5 |
| 2021 | Beyond Outlier Detection: Outlier Interpretation by Attention-Guided Triplet Deviation NetworkabstractOutlier detection is an important task in many domains and is intensively studied in the past decade. Further, how to explain outliers, i.e., outlier interpretation, is more significant, which can provide valuable insights for analysts to better understand, solve, and prevent these detected outliers. However, only limited studies consider this problem. Most of the existing methods are based on the score-and-search manner. They select a feature subspace as interpretation per queried outlier by estimating outlying scores of the outlier in searched subspaces. Due to the tremendous searching space, they have to utilize pruning strategies and set a maximum subspace length, often resulting in suboptimal interpretation results. Accordingly, this paper proposes a novel Attention-guided Triplet deviation network for Outlier interpretatioN (ATON). Instead of searching a subspace, ATON directly learns an embedding space and learns how to attach attention to each embedding dimension (i.e., capturing the contribution of each dimension to the outlierness of the queried outlier). Specifically, ATON consists of a feature embedding module and a customized self-attention learning module, which are optimized by a triplet deviation-based loss function. We obtain an optimal attention-guided embedding space with expanded high-level information and rich semantics, and thus outlying behaviors of the queried outlier can be better unfolded. ATON finally distills a subspace of original features from the embedding module and the attention coefficient. With the good generality, ATON can be employed as an additional step of any black-box outlier detector. A comprehensive suite of experiments is conducted to evaluate the effectiveness and efficiency of ATON. The proposed ATON significantly outperforms state-of-the-art competitors on 12 real-world datasets and obtains good scalability w.r.t. both data dimensionality and data size. Hongzuo Xu, Yijie Wang 0001, Songlei Jian, Ning Liu 0015, Fei Li 0040 |
WWW | 2 |
| 2020 | Outlier Detection Ensemble with Embedded Feature SelectionabstractFeature selection places an important role in improving the performance of outlier detection, especially for noisy data. Existing methods usually perform feature selection and outlier scoring separately, which would select feature subsets that may not optimally serve for outlier detection, leading to unsatisfying performance. In this paper, we propose an outlier detection ensemble framework with embedded feature selection (ODEFS), to address this issue. Specifically, for each random sub-sampling based learning component, ODEFS unifies feature selection and outlier detection into a pairwise ranking formulation to learn feature subsets that are tailored for the outlier detection method. Moreover, we adopt the thresholded self-paced learning to simultaneously optimize feature selection and example selection, which is helpful to improve the reliability of the training set. After that, we design an alternate algorithm with proved convergence to solve the resultant optimization problem. In addition, we analyze the generalization error bound of the proposed framework, which provides theoretical guarantee on the method and insightful practical guidance. Comprehensive experimental results on 12 real-world datasets from diverse domains validate the superiority of the proposed ODEFS. Li Cheng 0001, Yijie Wang 0001, Bin Li 0030 |
AAAI | 2 |
| 2020 | Feature Selection on Data Stream via Multi-Cluster Structure PreservationabstractThe modern data arrive continuously in a rapid and time-varying stream, which appears to generate unstable associations on the data structure. However, most of the existing methods focus on dealing with the static data, and they cannot fully take them into the structure construction. To address this issue, we propose an online unsupervised Feature Selection method via Multi-Cluster structure Preservation (FSMCP for short). FSMCP weighs all features by minimizing the differences between the Multi-Cluster structures in the original and the selected feature space. The structure integrates the three-level associations, i.e., the individual-level associations, the aggregation-level associations, and the streaming-level associations. To provide informative features in time, FSMCP check and update the associations as soon as new instances arrive. In comparison with the baseline methods, FSMCP holds better efficiency than offline methods, while still providing almost similar or even better quantitative feature subsets. It outperforms the existing online methods with average NMI improvement of 10.33%. Yijie Wang 0001, Li Cheng 0001 |
CIKM | 2 |
| 2020 | Reducing network cost of data repair in erasure-coded cross-datacenter storage
Han Bao 0006, Yijie Wang 0001, Fangliang Xu |
Future Gener. Comput. Syst. | 2 |
| 2020 | An end-to-end distance measuring for mixed data based on deep relevance learningabstractDistance Measuring between two mixed data objects is the basis of many learning algorithms. The complex relevance between heterogeneous – various types/scales – attributes has a significant influence on the measured results. In this paper, we propose an End-to-End Distance Measuring method for mixe d data based on deep relevance learning, called E2DM. Existing methods confuse the attributes space by mapping the discrete attribute values to new continuous values, or discretize continuous attributes values without considering the relevance. In contrast, E2DM directly manipulates on the original data with data conversion and relevance learning simultaneously to avoid information loss and attribute space confusion. E2DM firstly estimates internal relevance (i.e., relevance within the attribute) influenced distance by considering the categorical attribute value frequency and mapping numerical attribute values into multiple bins. Then it takes a wrapper approach to iteratively optimize relevance influenced distance and bin boundaries using a Frobenius-norm deviation as its objective function. Co-occurrence Mover’s Distance is proposed to explicitly explore relevance between attributes in each iteration. Finally, the distance for numerical attribute values is refined based on the original values and the fallen bin centers. Experimental results on a number of real-world datasets demonstrate that E2DM outperforms the state-of-the-art methods. Li Cheng 0001, Yijie Wang 0001, Xingkong Ma |
Intell. Data Anal. | 2 |
| 2020 | An Adaptive Erasure Code for JointCloud Storage of Internet of Things Big DataabstractJointCloud is a cross-cloud cooperation architecture for integrated Internet service customization. The customized cross-cloud storage service based on this architecture is called JointCloud storage. Storing the Internet of Things (IoT) big data in erasure-coded JointCloud storage systems ensures that data can be accessed when several cloud services interrupt. However, because existing erasure codes cannot adapt the generator matrix and data placement scheme to different network environments and encoding parameters, they usually incur a large network resource consumption (NRC) for repairing data in JointCloud storage systems. As a result, the availability of IoT applications running on JointCloud storage systems is impaired. In this article, to minimize the NRC of repairing data, we propose an adaptive erasure code for JointCloud storage of IoT big data called ACIoT. Specifically, we first propose the concept of average weighted locality (AWL) of a stripe of erasure-coded data, which is proportional to the average NRC of repairing this stripe in JointCloud storage systems. Then, we propose an active parallel trial-and-error algorithm to calculate the optimal generator matrix and data placement scheme to achieve the lowest AWL, under different network environments and encoding parameters. By encoding and placing each stripe of data with the optimal generator matrix and data placement scheme, ACIoT can achieve the minimum NRC. The experiments show that, compared with several state-of-the-art erasure codes, ACIoT reduces the NRC by 26.4%-44.7%. Han Bao 0006, Yijie Wang 0001, Fangliang Xu |
IEEE Internet Things J. | 2 |
| 2019 | Embedding-Based Complex Feature Value Coupling Learning for Detecting Outliers in Non-IID Categorical DataabstractNon-IID categorical data is ubiquitous and common in realworld applications. Learning various kinds of couplings has been proved to be a reliable measure when detecting outliers in such non-IID data. However, it is a critical yet challenging problem to model, represent, and utilise high-order complex value couplings. Existing outlier detection methods normally only focus on pairwise primary value couplings and fail to uncover real relations that hide in complex couplings, resulting in suboptimal and unstable performance. This paper introduces a novel unsupervised embedding-based complex value coupling learning framework EMAC and its instance SCAN to address these issues. SCAN first models primary value couplings. Then, coupling bias is defined to capture complex value couplings with different granularities and highlight the essence of outliers. An embedding method is performed on the value network constructed via biased value couplings, which further learns high-order complex value couplings and embeds these couplings into a value representation matrix. Bidirectional selective value coupling learning is proposed to show how to estimate value and object outlierness through value couplings. Substantial experiments show that SCAN (i) significantly outperforms five state-of-the-art outlier detection methods on thirteen real-world datasets; and (ii) has much better resilience to noise than its competitors. Hongzuo Xu, Zhiyue Wu, Yijie Wang 0001 |
AAAI | 4 |
| 2019 | Central-Diffused Instance Generation Method in Class Incremental Learning
Yijie Wang 0001 |
ICANN (2) | 2 |
| 2019 | Unsupervised Feature Selection via Local Total-Order Preservation
Yijie Wang 0001, Li Cheng 0001 |
ICANN (2) | 2 |
| 2019 | MIX: A Joint Learning Framework for Detecting Both Clustered and Scattered Outliers in Mixed-Type DataabstractMixed-type data are pervasive in real life, but very limited outlier detection methods are available for these data. Some existing methods handle mixed-type data by feature converting, whereas their performance is downgraded by information loss and noise caused by the transformation. Another kind of approaches separately evaluates outlierness in numerical and categorical features. However, they fail to adequately consider the behaviours of data objects in different feature spaces, often leading to suboptimal results. As for outlier form, both clustered outliers and scattered outliers are contained in many real-world data, but a number of outlier detectors are inherently restricted by their outlier definitions to simultaneously detect both of them. To address these issues, an unsupervised outlier detection method MIX is proposed. MIX constructs a joint learning framework to establish a cooperation mechanism to make separate outlier scoring constantly communicate and sufficiently grasp the behaviours of data objects in another feature space. Specifically, MIX iteratively performs outlier scoring in numerical and categorical space. Each outlier scoring phase can be iteratively and cooperatively enhanced by the prior knowledge given by another feature space. To target both clustered and scattered outliers, the outlier scoring phases capture the essential characteristic of outliers, i.e., evaluating outlierness via the deviation from the normal model. We show that MIX significantly outperforms eight state-of-the-art outlier detectors on twelve real-world datasets and obtains good scalability. Hongzuo Xu, Yijie Wang 0001, Zhiyue Wu |
ICDM | 2 |
| 2019 | Multi-Hierarchy Attribute Relationship Mining Based Outlier Detection for Categorical DataabstractOutlier detection for categorical data is very important in many practical scenarios, such as intrusion detection, fraud detection, early detection of diseases, etc. However, there is no inherent difference measure for categorical data. The differences are hidden in complex attribute value relationships. Existing methods do not properly handle the internal relationship and external relationship of attributes, resulting in low accuracy of outlier detection.This paper proposes a novel unsupervised outlier detection method for categorical data based on Multi-Hierarchy Attribute Relationship Mining (MHARM). It detects outliers by mining the hierarchical and complex relationships between attribute values. MHARM first calculates the internal relationship. It processes each attribute independently via an information-theoretic difference to get an internal distance matrix. Then it handles different subhierarchy of external relationship. It divides attributes into two clusters, using mutual information as the correlation measure. For the external relationship of intra-cluster attributes, it iteratively updates an external distance matrix by using an entropy weighted Earth Mover's Distance (EMD) and the internal distance until convergence; for the external relationship of inter-cluster attributes, the joint entropy weighted sum is obtained to be the whole difference between objects. Finally, MHARM uses the sum of whole difference between objects as the outlier score, sorting it for outlier detection. Experimental results show that MHARM has an average AUC value of 13.84% higher than the state-of-the-art methods and significantly reduced the detection volume multiples (dvM) for unearthing 90% outliers on the given nine data sets. Yijie Wang 0001, Li Cheng 0001 |
IJCNN | 2 |
| 2019 | LAR: Locality-Aware Reconstruction for erasure-coded distributed storage systemsabstractSummary Many modern distributed storage systems adopt erasure coding to protect data from frequent server failures for cost reason. Reconstructing data in failed servers efficiently is vital to these erasure‐coded storage systems. To this end, tree‐structured reconstruction mechanisms where blocks are transmitted and combined through a reconstruction tree have been proposed. However, existing tree‐structured reconstruction mechanisms build reconstruction trees from the perspective of available network bandwidths between servers, which are fluctuating and difficult to measure. Besides, these reconstruction mechanisms cannot reduce data transmission. In this study, we overcome these limitations by proposing LAR, a locality‐aware tree‐structured reconstruction mechanism. LAR builds reconstruction trees from the perspective of data locality, which is stable and easy to obtain. More importantly, by building reconstruction trees that combine blocks closer to each other first, LAR can reduce the data transmitted through the network core and hence speed up reconstruction. We prove that a minimum spanning tree is an optimal reconstruction tree that minimizes core bandwidth usage. We also design and implement a general reconstruction framework that supports all tree‐structured reconstruction mechanisms and nearly all erasure codes. Large‐scale simulations on commonly deployed network topologies show that LAR consumes 20%–61% less core bandwidth than previous reconstruction mechanisms. Thorough experiments on a testbed consisting of 40 physical servers show that LAR improves proactive recovery throughput by 23% at least and improves degraded read rate by up to 68%. Fangliang Xu, Yijie Wang 0001, Xiaoqiang Pei, Xingkong Ma |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Variational autoencoder-based outlier detection for high-dimensional dataabstractAnalysis of high-dimensional data often suffers from the curse of dimensionality and the complicated correlation among dimensions. Dimension reduction methods often are used to alleviate these problems. Existing outlier detection methods based on dimension reduction usually only rely on reconstruct ion error to detect outlier or apply conventional outlier detection methods to the reduced data, which could deteriorate the performance of outlier detection as only considering part of the information from data. Few studies have been done to combine these two strategies to do outlier detection. In this paper, we proposed an outlier detection method based on Variational Autoencoder (VAE), which combines low-dimensional representation and reconstruction error to detect outliers. Specifically, we first model the data use VAE, then extract four outlier scores from VAE model, finally propose an ensemble method to combine the four outlier scores. The experiments conducted on six real-world datasets show that the proposed method performs better than or at least comparable to state of the art methods. Yongmou Li, Yijie Wang 0001, Xingkong Ma |
Intell. Data Anal. | 2 |
| 2019 | A Neural Probabilistic outlier detection method for categorical data
Li Cheng 0001, Yijie Wang 0001, Xingkong Ma |
Neurocomputing | 2 |
| 2019 | FAAD: an unsupervised fast and accurate anomaly detection method for a multi-dimensional sequence over data streamabstractRecently, sequence anomaly detection has been widely used in many fields. Sequence data in these fields are usually multi-dimensional over the data stream. It is a challenge to design an anomaly detection method for a multi-dimensional sequence over the data stream to satisfy the requirements of accuracy and high speed. It is because: (1) Redundant dimensions in sequence data and large state space lead to a poor ability for sequence modeling; (2) Anomaly detection cannot adapt to the high-speed nature of the data stream, especially when concept drift occurs, and it will reduce the detection rate. On one hand, most existing methods of sequence anomaly detection focus on the single-dimension sequence. On the other hand, some studies concerning multi-dimensional sequence concentrate mainly on the static database rather than the data stream. To improve the performance of anomaly detection for a multi-dimensional sequence over the data stream, we propose a novel unsupervised fast and accurate anomaly detection (FAAD) method which includes three algorithms. First, a method called “information calculation and minimum spanning tree cluster” is adopted to reduce redundant dimensions. Second, to speed up model construction and ensure the detection rate for the sequence over the data stream, we propose a method called “random sampling and subsequence partitioning based on the index probabilistic suffix tree.” Last, the method called “anomaly buffer based on model dynamic adjustment” dramatically reduces the effects of concept drift in the data stream. FAAD is implemented on the streaming platform Storm to detect multi-dimensional log audit data. Compared with the existing anomaly detection methods, FAAD has a good performance in detection rate and speed without being affected by concept drift. Bin Li 0030, Yijie Wang 0001, Yongmou Li, Xingkong Ma |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2018 | Exploring a High-quality Outlying Feature Value Set for Noise-Resilient Outlier Detection in Categorical DataabstractUnavoidable noise in real-world categorical data presents significant challenges to existing outlier detection methods because they normally fail to separate noisy values from outlying values. Feature subspace-based methods inevitably mix noisy values when retaining an entire feature because a feature may contain both outlying values and noisy values. Pattern-based methods are normally based on frequency and are easily misled by noisy values, resulting in many faulty patterns. This paper introduces a novel unsupervised framework termed OUVAS, and its parameter-free instantiation RHAC to explore a high-quality outlying value set for detecting outliers in noisy categorical data. Based on the observation that the relations between values reflect their essence, OUVAS investigates value similarities to cluster values into different groups and combines cluster-level analysis and value-level refinement to identify an outlying value set. RHAC instantiates OUVAS by three successive modules (i.e., the combination of Ochiai coefficient and LOUVAIN algorithm to cluster values, hierarchical value coupling learning to perform cluster-level analysis, and a threshold to divide fake and real outlying values in value-level refinement). We show that (i) RHAC-based outlier detector significantly outperforms five state-of-the-art outlier detection methods; (ii) Extended RHAC-based feature selection method successfully improves the performance of existing outlier detectors and performs better than two latest outlying feature selection methods. Hongzuo Xu, Li Cheng 0001, Yijie Wang 0001, Xingkong Ma |
CIKM | 4 |
| 2018 | FROD: Fast and Robust Distance-Based Outlier Detection with Active-Inliers-Patterns in Data Streams
Zongren Li, Yijie Wang 0001, Guohong Zhao, Li Cheng 0001, Xingkong Ma |
ICANN (1) | 2 |
| 2018 | Incremental encoding for erasure-coded cross-datacenters cloud storage
Fangliang Xu, Yijie Wang 0001, Xingkong Ma |
Future Gener. Comput. Syst. | 2 |
| 2018 | Stochastic extra-gradient based alternating direction methods for graph-guided regularized minimizationabstractIn this study, we propose and compare stochastic variants of the extra-gradient alternating direction method, named the stochastic extra-gradient alternating direction method with Lagrangian function (SEGL) and the stochastic extra-gradient alternating direction method with augmented Lagrangian function (SEGAL), to minimize the graph-guided optimization problems, which are composited with two convex objective functions in large scale. A number of important applications in machine learning follow the graph-guided optimization formulation, such as linear regression, logistic regression, Lasso, structured extensions of Lasso, and structured regularized logistic regression. We conduct experiments on fused logistic regression and graph-guided regularized regression. Experimental results on several genres of datasets demonstrate that the proposed algorithm outperforms other competing algorithms, and SEGAL has better performance than SEGL in practical use. Qiang Lan, Linbo Qiao, Yijie Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2018 | TA-Update: An Adaptive Update Scheme with Tree-Structured Transmission in Erasure-Coded Storage SystemsabstractErasure coding has received considerable attentions due to the better tradeoff between the space efficiency and reliability. The frequent update of the stored data in the distributed storage systems has posed a new challenge for erasure codes: how to update the erasure-coded data in a general, efficient and adaptive way. However, existing update schemes of erasure codes are inadequate to meet these requirements, since their code-related update manners lead to a low generality, their star-structured data transmission manners lead to a low update efficiency, and their redo manners when encountering the node failure lead to a low adaptivity. In this paper, we propose an adaptive update scheme with the tree-structured transmission, called TA-Update, which consists of a code-independent update framework and three algorithms: the rack-aware tree construction algorithm, the top-down data processing algorithm and the rollback-based failure processing algorithm. For generality, we propose a code-independent update framework with the tree structure to support the MDS code with any coding parameter. For efficiency, a rack-aware tree construction algorithm is proposed to achieve the high available bandwidth, which organizes the data node and parity nodes as an update tree. Moreover, a top-down data processing algorithm is proposed to achieve the high transmission and computation efficiency, which pipelines the data transmission along the update tree and distributes the encoding computations among all the participating nodes. For adaptivity, we propose a rollback-based failure processing algorithm to achieve high adaptivity, which handles the node failure during update with the existing update tree in a rollback manner. To evaluate the performance of TA-Update, we conduct experiments on HDFS-RAID under various parameter settings on both 30 physical and 200 virtual machines. Extensive experiments confirm that TA-Update could support the various erasure codes with any parameter, improve the update efficiency by 30 percent and the adaptivity by 47 percent on average compared with the state-of-the-art approaches under various parameter settings. Yijie Wang 0001, Xiaoqiang Pei, Xingkong Ma, Fangliang Xu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | TMRCP: A Trend-Matching Resources Coupled Prediction Method over Data Stream
Runfan Wu, Yijie Wang 0001, Xingkong Ma, Li Cheng 0001 |
ICONIP (5) | 2 |
| 2017 | A cloud-assisted publish/subscribe service for time-critical dissemination of bulk contentabstractSummary Characterized by the increasing arrival rate of live content, emergency applications pose a great challenge: how to disseminate data with diverse sizes to interested users in a real‐time manner. Most file sharing applications focus on the dissemination of bulk content with less consideration of users' interests. On the other hand, existing publish/subscribes are designed for notifying interested users with small‐sized content. To bridge this gap, we propose CAPS, a cloud‐assisted publish/subscribe service for time‐critical bulk content dissemination. In CAPS, a hybrid 2‐layer architecture is proposed to knit servers in the cloud and clients in the internet. Through dividing each event into attribute‐value pairs and the data content, CAPS provides both event matching service and data distribution in a parallel manner. To improve the upload bandwidth of data distribution, we propose a helper‐based content distribution protocol, where the servers not only guide the clients with similar interests to exchange their received data blocks but also contribute their own upload capacities to clients. Moreover, a volume‐aware helper renting scheme is proposed to adaptively adjust the scale of servers according to the churn of data volume, leading to a high‐performance price ratio. So as to evaluate the performance of CAPS, about 1000 virtual machines are deployed in our Cloud‐Stack testbed. Extensive experiments confirm that CAPS can linearly reduce the download completion time with the growing number of servers, adaptively adjust the upload capacity in tens of seconds according to the change of the workloads, and ensure reliable data dissemination even if a large number of nodes frequently churn or instantaneously fail. Compared with the state‐of‐the‐art approaches, CAPS demonstrates better performance under various parameter settings. Xingkong Ma, Yijie Wang 0001, Xiaoqiang Pei, Fangliang Xu |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | A decentralized redundancy generation scheme for codes with locality in distributed storage systemsabstractSummary The increasing data volume in a large number of applications presents a dire need for supporting the reliable data management in distributed storage systems. Existing classical erasure codes, such as the Reed‐Solomon codes and locally reconstruction codes, are widely adopted by many distributed storage systems. However, existing researches mainly focus on proposing new optimized codes, ignoring the optimization of the encoding process with the classical codes, where inefficient encoding process greatly degrades the encoding performance of the distributed storage systems. Thus, how to complete the encoding process in an efficient way has become the challenge for adopting the classical codes. In this paper, we propose a decentralized redundancy generation scheme on the basis of the codes with locality, called D2CP, where a 2‐step framework is proposed to support both the data patterns (replication to encodinganddirect encoding) and codes with locality with any parameter set. For improving the insertion throughput, D2CP adopts a data placement technique with consistent hashing to guide the selection of nodes. For reducing the network traffic cost, D2CP adopts a data sending scheduling technique to schedule the transmission of the source nodes and a cooperative parity generation technique to generate the parity data cooperatively. To evaluate the performance of D2CP, we conduct experiments on our RAID distributed storage system under various parameter settings with both 30 physical and 200 virtual servers. Extensive experiments confirm that D2CP can improve the encoding throughput by 20% and 32% and reduce the network traffic cost by 16% and 33% compared with the typical approaches on average for the 2 data patterns respectively. Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Efficient in-place update with grouped and pipelined data transmission in erasure-coded storage systems
Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu |
Future Gener. Comput. Syst. | 2 |
| 2016 | A C-SVM Based Anomaly Detection Method for Multi-Dimensional Sequence over Data StreamabstractAnomaly detection over multi-dimensional data stream has attracted considerable attention recently in various fields, such as network, finance and aerospace. In many cases, anomalies are composed of a sequence of multi-dimensional data, and it's necessary to detect this type of anomalies accurately and efficiently over data stream. Existing online methods of anomaly detection merely focus on the single-dimensional sequence. What's more, current studies about multi-dimensional sequence are mainly concentrated on static database. However, the anomaly detection for multi-dimensional sequence over data stream is much more difficult, due to the complexity of multidimensional sequence processing, the dynamic nature of data stream and the unbalance between normal and abnormal data. Facing these challenges, we propose an anomaly detection method for multi-dimensional sequence over data stream based on cost sensitive support vector machine (C-SVM) called ADMS. First, to improve the accuracy and efficiency, the ADMS transforms multi-dimensional sequences into feature vectors in a lossless way and prunes worthless features of these vectors. And then, the ADMS can detect abnormal sequences over dynamically imbalanced data stream by lively testing these vectors based on C-SVM. Experiments show that the false negative rate (FNR) of the ADMS is lower than 5%, the false positive rate (FPR) is lower than 7%, and the throughput is improved 42% by pruning worthless features. In addition, the AMDS performs well when there are concept drifts over the data stream. Han Bao 0006, Yijie Wang 0001 |
ICPADS | 2 |
| 2016 | GDSW: A General Framework for Distributed Sliding Window over Data StreamsabstractThe big data era is characterized by the emergence of live data with high volume and fast arrival rate, it poses a new challenge to stream processing applications: how to process the unbounded live data in real time with high throughput. The sliding window technique is widely used to handle the unbounded live data by storing the most recent history of streams. However, existing centralized solutions cannot satisfy the requirements for high processing capacity and low latency due to the single-node bottleneck. Moreover, existing studies on distributed windows primarily focus on specific operators, while a general framework for processing various window-based operators is wanted. In this paper, we firstly classify the window-based operators to two categories: data-independent operators and data-dependent operators. Then, we propose GDSW, a general framework for distributed count-based sliding window, which can handle both of data-independent and data-dependent operators. Besides, in order to balance system load, we further propose a dynamic load balance algorithm called DAD based on buffer usage. Our framework is implemented on Apache Storm 0.10.0. Extensive evaluation shows that GDSW can achieve sub-second latency, and 10X improvement in throughput compared with centralized processing, when processing rapid data rate or big size window. Yijie Wang 0001, Xingkong Ma |
ICPADS | 2 |
| 2016 | T-Update: A tree-structured update scheme with top-down transmission in erasure-coded systemsabstractErasure coding has received considerable attention due to the better tradeoff between the space efficiency and reliability. However, it consumes large network traffic and long time to complete the update, involving updates of both data nodes and parity nodes. Existing solutions to this problem mainly focus on proposing new class of codes with lower update complexity to reduce the network traffic, ignoring the optimization of data transmission structure. In fact, the data transmission structure has great impact on the update. In this paper, we propose T-Update, a tree-structured update scheme with top-down transmission that minimizes the update time for erasure-coded data with no additional network traffic. Specially, we propose a rack-aware tree construction technique to construct an update tree to organize the data connections, with the data node as the root and the parity nodes as the children. To maximize the update efficiency, we propose a top-down data transmission technique to guide the data transmission and distribute the data computation for updating the parity nodes. To evaluate the performance of T-Update, we conduct experiments on HDFS-RAID under various parameter settings on both 30 physical and 200 virtual servers. Extensive experiments confirm that T-Update reduces the update time by 27% and 32% on average compared with two typical update schemes respectively. Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu |
INFOCOM | 2 |
| 2016 | A Variable Markovian Based Outlier Detection Method for Multi-Dimensional Sequence over Data StreamabstractNowadays sequence data tends to be multi-dimensional sequence over data stream, it has a large state space and arrives at unprecedented speed. It is a big challenge to design a multi-dimensional sequence outlier detection method to meet the accurate and high speed requirements. The traditional methods can't handle multi-dimensional sequence effectively as they have poor abilities for multi-dimensional sequence modeling, and can't detect outlier timely as they have high computational complexity. In this paper we propose a variable Markovian based outlier detection method for multi-dimensional sequence over data stream, VMOD, which consists of two algorithms: mutual information based feature selection algorithm (MIFS), variable Markovian based sequential analysis algorithm (VMSA). It uses MIFS algorithm to reduce the state space and redundant features, and uses VMSA algorithm to accelerate the outlier detection. Through VMOD method, we can improve the detection rate and detection speed. The MIFS algorithm uses mutual information as similarity measures and adopt clustering based strategy to select features, it can improve the abilities for sequence modeling through reducing the state space and redundant features, consequently, to improve the detection rate. The VMSA algorithm use random sample and index structure to accelerate the variable Markovian model construction and reduce the model complexity, consequently, to quicken the outlier detection. The experiments show that VMOD can detect outlier effectively, and reduce the detection time by at least 50% compared with the traditional methods. Yijie Wang 0001, Yongmou Li, Xingkong Ma |
PDCAT | 2 |
| 2016 | A User Behavior Anomaly Detection Approach Based on Sequence Mining over Data StreamsabstractHow to design a low-latency and accurate approach for user behavior anomaly detection over data streams has become a great challenge. However, existing studies cannot meet low-latency and accurate requirements, due to a large number of subsequences and sequential relationship in behaviors. This paper presents BADSM, a user behavior anomaly detection approach based on sequence mining over data streams that seeks to address such challenge. BADSM uses self-adaptive behavior pruning algorithm to adaptively divide data stream into behaviors and decrease the number of subsequences to improve the efficiency of sequence mining. Meanwhile, the top-k abnormal scoring algorithm is used to reduce the complexity of traversal and obtain quantitative detection result to improve accuracy. We design and implement a streaming anomaly detection system based on BADSM to perform online detection. Extensive experiments confirm that BADSM significantly reduces processing delay by at least 36.8% and false positive rate by 6.4% compared with the classic sequence mining approach PrefixSpan. Yijie Wang 0001, Xingkong Ma |
PDCAT | 2 |
| 2016 | Repairing multiple failures adaptively with erasure codes in distributed storage systemsabstractSummary Repairs of multiple failures in distributed storage systems have posed the challenges for erasure coding: how to minimize the repair time with the least extra repair network traffic cost. However, existing repair schemes designed for single failure suffer from the high network traffic cost due to the serial repairs for multiple failures. Repair schemes designed for multiple failures suffer from long repair time due to the centralized repair structure. In this paper, we propose a decentralized adaptive repair scheme, called DARS, to minimize the repair time with the least extra network traffic cost. Specially, we propose a three‐layer repair model to support the repairs for both the single and multiple failures. For low repair time, a bandwidth‐aware node selection technique is proposed to guide the selection of nodes, and a line‐structured data transmission technique is proposed to organize the data transmission between the providers and the newcomer. For the least extra network traffic cost, a core‐based data distribution technique is proposed to organize the data transmission between the coordinator and other newcomers, and an intersection provider adjustment technique is proposed to adaptively adjust the number of intersection providers. Moreover, we adopt the ‘lazy repair’ within a stripe to further reduce the repair network traffic cost. We implement and evaluate DARS on our raid distributed storage system under various parameter settings with 30 physical machines and 200 virtual machines. Experimental results confirm that DARS reduces the repair time by 29% and 55% on average compared with tree‐structured repair and CORE, respectively. Copyright © 2015 John Wiley & Sons, Ltd. Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Towards Latency-Optimal Distributed Relay SelectionabstractLatency-sensitive multiparty applications involve intensive communication between multiple participating nodes. Relays are usually adopted for matchmaking end hosts, filtering unwanted traffics, bypassing routing outages and so on. Speeding up the relay-communication becomes increasingly important to improve the QoE of clients. Currently, no rigorous guarantees have been made for the latency-optimal relay communication. We propose a novel framework to truthfully represent the relay communication in the latency space. Real-world data sets show that nearly 90% of node triples obey the average triangle inequality, while our new model allows for the asymmetry and triangle inequality violations to occur. We propose the general triangle to rigorously locate a candidate relay closer to multiple nodes, with which we systematically analyze the feasibility of finding an optimal relay node for arbitrarily sized groups. Our results show that distributed greedy methods are able to locate optimal relays with modest communication overhead and small search hops. Yongquan Fu, Yijie Wang 0001, Xiaoqiang Pei |
CCGRID | 2 |
| 2015 | Scalable and elastic total order in content-based publish/subscribe systems
Xingkong Ma, Yijie Wang 0001, Xiaoqiang Pei, Fangliang Xu |
Comput. Networks | 2 |
| 2015 | BLOR: An efficient bandwidth and latency sensitive overlay routing approach for flash data disseminationabstractSummary The problem of flash data dissemination refers to transmitting time‐critical data to a large group of distributed receivers in a timely manner, which widely exists in many mission‐critical applications and Web services. However, existing approaches for flash data dissemination fail to ensure the timely and efficient transmission, because of the unpredictability of the dissemination process. Overlay routing has been widely used as an efficient routing primitive for providing better end‐to‐end routing quality by detouring inefficient routing paths in the real networks. To improve the predictability of the flash data dissemination process, we propose a bandwidth and latency sensitive overlay routing approach named BLOR, by optimizing the overlay routing and avoiding inefficient paths in flash data dissemination. BLOR tries to select optimal routing paths in terms of network latency, bandwidth capacity, and available bandwidth in nature, which has never been studied before. Additionally, a location‐aware unstructured overlay topology construction algorithm, an unbiased top‐kdominance model, and an efficient semi‐distributed information management strategy are proposed to assist the routing optimization of BLOR. Extensive experiments have been conducted to verify the effectiveness and efficiency of the proposals with real‐world data sets. Copyright © 2014 John Wiley & Sons, Ltd. Xiaoyong Li 0002, Yijie Wang 0001, Yongquan Fu, Xiaoling Li 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | A general scalable and elastic matching service for content-based publish/subscribe systemsabstractSUMMARY Characterized by the emergence of a large number of live content, the emergency applications have received increasing attention in recent years. Providing a general and scalable event, matching service can precisely notify users latest information that they are interested in. However, because the live content arrival rate may churn significantly in a short time and subscriptions with various patterns tend to be skewed, it is challenging to increase the generality, scalability, and elasticity of the matching process. We propose a novel parallel event matching service based on the cloud computing environment, called GSEM, to satisfy these requirements. GSEM first presents a two‐hop framework and a general subscription pattern to handle various patterns of subscriptions. To provide scalable matching service, ahybrid content space partitionscheme is proposed to divide large skewed subscriptions into multiple small clusters managed by a group of parallel servers. To adapt to the sudden change of event arrival rate, GSEM elastically adjusts the scale of servers and rebalances their workloads through aperformance‐aware detectiontechnique. A prototype deployment on the OpenStack platform shows that GSEM achieves scalable matching throughput with the growth of servers, elastic service capacity with the change of event arrival rate, and significantly outperforms the existing cloud based systems in various workloads. Copyright © 2014 John Wiley & Sons, Ltd. Xingkong Ma, Yijie Wang 0001, Xiaoqiang Pei, Xiaoyong Li 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | A Scalable and Reliable Matching Service for Content-Based Publish/Subscribe SystemsabstractCharacterized by the increasing arrival rate of live content, the emergency applications pose a great challenge: how to disseminate large-scale live content to interested users in a scalable and reliable manner. The publish/subscribe (pub/sub) model is widely used for data dissemination because of its capacity of seamlessly expanding the system to massive size. However, most event matching services of existing pub/sub systems either lead to low matching throughput when matching a large number of skewed subscriptions, or interrupt dissemination when a large number of servers fail. The cloud computing provides great opportunities for the requirements of complex computing and reliable communication. In this paper, we propose SREM, a scalable and reliable event matching service for content-based pub/sub systems in cloud computing environment. To achieve low routing latency and reliable links among servers, we propose a distributed overlay SkipCloud to organize servers of SREM. Through a hybrid space partitioning technique HPartition, large-scale skewed subscriptions are mapped into multiple subspaces, which ensures high matching throughput and provides multiple candidate servers for each event. Moreover, a series of dynamics maintenance mechanisms are extensively studied. To evaluate the performance of SREM, 64 servers are deployed and millions of live content items are tested in a CloudStack testbed. Under various parameter settings, the experimental results demonstrate that the traffic overhead of routing events in SkipCloud is at least 60 percent smaller than in Chord overlay, the matching rate in SREM is at least 3.7 times and at most 40.4 times larger than the single-dimensional partitioning technique of BlueDove. Besides, SREM enables the event loss rate to drop back to 0 in tens of seconds even if a large number of servers fail simultaneously. Xingkong Ma, Yijie Wang 0001, Xiaoqiang Pei |
IEEE Trans. Cloud Comput. | 2 |
| 2015 | A General Scalable and Elastic Content-Based Publish/Subscribe ServiceabstractThe big data era is characterized by the emergence of live content with increasing complexities of data dimensionality and data sizes, which poses a new challenge to emergency applications: how to timely disseminate large-scale live content to users who are interested in. The publish/subscribe (pub/sub) model is widely used to disseminate data because of its possibility of expanding the system to Internet-scale size. However, existing pub/sub systems are inadequate to meet the requirement of disseminating live content in the big data era, since their multi-hop routing techniques and coarse-grained partitioning techniques lead to a low matching throughput, and their upload capacities do not scale well. In this paper, we propose a general scalable and elastic pub/sub service based on the cloud computing environment, called GSEC. For generality, we propose a two-layer pub/sub framework to support the dissemination with diverse data sizes and data dimensionality. For scalability, a hybrid space partitioningtechnique is proposed to achieve high matching throughput, which divides subscriptions into multiple clusters in a hierarchical manner. Moreover, a helper-based content distribution technique is proposed to achieve high upload bandwidth, where servers act as both providers and coordinators to fully explore the upload capacity of the system. For elasticity, we propose a performance-aware provisioningtechnique to adjust the scale of servers to adapt to the churn workloads. To evaluate the performance of GSEC, about 1,000 servers are deployed and hundreds of thousands of live content items are tested in our CloudStack-based testbed. Extensive experiments confirm that GSEC can linearly increase the capacities of event matching and content distribution with the growth of servers, adaptively adjust these capacities in tens of seconds according to the churn workloads, and significantly outperforms the state-of-the-art approaches under various parameter settings. Yijie Wang 0001, Xingkong Ma |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | MCRTREE: A Mutually Cooperative Recovery Scheme for Multiple Losses in Distributed Storage Systems Based on Tree StructureabstractTo guarantee the reliability of distributed storage systems, erasure coding, as a redundant scheme, has received increasingly attention because it can greatly improve the space efficiency compared with the replica schemes. However, it takes a long time and consumes a lot of network bandwidth for erasure coding to repair the lost data on failed nodes. The state-of-art studies focus on the repairing optimization for the single-node-failure context. Real-world experiments have clearly shown that multi-node failures indeed happen in cloud storage systems. Borrowing single-node repairing techniques to the multi-node setting faces challenges on the efficiency. We propose a mutually cooperative recovery scheme MCRTREE based on the tree structure for multiple node failures. MCRTREE improves the bandwidth utilization and reduces the repair time by the construction of regeneration trees between each new node (denoted as newcomers) and alive nodes (denoted as providers). Further, MCRTREE reduces the size of the data volumes to be transmitted for the repair process. Numerical experiments show that MCRTREE consumes less storage cost and the maintenance bandwidth compared with other redundancy recovery schemes. Trace-driven simulation results reveal that the MCRTREE reduces the regeneration time by 30% - 50%, improves the successful regeneration probability by 10% - 20% and the data availability by 10% - 20% compared with the typical repair schemes. Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Yongquan Fu, Fangliang Xu |
NAS | 2 |
| 2014 | Feverfew: a scalable coverage-based hybrid overlay for Internet-scale pub/sub networks
Xingkong Ma, Yijie Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2014 | CommonFinder: A decentralized and privacy-preserving common-friend measurement method for the distributed online social networks
Yongquan Fu, Yijie Wang 0001, Wei Peng 0005 |
Comput. Networks | 2 |
| 2014 | Scalable and elastic event matching for attribute-based publish/subscribe systems
Xingkong Ma, Yijie Wang 0001, Qing Qiu, Xiaoqiang Pei |
Future Gener. Comput. Syst. | 2 |
| 2014 | Parallelizing skyline queries over uncertain data streams with sliding window partitioning and grid index
Xiaoyong Li 0002, Yijie Wang 0001, Xiaoling Li 0002 |
Knowl. Inf. Syst. | 2 |
| 2013 | Parallelizing Probabilistic Streaming Skyline Operator in Cloud Computing EnvironmentsabstractThe skyline query processing over uncertain data streams has received considerable attention, due to its importance in helping users make intelligent decisions over complex data. Nevertheless, existing studies only focus on retrieving the skylines over data streams in a centralized environment typically with one processor, which limits the scalability of algorithms and cannot meet the requirement for massive data analysis. The emerging cloud computing environment provides much more reliable and stable environments than the traditional distributed environments, which can be well adapted to the massive data management and complex queries. Unfortunately, existing parallel frameworks in clouds such as MapReduce and its variants are not suitable for the skyline queries over uncertain data streams. In this paper, we propose a general framework for parallelizing the probabilistic streaming skyline operator with the sliding window partitioning. Particularly, we propose four items mapping strategies CMS, AMS, DMS and APS to optimize the queries based on the proposed parallel framework. Extensive experiments with real deployment are conducted to demonstrate the effectiveness and efficiency of the proposals. Xiaoyong Li 0002, Yijie Wang 0001, Xiaoling Li 0002, Rubing Huang |
COMPSAC | 2 |
| 2013 | DKNNS: Scalable and accurate distributed K nearest neighbor search for latency-sensitive applications
Yongquan Fu, Yijie Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2013 | A general scalable and accurate decentralized level monitoring method for large-scale dynamic service provision in hybrid clouds
Yongquan Fu, Yijie Wang 0001, Ernst W. Biersack |
Future Gener. Comput. Syst. | 2 |
| 2013 | HybridNN: An accurate and scalable network location service based on the inframetric model
Yongquan Fu, Yijie Wang 0001, Ernst W. Biersack |
Future Gener. Comput. Syst. | 2 |
| 2013 | A survey of queries over uncertain data
Yijie Wang 0001, Xiaoyong Li 0002, Xiaoling Li 0002 |
Knowl. Inf. Syst. | 1 |
| 2011 | Towards estimating expected sizes of probabilistic skylines
Yongtao Yang, Yijie Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2010 | CANSE: A Churn Adaptive Approach to Network Size EstimationabstractNetwork size is one of the fundamental information of distributed applications. The approach to estimate network size must feature both high accuracy and robustness in order to adapt to the dynamic environment in different topologies. However, existing approaches fail to guarantee accuracy and robustness simultaneously in dynamic topologies due to the randomness of nodes sampled. In this paper, we propose a churn adaptive approach to network size estimation – CANSE, which collects closest nodes in identification to each node’s identification by sampling nodes periodically. Each node collects closest identifications by two schemes. One scheme is sampling random nodes from random walks along the topology. The other one is exchanging the closest identifications with other nodes. Finally, each node calculates the average spacing of the closest identifications collected to estimate network size. Compared with existing approaches, extensive experiments show that CANSE provides accurate estimation values quickly in various dynamic topologies. Xingkong Ma, Yijie Wang 0001 |
ICPADS | 2 |
| 2010 | SemanticCast: Content-Based Data Distribution over Self-Organizing Semantic Overlay NetworksabstractMany applications demand distributing data with different contents efficiently in the network environment with unreliable links and a high node churn. Existing approaches mostly focus on optimizing either efficiency or robustness of data distribution, and fail to ensure both of them simultaneously. In this paper, we propose Semantic Cast - a content-based data distribution approach over self-organizing semantic overlay networks. Semantic Cast maintains a self-organizing semantic overlay based on view exchange (called Crowd). In Crowd, each node seeks neighbors with more similar interests by periodically exchanging its neighbor list (called view) with a chosen neighbor. Through these nodes' self-organizing behavior, various interest communities emerge in the overlay. For data distribution over Crowd, Semantic Cast adopts random walk to route data between interest communities, and adopts flooding to disseminate data inside the interested communities. The experimental results show that compared to existing approaches, Semantic Cast can support efficient content-based data distribution in the unreliable and highly dynamic network environment. Yijie Wang 0001 |
PDCAT | 2 |
| 2009 | iRank: Supporting Proximity Ranking for Peer-to-Peer ApplicationsabstractProximity ranking according to end-to-end network distances (e.g., Round-Trip Time, RTT) can reveal detailed proximity information, which is important in network management and performance diagnosis in distributed systems. However, to the best of our knowledge, there has been no similar work on this subject in the P2P computing field. We present a distributed rating method iRank, that enables proximity rankings by providing discrete ratings in a distributed manner. It formulates the proximity ranking as a rating problem that faithfully captures the proximity based on noisy distance measurements scalably and practically. The primary challenge in inferring proximity rankings is enforcing distributed ratings with complex rating policies. Our solution is based on reconstructing ratings by decomposing a centralized rating method Maximum Margin Matrix Factorization (MMMF) into independent sub-problems, that can be efficiently solved in a decentralized manner. By relaxing the dependence on infrastructure nodes that are a single point of failure and limit scalability, iRank can gracefully handle network churns. Through real network latency data sets, we demonstrate that iRank can predict ratings with low distortion, which are smaller than 20 percentage worse than the centralized method, in the context of synthetic complex rating policies. Yongquan Fu, Yijie Wang 0001 |
ICPADS | 2 |
| 2009 | HyperSpring: Accurate and Stable Latency Estimation in the Hyperbolic SpaceabstractPredicting network latencies between Internet hosts can efficiently support large-scale Internet applications, e.g., file sharing service and the overlay construction. Several study use the hyperbolic space to model the Internet dense-core and many-tendril structure. However, existing hyperbolic space based embedding approaches are not designed for accurate latency estimation in the distributed context. We present HyperSpring, which estimates latency by modelling a mass spring system in the hyperbolic similar with Vivaldi. HyperSpring adopts coordinate initialization to speed up the convergence of coordinate computation, uses multiple-round symmetric updates to escape from bad local minima, and stabilizes coordinates by compensating RTT measurements to reduce the coordinate drifts. Evaluation results based on a network trace of 226 PlanetLab nodes indicate that, compared to Euclidean-space based Vivaldi, hyperspring provides performance improvements for most nodes, and incurs slightly higher distortions for a small number of nodes. Yongquan Fu, Yijie Wang 0001 |
ICPADS | 2 |
| 2009 | Anadem: A Hybrid Overlay Network for Content-Based Data DistributionabstractAs an infrastructure for data distribution, overlay networks have to feature efficient routing and adequate robustness to achieve fast and accurate data distribution in the environment with node churn. Considering that the existing overlay networks mostly focus on single optimization objective and fail to ensure routing efficiency and robustness simultaneously, a hybrid overlay network for content-based data distribution - Anadem is proposed in this paper. Anadem achieves a better compromise between routing efficiency and robustness by combining the inter-cluster multiple structured topologies with the intra-cluster unstructured topologies. Anadem also provides mechanisms for dynamic concurrent cluster creation, cluster departure and load balance to make data distribution more adaptive to the dynamic network environment. Experimental results reveal that compared with existing overlay networks, Anadem can support fast and accurate content-based data distribution even when large amount of nodes fail in the system. Yijie Wang 0001 |
ICPADS | 2 |
| 2007 | Research of Routing Algorithm in Hierarchy-Adaptive P2P Systems
Xiaoming Zhang 0001, Yijie Wang 0001, Zhoujun Li 0001 |
ISPA | 2 |
| 2007 | Key-Attributes Based Optimistic Data Consistency Maintenance Method
Yijie Wang 0001, Sikun Li |
ISPA | 2 |
| 2005 | Design and Evaluation of Network-Bandwidth-Based Parallel Replication Algorithm
Yijie Wang 0001, Yongjin Qin |
ISPA | 1 |
| 2005 | Research of Power-Aware Dynamic Adaptive Replica Allocation Algorithm in Mobile Ad Hoc Networks
Yijie Wang 0001, Kan Yang 0002 |
ISPA | 1 |
| 2005 | Research of Survival-Time-Based Dynamic Adaptive Replica Allocation Algorithm in Mobile Ad Hoc Networks
Yijie Wang 0001, Kan Yang 0002 |
NPC | 1 |
| 2003 | Dynamic Self-Adaptive Replica Location Method in Data GridsabstractWithin data grid environments, data replication is a general mechanism to improve performance and availability for distributed applications. However, it is a challenging problem to find the physical locations of multiple replicas of desired data efficiently in large-scale wide area data grid systems. In this paper, we proposed a new dynamic self-adaptive distributed replica location method - DSRL to solve the problem. In DSRL, each data element has a home node, which maintains the indices of the location information replicas. Home nodes are used to support locating multiple replicas of the same data element efficiently. Meanwhile, DSRL employs local location nodes which maintain the local replica information of data elements to support local query for local replicas. A dynamic mapping technique that can adapt to the joining or departing of home nodes is utilized to spread global replica location information evenly on location nodes. The correctness and properties of DSRL are presented and proved. Analysis and experiments show that DSRL can achieve low latency, good scalability, reliability, adaptability and ease of implementation. Dongsheng Li 0001, Nong Xiao 0001, Xicheng Lu, Yijie Wang 0001, Kai Lu 0001 |
CLUSTER | 4 |