Daniel Sun 0004

dblp:09/5042-4 · also Daniel W. Sun, Wei Sun 0004 · DBLP profile ↗
← Back
43ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-2342-7421ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 9 · 4 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Security and privacy · 6 · 3 first-authorArtificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Efficient data structures for fast and low-cost first-order logic rule mining
abstract
Logic rule mining discovers association patterns in the form of logic rules from structured data. Logic rules are widely applied in information systems to assist decisions in an interpretable way. However, too many computational resources are required in state-of-the-art systems, as most of these systems optimize rule mining algorithms from the perspectives of algorithms and architecture, while data efficiency has been overlooked. Although some start-of-the-art systems implement customized data structures to improve mining speed, the space overhead of the data structures is unaffordable when processing large-scale knowledge bases. Therefore, in this article, we propose data structures to improve data efficiency and accelerate logic rule mining. Our techniques implicitly represent the Cartesian product of variable substitutions in logic rules and build compact indices for a logic entailment cache. Furthermore, we create a pool and a lookup table for the cache so that cache components will not be repeatedly created. The evaluation results show that over 95% of memory can be reduced by our techniques, and mining procedures have been accelerated by about 20x on average. Most importantly, mining on large-scale knowledge bases is practical on normal hardware where only one thread and 20GB of memory are sufficient even for large-scale knowledge bases.
Ruoyu Wang 0004, Raymond Wong, Daniel Sun 0004
Inf. Syst.3
2025 SIB: Sorted-Integers-Based Index for Compact and Fast Caching in Top-Down Logic Rule Mining Targeting KB Compression
abstract
Background Mining logic rules from structured knowledge bases is the basis of knowledge engineering. Due to the NP‐hardness of the rule mining problem, logic rules cannot be efficiently induced from knowledge bases, especially large‐scale ones. Idea In this article, we propose a compact and efficient index structure for the maintenance of the intermediate data during top‐down rule mining, such that the memory consumption can be reduced and mining efficiency can be improved. Developing Points The index is based on a mapping from constant symbols to integers and the sorting of the mapped integers. Index update has been dissembled into four basic operations. Moreover, the index itself acts as the cache during top‐down mining. Value Most contributions in existing works employ algorithmic and architectural optimizations to improve efficiency. Data‐oriented optimizations have also been explored to some extent, but the data efficiency is relatively low, and the memory consumption is thus becoming a new challenge for state‐of‐the‐art systems. We tackle this challenge in this article, and our technique has been proven more efficient than state‐of‐the‐art systems. We evaluate our method on six datasets which contain up to 160 K records and are frequently used as benchmarks in tasks related to knowledge engineering. The experimental results show that the proposed technique speeds up the rule mining procedure by on average and reduces memory consumption by up to 70%. The space overhead of the data structure is about twice that of the indexed records, which is more than 80% lower than that of the state‐of‐the‐art technique.
Ruoyu Wang 0004, Raymond K. Wong 0001, Daniel Sun 0004, Rajiv Ranjan 0001
Softw. Pract. Exp.3
2024 Estimation-based optimizations for the semantic compression of RDF knowledge bases
abstract
Structured knowledge bases are critical for the interpretability of AI techniques. RDF KBs, which are the dominant representation of structured knowledge, are expanding extremely fast to increase their knowledge coverage, enhancing the capability of knowledge reasoning while bringing heavy burdens to downstream applications. Recent studies employ semantic compression to detect and remove knowledge redundancies via semantic models and use the induced model for further applications, such as knowledge completion and error detection. However, semantic models that are sufficiently expressive for semantic compression cannot be efficiently induced, especially for large-scale KBs, due to the hardness of logic induction. In this article, we present estimation-based optimizations for the semantic compression of RDF KBs from the perspectives of input and intermediate data involved in the induction of first-order logic rules. The negative sampling technique selects a representative subset of all negative tuples with respect to the closed-world assumption, reducing the cost of evaluating the quality of a logic rule used for knowledge inference. The number of logic inference operations used during a compression procedure is reduced by a statistical estimation technique that prunes logic rules of low quality. The evaluation results show that the two techniques are feasible for the purpose of semantic compression and accelerate the compression algorithm by up to 47x compared to the state-of-the-art system. • Negative sampling reduces the cost of a single logic inference. • Negative sampling speeds up semantic compression of RDF KBs by more than 2x. • Entailment estimation reduces the number of logic inferences in logic rule mining. • Entailment estimation speeds up semantic compression of RDF KBs by up to 47x. • Entailment estimation reduces up to 99% of memory cost of semantic compression.
Ruoyu Wang 0004, Raymond K. Wong 0001, Daniel Sun 0004
Inf. Process. Manag.3
2023 Horn rule discovery with batched caching and rule identifier for proficient compressor of knowledge data
abstract
Abstract Knowledge data has been widely applied to artificial intelligence applications for interpretable and complex reasoning. Modern knowledge bases are constructed via automatic knowledge extraction from open‐accessible sources. Thus the sizes of KBs are continuously growing, heavily burdening the maintenance and application of the knowledge data. Besides the grammatical redundancies, semantically repeated information also frequently appears in knowledge bases but is still under‐explored. Existing semantic compressors fail to efficiently discover expressive patterns and thus perform unsatisfyingly on knowledge data. This article proposes SInC, a semantic inductive compressor, to efficiently induce first‐order Horn rules and semantically compress knowledge bases. SInC improves the scalability of top‐down rule mining by batching correlated records in the cache and further optimizes the pruning of duplication and specialization via an identifier structure of Horn rules. SInC was evaluated on real‐world and synthetic datasets and compared against the state‐of‐the‐art. The results show that the batched caching speed up the rule mining procedure by more than two orders while consuming fewer than three times memory space. The identifier technique speeds up the duplication and specialization pruning by orders of magnitude with less than 5‰ and 15% error rates, respectively. SInC outperforms the state‐of‐the‐art from the perspective of overall compression on both scalability and compression effect.
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001, Rajiv Ranjan 0001
Softw. Pract. Exp.2
2023 Symbolic Minimization on Relational Data
abstract
The current wave of AI is heavily driven by data, especially for cognitive capabilities. Minimization of data semantics not only reveals core information but also becomes a guide in a wide range of domains. However, scalability is theoretically weak in pure semantic methodologies. In order to cooperate with large DBs, expressiveness is over-sacrificed in existing techniques. Thus, the quality of discovered patterns and redundancies are far from satisfactory. In this article, we formalize symbolic minimization on relational DBs and prove its NP-Completeness. A lossless technique is proposed by inducing generic first-order Horn rules that infer a subset of records from the others. More importantly, we further improve the scalability via effective caching and pruning without sacrificing the expressiveness of first-order Horn rules. A concrete system is implemented and comprehensively evaluated. Experiments show that our technique removes up to 70% contents and outperforms the state-of-the-art on minimization and scalability. The optimizations reduce up to 96% memory consumption and accelerate the performance by two orders. Our technique shows the practicality of pure semantic approaches in database mining.
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001
IEEE Trans. Knowl. Data Eng.2
2022 RDF Knowledge Base Summarization by Inducing First-Order Horn Rules
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001
ECML/PKDD (2)2
2022 SInC: Semantic approach and enhancement for relational data compression
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001, Rajiv Ranjan 0001, Albert Y. Zomaya
Knowl. Based Syst.2
2021 DuroNet: A Dual-robust Enhanced Spatial-temporal Learning Network for Urban Crime Prediction
abstract
Urban crime is an ongoing problem in metropolitan development and attracts general concern from the international community. As an effective means of defending urban safety, crime prediction plays a crucial role in patrol force allocation and public safety. However, urban crime data is a macro result of crime patterns overlapped by various irrelevant factors that cause inhomogeneous noises—local outliers and irregular waves. These noises might obstruct the learning process of crime prediction models and result in a deviation of performance. To tackle the problem, we propose a novel paradigm of Dual-robust Enhanced Spatial-temporal Learning Network (DuroNet), an encoder-decoder architecture that possesses an adaptive robustness for reducing the effect of outliers and waves. The robustness is mainly reflected on two aspects. One is a locality enhanced module that employs local temporal context information to smooth the deviation of outliers and dynamic spatial information to assist in understanding normal points. The other is a self-attention-based pattern representation module to weaken the effect of irregular waves by learning attentive weights. Finally, extensive experiments are conducted on two real-world crime datasets before and after adding Gaussian noises. The results demonstrate the superior performance of our DuroNet over the state-of-the-art methods.
Kaixi Hu, Lin Li 0001, Jianquan Liu, Daniel Sun 0004
ACM Trans. Internet Techn.4
2021 Dynamic Resource Provisioning for Sustainable Cloud Computing Systems in the Presence of Correlated Failures
abstract
Dependence of computing resources on each other in cloud computing systems (CCS) makes them prone to fail in correlated manner which significantly impacts their service reliability and energy efficiency. Focusing on these two metrics of CCS while considering correlated failures remained an open question, which is the focus of this work. This paper proposes mechanisms for improving reliability and energy efficiency jointly under correlated failures in CCS. In order to model failure correlation, statistical cluster analysis techniques are applied to real failure traces. Then, mathematical models are built to calculate reliability and energy consumption of failure prone CCS. These mathematical models are used to design fault-tolerant and energy-aware resource provisioning mechanisms/policies. In order to further reduce the energy consumption, a correlated failure-aware VM consolidation policy is also proposed in this paper. A simulation based study of the proposed resource management policies and fault tolerance mechanisms is conducted by using real failure traces and Bag-of-Tasks workload. The results demonstrate that by exploiting failure correlation with the proposed resource management policies, we reduce the occurrence of failures on tasks by 34 percent and increase the energy efficiency of the system by 20 percent, approximately in comparison to the environments where failures are handled independently.
Javid Taheri, Weisheng Si, Daniel Sun 0004, Bahman Javadi
IEEE Trans. Sustain. Comput.4
2020 Statistical Detection Of Collective Data Fraud
abstract
Statistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and therefore a more general approach is required. In data detection, statistical divergence can be used as a similarity measurement based on collective features. In this paper, we present a collective detection technique based on statistical divergence. The technique extracts distribution similarities among data collections, and then uses the statistical divergence to detect collective anomalies. Evaluation shows that it is applicable in the real world.
Ruoyu Wang 0004, Daniel Sun 0004, Guoqiang Li 0001, Raymond K. Wong 0001, Shiping Chen 0001, Jianquan Liu
ICME3
2020 Pipeline provenance for cloud-based big data analytics
abstract
Summary Provenance is information about the origin and creation of data. In data science and engineering related with cloud environment, such information is useful and sometimes even critical. In data analytics, it is necessary for making data‐driven decisions to trace back history and reproduce final or intermediate results, even to tune models and adjust parameters in a real‐time fashion. Particularly, in cloud, users need to evaluate data and pipeline trustworthiness. In this paper, we propose a solution: LogProv, toward realizing these functionalities for big data provenance, which needs to renovate data pipelines or some of big data software infrastructure to generate structured logs for pipeline events, and then stores data and logs separately in cloud space. The data are explicitly linked to the logs, which implicitly record pipeline semantics. Semantic information can be retrieved from the logs easily since they are well defined and structured beforehand. We implemented and deployed LogProv in Nectar Cloud,* associated with Apache Pig, Hadoop ecosystem, and adopted Elasticsearch to provide query service. LogProv was evaluated and empirically case studied. The results show that LogProv is efficient since the performance overhead is no more than 10%; the query can be responded within 1 second; the trustworthiness is marked clearly; and there is no impact on the data processing logic of original pipelines.
Ruoyu Wang 0004, Daniel Sun 0004, Guoqiang Li 0001, Raymond K. Wong 0001, Shiping Chen 0001
Softw. Pract. Exp.2
2020 IoT-Enabled Service for Crude-Oil Production Systems Against Unpredictable Disturbance
abstract
Internet of Things (IoT) has become a new paradigm of communication to reform traditional industries, in which distributed data automatically collected via IoT in a low cost enables many new IT services that were even impossible decades ago. This research reports on an IoT-enabled production management service for crude-oil industry. In practice, even if an optimal management decision is achieved, disruptions, such as possible oil-device failures, inclement weathers and other disturbances, arise frequently and then weaken efficiency and stability of supplement. With the help of IoT, a near-real-time management service comes into being, although the adoption of IoT brings new challenges to management of disruptions. The contributions of this article are as follows: First, a service framework is proposed for refinery which combines MQTT and Azure cloud, enabling reliable data/command delivery. Second, a smart disruption management service system is developed, which consists of monitor and alarm module, disruption management module, and rescheduling procedure module. The rescheduling procedure module takes into account the network of the refinery operations, and is easy to accommodate changes in the refinery configuration for unforeseen disruptions. The experimental results show that the proposed disruption management method balances efficiency and stability compared to traditional methods.
Qianqian Duan, Daniel Sun 0004, Guoqiang Li 0001, Genke Yang, Weiwu Yan
IEEE Trans. Serv. Comput.2
2019 Video denoising for security and privacy in fog computing
abstract
Summary To reduce heavy noise from degraded video in low or predictable latency and preserve privacy, a powerful and efficient video denoising algorithm is proposed based on fog computing for Visual Internet of Things. The conventional method is to remove noise in the cloud; however, this may overload computation and communication and raise security and privacy issues. The proposed denoising algorithm is distributed to heterogeneous devices at network edges to preserve privacy and avoid security risks as noise can be reduced in the fog rather than the cloud. To address the problems of latency, communication rate, and extremely heavy noise, structure registration, inter‐frame and inner‐frame filters, and distribution compensation are applied in the proposed algorithm. A scheme for encrypting the denoised data at network edges is provided so that security and privacy issues may be avoided during transmission and storage. Compared with other denoising approaches under extremely heavy noise conditions, the experimental results demonstrate that the proposed approach achieves superior denoising performance in terms of peak signal‐noise ratio and visual quality at low computational cost, high bandwidth efficiency, and low‐latency response in a fog computing manner.
Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001, Daniel Sun 0004, Jun Zhang 0010, Guoqiang Li 0001, Mingui Sun
Concurr. Comput. Pract. Exp.4
2019 Failure-aware energy-efficient VM consolidation in cloud computing systems
Weisheng Si, Daniel Sun 0004, Bahman Javadi
Future Gener. Comput. Syst.3
2019 Statistically managing cloud operations for latency-tail-tolerance in IoT-enabled smart cities
Daniel Sun 0004, Guoqiang Li 0001, Yuanyuan Zhang 0012, Liming Zhu 0001, Raj Gaire 0001
J. Parallel Distributed Comput.1
2019 Ada-Things: An adaptive virtual machine monitoring and migration strategy for internet of things applications
Zhong Wang 0013, Daniel Sun 0004, Guangtao Xue, Shiyou Qian, Guoqiang Li 0001, Minglu Li 0001
J. Parallel Distributed Comput.2
2019 Unsupervised blocking and probabilistic parallelisation for record matching of distributed big data
Chenxiao Dou, Daniel Sun 0004, Raymond K. Wong 0001, Muhammad Atif 0003, Guoqiang Li 0001, Rajiv Ranjan 0001
J. Supercomput.3
2019 Multi-objective Optimisation of Online Distributed Software Update for DevOps in Clouds
abstract
This article studies synchronous online distributed software update, also known as rolling upgrade in DevOps, which in clouds upgrades software versions in virtual machine instances even when various failures may occur. The goal is to minimise completion time, availability degradation, and monetary cost for entire rolling upgrade by selecting proper parameters. For this goal, we propose a stochastic model and a novel optimisation method. We validate our approach to minimise the objectives through both experiments in Amazon Web Service (AWS) and simulations.
Daniel Sun 0004, Shiping Chen 0001, Guoqiang Li 0001, Yuanyuan Zhang 0012, Muhammad Atif 0003
ACM Trans. Internet Techn.1
2018 Examine Manipulated Datasets with Topology Data Analysis: A Case Study
Yun Guo, Daniel Sun 0004, Guoqiang Li 0001, Shiping Chen 0001
ICICS2
2018 R2C: Robust Rolling-Upgrade in Clouds
abstract
Rolling upgradeis a widely-used industry technique for updating software while a service provided by multiple instances of the software remains available. In cloud deployments of software, it is usual to implement the update step for rolling upgrade by replacing entire virtual machine instances. During the process of rolling upgrade, various failures may occur due to the complexity of software stack and the uncertainties of cloud platforms. Instance health checking and replacement are standard functionalities in most cloud infrastructures, though these create uncertainty in the duration of the whole upgrade procedure. In contrast, software and configuration errors are not usually detected by infrastructure functionalities, and if these happen, the entire rolling upgrade normally is unsuccessful and the system is left in an unsuitable state. In this paper, we propose an approach, named R2C, which innovates the stat of the art with our early error detection and predictability to increase the robustness of rolling upgrade on cloud platforms. We evaluate our techniques through real life testing in Amazon Web Service (AWS) and through a simulation.
Daniel Sun 0004, Alan D. Fekete, Vincent Gramoli, Guoqiang Li 0001, Xiwei Xu 0001, Liming Zhu 0001
IEEE Trans. Dependable Secur. Comput.1
2017 Efficient Density-Based Blocking for Record Matching
abstract
Record Matching in data engineering refers to searching for data records originating from the same entities across different data sources. In practice, the main challenge of record matching is that the amount of non-matches typically far exceeds the amount of matches. This is called imbalance problem, which notoriously affects efficiency and effectiveness of matching algorithms. To solve the imbalance problem, recently, density-based blocking algorithms have been studied and demonstrated an effective blocking performance. However, the efficiency of density-based blocking approaches is not good as their effectiveness. In this paper, we improve the efficiency of density-based blocking by exploiting the idea of pre-computing and pruning. Our approach optimizes the method of computing density to speed up the blocking process. Throughout experiments on real-world datasets, the proposed approach demonstrated a high performance on both blocking efficiency and blocking effectiveness.
Chenxiao Dou, Ruoyu Wang 0004, Daniel Sun 0004, Muhammad Atif 0003
IDEAS3
2017 Active Learning with Density-Initialized Decision Tree for Record Matching
abstract
One of the fundamental problem in data management and data integration fields is Record Matching, which refers to identifying records that relate to the same entities across different data sources. In recent literature, active learning has demonstrated to be effective for record matching. One of the key steps of active learning is to build a proper initial classifier, with which active learning algorithms can quickly locate informative examples for training accurate models. However, in this process, example labelling for model training is usually expensive. Even worse, if a weak initial classifier is used, the labelling cost can be significantly increased. In this paper, we propose an unsupervised algorithm to determine the initial classifier. The process of classifier initialization requires no labelling cost. Then on our proposed algorithm, we present an active sampling method for selecting informative examples. The experiments show that our approach achieves competitive learning performance with much less labelling cost than other approaches of active learning.
Chenxiao Dou, Daniel Sun 0004, Guoqiang Li 0001, Raymond K. Wong 0001
SSDBM2
2017 A game-theoretic model and analysis of data exchange protocols for Internet of Things in clouds
Xiuting Tao, Guoqiang Li 0001, Daniel Sun 0004, Hongming Cai 0001
Future Gener. Comput. Syst.3
2017 Verifying cooperative software: A SMT-based bounded model checking approach for deterministic scheduler
Guoqiang Li 0001, Daniel Sun 0004, Yonggang Lu, Ching-Hsien Hsu
J. Syst. Archit.3
2017 Parameterization of LSB in Self-Recovery Speech Watermarking Framework in Big Data Mining
abstract
The privacy is a major concern in big data mining approach. In this paper, we propose a novel self-recovery speech watermarking framework with consideration of trustable communication in big data mining. In the framework, the watermark is the compressed version of the original speech. The watermark is embedded into the least significant bit (LSB) layers. At the receiver end, the watermark is used to detect the tampered area and recover the tampered speech. To fit the complexity of the scenes in big data infrastructures, the LSB is treated as a parameter. This work discusses the relationship between LSB and other parameters in terms of explicit mathematical formulations. Once the LSB layer has been chosen, the best choices of other parameters are then deduced using the exclusive method. Additionally, we observed that six LSB layers are the limit for watermark embedding when the total bit layers equaled sixteen. Experimental results indicated that when the LSB layers changed from six to three, the imperceptibility of watermark increased, while the quality of the recovered signal decreased accordingly. This result was a trade-off and different LSB layers should be chosen according to different application conditions in big data infrastructures.
Zhanjie Song, Wenhuan Lu, Daniel Sun 0004, Jianguo Wei
Secur. Commun. Networks4
2017 Exploiting long-term and short-term preferences and RFID trajectories in shop recommendation
abstract
Summary Shop recommendation in large shopping malls is useful in the mobile internet era. With the maturity of indoor positioning technology, customers' indoor trajectories can be captured by radio frequency identification devices readers, which provides a new way to analyze customers' potential preferences. In this paper, we design three methods for the top‐N shop recommendation problem. The first method is an improved matrix factorization method fusing estimated prior customer preference matrix that is constructed by Session‐based Temporal Graph computing. The second method is a Bayesian personalized ranking method based on the first method. The third method is by tensor decomposition combined with Session‐based Temporal Graph. Besides, we exploit customer history radio frequency identification devices trajectory information to find customers' frequent paths and revise predicted rating values to improve recommendation accuracy. Our methods are effective in modeling customers' temporal dynamics. At the same time, our approach considers repeated recommendation of the same shop by designing rating update rules. The test dataset is formed byJoyCitycustomer behavior records.JoyCityis a large‐scale modern shopping center in downtown Shanghai, China. The results show that our approaches are effective and outperform previous state‐of‐the‐art approaches. Copyright © 2016 John Wiley & Sons, Ltd.
Yue Ding 0001, Dong Wang 0024, Guoqiang Li 0001, Daniel Sun 0004, Xin Xin 0003, Shiyou Qian
Softw. Pract. Exp.4
2017 Runtime recovery actions selection for sporadic operations on public cloud
abstract
Sporadic operations such as rolling upgrade or machine instance redeployment are prone to unpredictable failures in the public cloud largely because of the inherent high variability nature of public cloud. Previous dependability research has established several recovery methods for cloud failures. In this paper, we first propose eight recovery patterns for sporadic operations on public cloud. We then present the filtering process which filters applicable recovery patterns. We propose an automation mechanism to automatically generate recovery actions for those applicable recovery patterns based on our resource state transition algorithm. We also propose a methodology to evaluate the recovery actions generated for the applicable recovery patterns based on the recovery evaluation metrics of Recovery Time, Recovery Cost, and Recovery Impact. This quantitative evaluation will lead to selection of the acceptable recovery actions. We propose two recovery actions selection mechanisms: one is based on user constraints of the recovery evaluation metrics, and the other one is based on Pareto set searching algorithm. We implement a recovery service and illustrate its applicability by recovering from errors occurring in the rolling upgrade operation on AWS cloud.
Min Fu 0001, Liming Zhu 0001, Daniel Sun 0004, Anna Liu, Leonard J. Bass, Qinghua Lu 0001
Softw. Pract. Exp.3
2016 Probabilistic parallelisation of blocking non-matched records for big data
abstract
Blocking is a technique of filtering unlikely matched pairs for record matching, which aims to collect all pairs of records that relate to the same entities across different data sources. Blocking has been broadly adopted in data mining and database. However, for big data, there is no fast and effective blocking algorithm yet, because the number of candidate pairs is tremendous between large data sets. In this paper, we report on a probabilistic parallelisation of a recently proposed blocking that is a sequential algorithm for efficient record matching in single machines. Our approach runs blocking processes distributedly on partitioned input data. In order to reduce data exchange among those blocking processes, we adopt a probabilistic technique to assure that the processes can run independently and meanwhile the aggregated result is correct with respect to common metrics. Our experimental analysis endorses the advantage of our technique and shows its novel scalability on a Hadoop MapReduce system deployed physically in a cloud.
Chenxiao Dou, Daniel Sun 0004, Guoqiang Li 0001, Jianquan Liu
IEEE BigData2
2016 LogProv: Logging events as provenance of big data analytics pipelines with trustworthiness
abstract
Provenance is information about the origin and creation of data. In data science and engineering, such information is useful and sometimes even critical. In spite of that, provenance for big data is under-explored due to the challenges from the `Vs' of big data. In data analytics, users need to query history, reproduce intermediate or final results, tune models, and adjust parameters in runtime for making data-driven decisions. In addition, users need to evaluate data and pipeline trustworthiness. Towards realising these functionalities for big data provenance, we propose a solution, called LogProv, which needs to renovate data pipelines or even some of big data software infrastructure to generate structured logs for pipeline events, and then stores data and logs separately. The data are explicitly linked to the logs, which implicitly record pipeline semantics. Semantic information can be retrieved from the logs easily since the logs are well defined and structured beforehand. We implemented LogProv in Apache Pig, and adopted ElasticSearch to provide query service. In this paper LogProv is evaluated in a Hadoop ecosystem hosted by a cloud and empirically case-studied. The results show that LogProv is efficient since the performance overhead is no more than 10%, the query can be responded within 1 second, the trustworthiness is marked clearly, and there is no impact on the data processing logic of original pipelines.
Ruoyu Wang 0004, Daniel Sun 0004, Guoqiang Li 0001, Muhammad Atif 0003, Surya Nepal
IEEE BigData2
2016 GLAP: Distributed Dynamic Workload Consolidation through Gossip-Based Learning
abstract
Dynamic virtual machine consolidation (DVMC) using live migration is one of the most promising solutions to reduce energy consumption in cloud data centers. Distributed DVMC often aggressively consolidates virtual machines (VMs) at the high expense of Service Level Agreement (SLA) of customers due to virtual machines (VMs) workload fluctuations. To alleviate this, static and adaptive threshold algorithms were proposed. However, the former is unable to predict the VMs future resource demands and the latter calculates a fixed threshold value for all physical machines (PMs) while each PM hosts VMs with different workload patterns. Moreover, both methods cannot consolidate VMs to maintain PMs in a long-term safe state. To overcome these problems, we propose a fully distributed and threshold-free DVMC algorithm called, GLAP. We combine Q-Learning with a gossip-based protocol to characterize workload patterns of VMs and take consolidation decisions. We also propose a novel two-phase distributed algorithm by which PMs unify the learned pattern which is vital for efficient execution of the algorithm. Finally, we compare GLAP experimentally against three existing techniques and show that GLAP reduces by from 43% to 78% the number of overloaded PMs under the Google Cluster VMs workload traces.
Mansour Khelghatdoust, Vincent Gramoli, Daniel Sun 0004
CLUSTER3
2016 Schedulability Analysis of Timed Regular Tasks by Under-Approximation on WCET
Bingbing Fang, Guoqiang Li 0001, Daniel Sun 0004, Hongming Cai 0001
SETTA3
2016 Unsupervised Blocking of Imbalanced Datasets for Record Matching
Chenxiao Dou, Daniel Sun 0004, Raymond K. Wong 0001
WISE (2)2
2016 Reliability and energy efficiency in cloud computing systems: Survey and taxonomy
Bahman Javadi, Weisheng Si, Daniel Sun 0004
J. Netw. Comput. Appl.4
2016 An online greedy allocation of VMs with non-increasing reservations in clouds
Yonggen Gu, Jie Tao 0001, Guoqiang Li 0001, Prem Prakash Jayaraman, Daniel Sun 0004, Rajiv Ranjan 0001, Albert Y. Zomaya, Jingti Han
J. Supercomput.6
2016 Rollup: Non-Disruptive Rolling Upgrade with Fast Consensus-Based Dynamic Reconfigurations
abstract
Rolling upgrade consists of upgrading progressively the servers of a distributed system to reduce service downtime.Upgrading a subset of servers requires a well-engineered cluster membership protocol to maintain, in the meantime, the availability of the system state. Existing cluster membership reconfigurations, like CoreOS etcd, rely on a primary not only for reconfiguration but also for storing information. At any moment, there can be at most one primary, whose replacement induces disruption. We propose Rollup, a non-disruptive rolling upgrade protocol with a fast consensus-based reconfiguration. Rollup relies on a candidate leader only for the reconfiguration and scalable biquorums for service requests. While Rollup implements a non-disruptive cluster membership protocol, it does not offer a full-fledged coordination service. We analyzed Rollup theoretically and experimentally on an isolated network of 26 physical machines and an Amazon EC2 cluster of 59 virtual machines. Our results show an 8-fold speedup compared to a rolling upgrade based on a primary for reconfiguration.
Vincent Gramoli, Leonard J. Bass, Alan D. Fekete, Daniel Sun 0004
IEEE Trans. Parallel Distributed Syst.4
2016 A Unified Business-Driven Cloud Management Framework
abstract
Cloud system management is complex due to their diversity and frequent runtime changes. Cloud systems were previously managed through cloud specific management tools that focus on optimising technical metrics, such as performance. However, business users care business metrics (such as cost and revenue) more than technical metrics. To address these issues, this paper proposes a unified business-driven cloud management framework, which enables optimisation of business metrics without limiting business to a specific cloud provider. The main contributions include: (1) a taxonomy which defines a set of actions, events and metrics for unified cloud management; (2) a cloud management policy language that specifies cloud management policies from a business perspective; and (3) middleware architecture that allows business-driven management of diverse clouds. The proposed solutions are evaluated in terms of feasibility, functional correctness, generality, usefulness, and performance.
Qinghua Lu 0001, Liming Zhu 0001, Xiwei Xu 0001, Vladimir Tosic, Dipesh Chauhan, Weishan Zhang, Daniel Sun 0004
IEEE Trans. Serv. Comput.7
2015 Multi-objective Optimisation of Rolling Upgrade Allowing for Failures in Clouds
abstract
Rolling upgrade is a practical industry technique for online updating of software in distributed systems. This paper focuses on rolling upgrade of software versions in virtual machine instances on cloud computing platforms, when various failures may occur. An operator can choose the number of instances that are updated in one round and system environments to minimise completion time, availability degradation, and monetary cost for entire rolling upgrade, and hence this is a multi-objective optimisation problem. To predict completion time in the presence of failures, we offer a stochastic model that represents the dynamics of rolling upgrade. To reduce the computational effort of decision making for large scale complex systems, we propose a technique that can find a Pareto set quickly via an upper bound of the expected completion time. Then an optimum of the original problem can be chosen from this set of potential solutions. We validate our approach to minimise the objectives, through both experiments in Amazon Web Service (AWS) and simulations.
Daniel Sun 0004, Daniel Guimarans, Alan D. Fekete, Vincent Gramoli, Liming Zhu 0001
SRDS1
2014 POD-Diagnosis: Error Diagnosis of Sporadic Operations on Cloud Applications
abstract
Applications in the cloud are subject to sporadic changes due to operational activities such as upgrade, redeployment, and on-demand scaling. These operations are also subject to interferences from other simultaneous operations. Increasing the dependability of these sporadic operations is non-trivial, particularly since traditional anomaly-detection-based diagnosis techniques are less effective during sporadic operation periods. A wide range of legitimate changes confound anomaly diagnosis and make baseline establishment for "normal" operation difficult. The increasing frequency of these sporadic operations (e.g. due to continuous deployment) is exacerbating the problem. Diagnosing failures during sporadic operations relies heavily on logs, while log analysis challenges stemming from noisy, inconsistent and voluminous logs from multiple sources remain largely unsolved. In this paper, we propose Process Oriented Dependability (POD)-Diagnosis, an approach that explicitly models these sporadic operations as processes. These models allow us to (i) determine orderly execution of the process, and (ii) use the process context to filter logs, trigger assertion evaluations, visit fault trees and perform on-demand assertion evaluation for online error diagnosis and root cause analysis. We evaluated the approach on rolling upgrade operations in Amazon Web Services (A WS) while performing other simultaneous operations. During our evaluation, we correctly detected all of the 160 injected faults, as well as 46 interferences caused by concurrent operations. We did this with 91.95% precision. Of the correctly detected faults, the accuracy rate of error diagnosis is 96.55%.
Xiwei Xu 0001, Liming Zhu 0001, Ingo Weber, Leonard J. Bass, Daniel Sun 0004
DSN5
2009 A Novel Genetic Admission Control for Real-Time Multiprocessor Systems
abstract
In real-time multiprocessor systems, admission control must respond to the requests of tasks quickly, otherwise some requests have to be rejected even if there are enough resources to meet the requirements of tasks running. Real-time task scheduling in multiprocessor systems has been proved to be NP-hard problems. Genetic algorithms (GAs) are known as a class of effective tools to solve NP problems, but the execution time of GAs is usually very long. In this paper, we present a novel approach of using genetic algorithms in real-time task admission control and scheduling. First our approach preserves the population of the GA and tries to use the historical information, i. e., the previous task schedules, to shorten the search time, instead of destroying and creating a population respectively after and before dealing with a new task in the standard GA procedure; second our approach dynamically updates the chromosomes in the population in terms of task arrivals and departures in order to repeatedly reuse the preserved population. Through simulations, it is demonstrated that our approach can rapidly make admission decisions and produce task schedules, meanwhile with satisfying task acceptance ratio and low preemption frequency.
Daniel Sun 0004
PDCAT1
2008 Predict task running time in grid environments based on CPU load predictions
Yuanyuan Zhang 0012, Daniel Sun 0004, Yasushi Inoguchi
Future Gener. Comput. Syst.2
2007 Real-time Task Scheduling Using Extended Overloading Technique for Multiprocessor Systems
abstract
The scheduling of real-time tasks with fault-tolerant requirements has been an important problem in multiprocessor systems. Primary-backup (PB) approach is often used as a fault-tolerant technique to guarantee the deadlines of tasks despite the presence of faults. In this paper we propose a PB-based task scheduling approach, wherein an allocation parameter is used to search the available time slots for a newly arriving task, and the previously scheduled tasks can be rescheduled when there is no available time slot for the newly arriving task. In order to improve the schedulability we extend the existing PB-overloading and the Backup-backup (BB) overloading. Our proposed task scheduling algorithm is compared with some existing scheduling algorithms in the literature through simulation studies. The results have shown that the task rejection ratio of our real-time task scheduling algorithm is lower than the compared algorithms.
Daniel Sun 0004, Yuanyuan Zhang 0012, Xavier Défago, Yasushi Inoguchi
DS-RT1
2007 Hybrid Overloading and Stochastic Analysis for Redundant Real-time Multiprocessor Systems
abstract
In multiprocessor systems, redundant scheduling is a technique that trades processing power for increased reliability. One approach, called primary-backup task scheduling, is often used in real-time multiprocessor systems to ensure that deadlines are met in spite of faults. Briefly, it consists in scheduling a secondary task conditionally, in such a way that the secondary task actually gets executed only if the primary task (or the processor executing it) fails to terminate properly. Doing so avoids wasting CPU resources in the failure-free case, but primary and secondary tasks must then compete for resources in case of failure. To overcome this, overloading strategies, such as primary and backup overloading (PB) and backup-backup overloading (BB), aim at improving schedulability while retaining a certain level of reliability. In this paper, we propose a hybrid overloading technique based on extended PB overloading, which combines advantages of both PB and BB overloading. The three overloading strategies are then compared through a stochastic analysis, and by simulating them under diverse system conditions. The analysis shows that hybrid overloading provides an excellent tradeoff between schedulability and reliability.
Daniel Sun 0004, Yuanyuan Zhang 0012, Xavier Défago, Yasushi Inoguchi
SRDS1
2006 CPU Load Predictions on the Computational Grid
abstract
The ability to accurately predict future resource capabilities is of great importance for applications and scheduling algorithms which need to determine how to use time-shared resources in a dynamic grid environment. In this paper we present and evaluate a new and innovative method to predict the one-stepahead CPU load in a grid. Our prediction strategy forecasts the future CPU load based on the tendency in several past steps and in previous similar patterns, and uses a polynomial fitting method. Our experimental results demonstrate that this new prediction strategy achieves average prediction errors that are between 37% and 86% lower than those incurred by the previously best tendency-based method.
Yuanyuan Zhang 0012, Daniel Sun 0004, Yasushi Inoguchi
CCGRID2