VLDB 2026 Research / reviewers in the wild / expert
Anna Liu
dblp:89/5066
· DBLP profile ↗
57ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 17 · 2 since 2021Systems, architecture and hardware · 8Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Artificial intelligence and machine learning · 4Security and privacy · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Query processing and optimization · 47% Machine learning and data management · 27% Data stream processing · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 91% Distributed systems · 5% Performance modeling and evaluation · 4% | |
| Software engineering, system software, and programming languages
5 papers |
Services computing and microservices · 52% Requirements engineering and software design · 47% Software testing · 2% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 28 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management
active learning |
1.1 | 2 | 2024 | Efficient and robust active learning methods for interactive database exploration · VLDB J. 2024 A Factorized Version Space Algorithm for "Human-In-the-Loop" Data Exploration · ICDM 2019 |
Query processing and optimization
interactive data exploration |
1.1 | 2 | 2024 | Efficient and robust active learning methods for interactive database exploration · VLDB J. 2024 Optimization for Active Learning-based Interactive Database Exploration · Proc. VLDB Endow. 2018 |
Query processing and optimization
query optimization |
0.3 | 1 | 2018 | Optimization for Active Learning-based Interactive Database Exploration · Proc. VLDB Endow. 2018 |
Data stream processing
uncertain data stream |
0.3 | 2 | 2012 | CLARO: modeling and processing uncertain data streams · VLDB J. 2012 Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Data mining
exploratory data analysis |
0.2 | 1 | 2024 | Efficient and robust active learning methods for interactive database exploration · VLDB J. 2024 |
Cloud and datacenter computing › resource allocation › dynamic resource allocation
adaptive resource allocation |
0.2 | 1 | 2015 | A Framework for Consumer-Centric SLA Management of Cloud-Hosted Databases · IEEE Trans. Serv. Comput. 2015 |
Cloud and datacenter computing › resource management
cloud resource management |
0.2 | 1 | 2015 | A Framework for Consumer-Centric SLA Management of Cloud-Hosted Databases · IEEE Trans. Serv. Comput. 2015 |
Cloud and datacenter computing › cloud service management
service level agreement management |
0.2 | 1 | 2015 | A Framework for Consumer-Centric SLA Management of Cloud-Hosted Databases · IEEE Trans. Serv. Comput. 2015 |
Requirements engineering and software design
software architecture |
0.2 | 5 | 2011 | An architects guide to enterprise application integration with J2EE and .NET · ICSE 2005 Architectures and Technologies for Enterprise Application Integration · ICSE 2004 Third international workshop on principles of engineering service-oriented systems: (PESOS 2011) · ICSE 2011 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.2 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Data models and query languages
uncertain data |
0.2 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Query processing and optimization › query execution
user-defined function execution |
0.2 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Data stream processing
continuous query processing |
0.1 | 1 | 2012 | CLARO: modeling and processing uncertain data streams · VLDB J. 2012 |
Query processing and optimization
probabilistic query processing |
0.1 | 1 | 2011 | Optimizing Probabilistic Query Processing on Continuous Uncertain Data · Proc. VLDB Endow. 2011 |
Query processing and optimization
uncertain data query processing |
0.1 | 1 | 2011 | Optimizing Probabilistic Query Processing on Continuous Uncertain Data · Proc. VLDB Endow. 2011 |
Machine learning › Probabilistic and Bayesian machine learning
distribution approximation |
0.1 | 1 | 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Query processing and optimization
aggregate query processing |
0.1 | 1 | 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Services computing and microservices
enterprise application integration |
0.1 | 2 | 2005 | An architects guide to enterprise application integration with J2EE and .NET · ICSE 2005 Architectures and Technologies for Enterprise Application Integration · ICSE 2004 |
Data models and query languages
uncertain data management |
0.0 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Services computing and microservices
service-oriented architecture |
0.0 | 1 | 2004 | Architectures and Technologies for Enterprise Application Integration · ICSE 2004 |
Requirements engineering and software design › software architecture
component-based software engineering |
0.0 | 1 | 2002 | Software component quality assessment in practice: successes and practical impediments · ICSE 2002 |
Data mining
anomaly detection |
0.0 | 1 | 2010 | PODS: a new model and processing algorithms for uncertain data streams · SIGMOD Conference 2010 |
Query processing and optimization
approximate query processing |
0.0 | 1 | 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Distributed systems
distributed object systems |
0.0 | 1 | 2000 | DeBOT - an approach for constructing high performance, scalable distributed object systems (poster) · ICSE 2000 |
Requirements engineering and software design
requirements elicitation |
0.0 | 1 | 2005 | An architects guide to enterprise application integration with J2EE and .NET · ICSE 2005 |
Services computing and microservices
middleware |
0.0 | 1 | 2002 | Software component quality assessment in practice: successes and practical impediments · ICSE 2002 |
Software testing › test generation
automated test generation |
0.0 | 1 | 2001 | Generation of Distributed System Test-Beds from High-Level Software Architecture Descriptions · ASE 2001 |
Distributed systems
middleware |
0.0 | 1 | 2001 | Generation of Distributed System Test-Beds from High-Level Software Architecture Descriptions · ASE 2001 |
Methods — techniques the papers use, named apart from their topics
active learning · 1.1version space factorization · 0.4online algorithm · 0.3gaussian process · 0.3virtualization-based replication · 0.2randomized approximation · 0.2policy-based scaling · 0.2deterministic approximation · 0.2probabilistic modeling · 0.1statistical approximation · 0.1sampling · 0.1code generation · 0.1experience report · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Efficient Version Space Algorithms for Human-in-the-loop Model DevelopmentabstractWhen active learning (AL) is applied to help users develop a model on a large dataset through interactively presenting data instances for labeling, existing AL techniques often suffer from two main drawbacks: First, to reach high accuracy they may require the user to label hundreds of data instances, which is an onerous task for the user. Second, retrieving the next instance to label from a large dataset can be time-consuming, making it incompatible with the interactive nature of the human exploration process. To address these issues, we introduce a novel version-space-based active learner for kernel classifiers, which possesses strong theoretical guarantees on performance and efficient implementation in time and space. In addition, by leveraging additional insights obtained in the user labeling process, we can factorize the version space to perform active learning in a set of subspaces, which further reduces the user labeling effort. Evaluation results show that our algorithms significantly outperform state-of-the-art version space strategies, as well as a recent factorization-aware algorithm, for model development over large datasets. Luciano Di Palma, Yanlei Diao, Anna Liu |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Efficient and robust active learning methods for interactive database exploration
Enhui Huang, Yanlei Diao, Anna Liu, Liping Peng, Luciano Di Palma |
VLDB J. | 3 |
| 2022 | Message from the ICSA 2022 General Chairs and Program ChairsabstractThe IEEE International Conference on Software Architecture is the premier venue for practitioners and researchers interested in software architecture, in component-based software engineering and in quality aspects of software and how these relate to the design of software. ICSA has a strong tradition as a working conference, where researchers meet practicing software architects who can explain the problems they face in their day-to-day duties, and who can influence the future of the field. Rick Kazman, Patrizio Pelliccione, Anna Liu, Ingo Weber |
ICSA | 3 |
| 2022 | Integrating bioinformatic strategies in spatial life science researchabstractAs space exploration programs progress, manned space missions will become more frequent and farther away from Earth, putting a greater emphasis on astronaut health. Through the collaborative efforts of researchers from various countries, the effect of the space environment factors on living systems is gradually being uncovered. Although a large number of interconnected research findings have been produced, their connection seems to be confused, and many unknown effects are left to be discovered. Simultaneously, several valuable data resources have emerged, accumulating data measuring biological effects in space that can be used to further investigate the unknown biological adaptations. In this review, the previous findings and their correlations are sorted out to facilitate the understanding of biological adaptations to space and the design of countermeasures. The biological effect measurement methods/data types are also organized to provide references for experimental design and data analysis. To aid deeper exploration of the data resources, we summarized common characteristics of the data generated from longitudinal experiments, outlined challenges or caveats in data analysis and provided corresponding solutions by recommending bioinformatics strategies and available models/tools. Yangyang Hao, Anna Liu, Xiaoyue Kong, Fengji Liang, Jianghui Xiong, Lina Qu |
Briefings Bioinform. | 3 |
| 2022 | HLA3D: an integrated structure-based computational toolkit for immunotherapyabstractMOTIVATION: The human major histocompatibility complex (MHC), also known as human leukocyte antigen (HLA), plays an important role in the adaptive immune system by presenting non-self-peptides to T cell receptors. The MHC region has been shown to be associated with a variety of diseases, including autoimmune diseases, organ transplantation and tumours. However, structural analytic tools of HLA are still sparse compared to the number of identified HLA alleles, which hinders the disclosure of its pathogenic mechanism. RESULT: To provide an integrative analysis of HLA, we first collected 1296 amino acid sequences, 256 protein data bank structures, 120 000 frequency data of HLA alleles in different populations, 73 000 publications and 39 000 disease-associated single nucleotide polymorphism sites, as well as 212 modelled HLA heterodimer structures. Then, we put forward two new strategies for building up a toolkit for transplantation and tumour immunotherapy, designing risk alignment pipeline and antigenic peptide prediction pipeline by integrating different resources and bioinformatic tools. By integrating 100 000 calculated HLA conformation difference and online tools, risk alignment pipeline provides users with the functions of structural alignment, sequence alignment, residue visualization and risk report generation of mismatched HLA molecules. For tumour antigen prediction, we first predicted 370 000 immunogenic peptides based on the affinity between peptides and MHC to generate the neoantigen catalogue for 11 common tumours. We then designed an antigenic peptide prediction pipeline to provide the functions of mutation prediction, peptide prediction, immunogenicity assessment and docking simulation. We also present a case study of hepatitis B virus mutations associated with liver cancer that demonstrates the high legitimacy of our antigenic peptide prediction process. HLA3D, including different HLA analytic tools and the prediction pipelines, is available at http://www.hla3d.cn/. Xueyin Mei, Pin Chen, Anna Liu, Weicheng Liang, Shan Chang |
Briefings Bioinform. | 5 |
| 2019 | A Factorized Version Space Algorithm for "Human-In-the-Loop" Data ExplorationabstractWhile active learning (AL) has been recently applied to help the user explore a large database to retrieve data instances of interest, existing methods often require a large number of instances to be labeled in order to achieve good accuracy. To address this slow convergence problem, our work augments version space-based AL algorithms, which have strong theoretical results on convergence but are very costly to run, with additional insights obtained in the user labeling process. These insights lead to a novel algorithm that factorizes the version space to perform active learning in a set of subspaces. Our work offers theoretical results on optimality and approximation for this algorithm, as well as optimizations for better performance. Evaluation results show that our factorized version space algorithm significantly outperforms other version space algorithms, as well as a recent factorization-aware algorithm, for large database exploration. Luciano Di Palma, Yanlei Diao, Anna Liu |
ICDM | 3 |
| 2018 | Optimization for Active Learning-based Interactive Database ExplorationabstractThere is an increasing gap between fast growth of data and limited human ability to comprehend data. Consequently, there has been a growing demand of data management tools that can bridge this gap and help the user retrieve high-value content from data more effectively. In this work, we aim to build interactive data exploration as a new database service, using an approach called "explore-by-example". In particular, we cast the explore-by-example problem in a principled "active learning" framework, and bring the properties of important classes of database queries to bear on the design of new algorithms and optimizations for active learning-based database exploration. These new techniques allow the database system to overcome a fundamental limitation of traditional active learning, i.e., the slow convergence problem. Evaluation results using real-world datasets and user interest patterns show that our new system significantly outperforms state-of-the-art active learning techniques and data exploration systems in accuracy while achieving desired efficiency for interactive performance. Enhui Huang, Liping Peng, Luciano Di Palma, Ahmed Abdelkafi, Anna Liu, Yanlei Diao |
Proc. VLDB Endow. | 5 |
| 2017 | Runtime recovery actions selection for sporadic operations on public cloudabstractSporadic operations such as rolling upgrade or machine instance redeployment are prone to unpredictable failures in the public cloud largely because of the inherent high variability nature of public cloud. Previous dependability research has established several recovery methods for cloud failures. In this paper, we first propose eight recovery patterns for sporadic operations on public cloud. We then present the filtering process which filters applicable recovery patterns. We propose an automation mechanism to automatically generate recovery actions for those applicable recovery patterns based on our resource state transition algorithm. We also propose a methodology to evaluate the recovery actions generated for the applicable recovery patterns based on the recovery evaluation metrics of Recovery Time, Recovery Cost, and Recovery Impact. This quantitative evaluation will lead to selection of the acceptable recovery actions. We propose two recovery actions selection mechanisms: one is based on user constraints of the recovery evaluation metrics, and the other one is based on Pareto set searching algorithm. We implement a recovery service and illustrate its applicability by recovering from errors occurring in the rolling upgrade operation on AWS cloud. Min Fu 0001, Liming Zhu 0001, Daniel Sun 0004, Anna Liu, Leonard J. Bass, Qinghua Lu 0001 |
Softw. Pract. Exp. | 4 |
| 2016 | Process-Oriented Non-intrusive Recovery for Sporadic Operations on CloudabstractCloud-based systems get changed more frequently than traditional systems. These frequent changes involve sporadic operations such as installation and upgrade. Sporadic operations may fail due to the uncertainty of cloud platforms. Each sporadic operation manipulates a number of cloud resources. The accessibility of resources manipulated makes it possible to build an accurate process model of the correct behavior for an operation and its desired effects. This paper proposes a non-intrusive recovery approach for sporadic operations on cloud, called POD-Recovery. POD-Recovery utilizes the above-mentioned process model of the operation. When needed, it triggers recovery actions based on the model through non-intrusive means, i.e., without modifying the code which implements the sporadic operation. POD-Recovery employs an efficient artificial intelligence (AI) planning technique for generating recovery plans. We implement POD-Recovery and evaluate it by recovering from faults injected into 920 runs of five representative sporadic operations. Min Fu 0001, Liming Zhu 0001, Ingo Weber, Leonard J. Bass, Anna Liu, Xiwei Xu 0001 |
DSN | 5 |
| 2015 | Evaluating the impact of fine-scale burstiness on cloud elasticityabstractElasticity is the defining feature of cloud computing. Performance analysts and adaptive system designers rely on representative benchmarks for evaluating elasticity for cloud applications under realistic reproducible workloads. A key feature of web workloads is burstiness or high variability at fine timescales. In this paper, we explore the innate interaction between fine-scale burstiness and elasticity and quantify the impact from the cloud consumer's perspective. We propose a novel methodology to model workloads with fine-scale burstiness so that they can resemble the empirical stylized facts of the arrival process. Through an experimental case study, we extract insights about the implications of fine-scale burstiness for elasticity penalty and adaptive resource scaling. Our findings demonstrate the detrimental effect of fine-scale burstiness on the elasticity of cloud applications. Sadeka Islam, Srikumar Venugopal, Anna Liu |
SoCC | 3 |
| 2015 | A Framework for Consumer-Centric SLA Management of Cloud-Hosted DatabasesabstractService Level Agreements (SLA) represent the contract which captures the agreed upon guarantees between a service provider and its customers. The specifications of existing service level agreements (SLA) for cloud services are not designed to flexibly handle even relatively straightforward performance and technical requirements of consumer applications. In this article, we present a novel approach for SLA-based management of cloud-hosted databases from the consumer perspective. We present an end-to-end framework for consumer-centric SLA management of cloud-hosted databases. The framework facilitates adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. In this framework, the SLA of the consumer applications are declaratively defined in terms of goals which are subjected to a number of constraints that are specific to the application requirements. The framework continuously monitors the application-defined SLA and automatically triggers the execution of necessary corrective actions (scaling out/in the database tier) when required. The framework is database platform-agnostic, uses virtualization-based database replication mechanisms, and requires zero source code changes of the cloud-hosted software applications. The experimental results demonstrate the effectiveness of our SLA-based framework in providing the consumer applications with the required flexibility for achieving their SLA requirements. Liang Zhao 0009, Sherif Sakr, Anna Liu |
IEEE Trans. Serv. Comput. | 3 |
| 2014 | Towards a Taxonomy of Cloud Recovery StrategiesabstractRecovering from failures of sporadic operations such as rolling upgrade or migration is complicated by the fact that the application being upgraded or migrated must continue to provide service. This means that recovery strategies for sporadic operations must include facilities for recovering from normal operations as well. As a step in deriving methods for recovering from failures in sporadic operations, we classify existing methods into four categories according to their purposes and the life cycle phase for which they are applicable. Not only does this taxonomy facilitate the research on recoverability of cloud sporadic operations but also it can help better understand the existing cloud recovery strategies. Min Fu 0001, Leonard J. Bass, Anna Liu |
DSN | 3 |
| 2014 | Recovery for Failures in Rolling Upgrade on CloudsabstractWhen cloud consumers perform rolling upgrade operations on cloud applications, they may encounter failures due to cloud uncertainty, interfering operations and incorrect configurations. For example, unreliable cloud API calls can make the rolling upgrade operation fail in unpredictable ways due to a long time delay to respond to the API call. This paper proposes two recovery strategies for recovering from rolling upgrade failures. The strategies are Compensated Undo & Redo and Reparation. We evaluated our recovery strategies on Asgard-based rolling upgrade operation on Amazon Cloud based on two evaluation metrics: MTTR and Service Performance. The experiment results show that our strategies perform better than the recovery mechanisms provided by Asgard itself. We also conduct a comparison between the two recovery strategies based on the metrics. Min Fu 0001, Liming Zhu 0001, Leonard J. Bass, Anna Liu |
DSN | 4 |
| 2014 | Consumer Monitoring of Infrastructure Performance in a Public Cloud
Rabia Chaudry, Adnene Guabtni, Alan D. Fekete, Leonard J. Bass, Anna Liu |
WISE (2) | 5 |
| 2014 | GEAP: A Generic Approach to Predicting Workload Bursts for Web Hosted Events
Matthew Sladescu, Alan D. Fekete, Anna Liu |
WISE (2) | 4 |
| 2013 | Improving Availability of Cloud-Based Applications through Deployment ChoicesabstractDeployment choices are critical in determining the availability of applications running in a cloud. But choosing good deployment for various software application components into virtual machines is a challenging task because of potential sharing of components among applications and potential interference from multi-tenancy. This paper presents an approach for improving the availability guarantee of software applications by optimizing the availability, performance and monetary cost trade-offs of different deployment choices. Our approach explicitly considers different classes of application requests during the decision process. The results of our experimental evaluation show that the approach can effectively improve the availability guarantees with little or negligible increase in the performance and monetary cost of the deployment choice. Jim Zhanwen Li, Qinghua Lu 0001, Liming Zhu 0001, Leonard J. Bass, Xiwei Xu 0001, Sherif Sakr, Paul L. Bannerman, Anna Liu |
IEEE CLOUD | 8 |
| 2013 | Incorporating Uncertainty into In-Cloud Application Deployment Decisions for AvailabilityabstractCloud consumers have a variety of deployment related techniques, such as auto-scaling policies and recovery strategies, for dealing with the uncertainties in the cloud. Uncertainties can be characterized as stochastic (such as failures, disasters, and workload spikes) and subjective (such as choice among various deployment options). Cloud consumers must consider both stochastic and subjective uncertainties. Analytic support for consumers in selecting appropriate techniques and setting the required parameters in the face of different types of uncertainty is currently limited. In this paper, we propose a set of application availability analysis models that capture subjective uncertainties in addition to stochastic uncertainties. We built and validated the models by using industry best practices on deployment, and actual commercial products for disaster recovery and live migration. Our results show that the models permit more informed and quantitative availability analysis than industry best practices under a wide range of scenarios. Qinghua Lu 0001, Xiwei Xu 0001, Liming Zhu 0001, Leonard J. Bass, Jim Zhanwen Li, Sherif Sakr, Paul L. Bannerman, Anna Liu |
IEEE CLOUD | 8 |
| 2013 | Consumer-centric SLA manager for cloud-hosted databasesabstractWe present an end-to-end framework for consumer-centric SLA management of virtualized database servers. The framework facilitates adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. In this framework, the SLA of the consumer applications are declaratively defined in terms of goals which are subjected to a number of constraints that are specific to the application requirements. The framework continuously monitors the application-defined SLA and automatically triggers the execution of necessary corrective actions (scaling out/in the database tier) when required. The framework is database platform-agnostic, uses virtualization-based database replication mechanisms and requires zero source code changes of the cloud-hosted application. Liang Zhao 0009, Sherif Sakr, Anna Liu |
CIKM | 3 |
| 2013 | Process-oriented recovery for operations on cloud applicationsabstractA large number of cloud application failures happen during sporadic operations on cloud applications, such as upgrade, deployment reconfiguration, migration and scaling-out/in. Most of them are caused by operator and process errors [1]. From a cloud consumer's perspective, recovery from these failures relies on the limited control and visibility provided by the cloud providers. In addition, a large-scale system often has multiple operation processes happening simultaneously, which exacerbates the problem during error diagnosis and recovery. Existing built-in or infrastructure-based recovery mechanisms often assume random component failures and use checkpoint-based rollback, compensation actions [2], redundancy and rejuvenation to handle recovery [3]. These recovery mechanisms do not consider the characteristics of a specific operation process that consists of a set of steps carried out by scripts and humans interacting with fragile cloud infrastructure APIs and uncertain resources [4]. Other approaches such as FATE/DESTINI [5] look at the process implied by a system's internal protocols and rely on the built-in recovery protocol to detect and recover from bugs. The problem we target is at a different level related to the external sporadic activities operating on a hosted cloud application. Min Fu 0001, Liming Zhu 0001, Anna Liu, Xiwei Xu 0001, Leonard J. Bass |
SoCC | 3 |
| 2013 | Challenges to Error Diagnosis in Hadoop Ecosystems
Jim Zhanwen Li, Liming Zhu 0001, Xiwei Xu 0001, Min Fu 0001, Leonard J. Bass, Anna Liu, An Binh Tran |
LISA | 7 |
| 2013 | Supporting Undoability in Systems Operations
Ingo Weber, Hiroshi Wada, Alan D. Fekete, Anna Liu, Leonard J. Bass |
LISA | 4 |
| 2013 | Is Your Cloud-Hosted Database Truly Elastic?abstractElasticity has been recognized as one of the most appealing features for users of cloud services. It represents the ability to dynamically and rapidly scale up or down the allocated computing resources on demand. In practice, it is difficult to understand the elasticity requirements of a given application and workload, and to assess if the elasticity provided by a cloud service will meet these requirements. In this experience paper, we take the position that a deep understanding of the capabilities of cloud-hosted database services is a crucial requirement for cloud users in order to bring forward the vision of deploying data-intensive applications on cloud platforms. We argue that it is important that cloud users become able to paint a comprehensive picture of the relationship between the capabilities of the different type of cloud database services, the application characteristics and workloads, and the geographical distribution of the application clients and the underlying database replicas. We discuss the current elasticity capabilities of the different categories of cloud database services and identify some of the main challenges for deploying a truly elastic database tier on cloud environments. Finally, we propose a benchmarking mechanism that can evaluate the elasticity capabilities of cloud database services in different application scenarios and workloads. Sherif Sakr, Anna Liu |
SERVICES | 2 |
| 2013 | Supporting User-Defined Functions on Uncertain DataabstractUncertain data management has become crucial in many sensing and scientific applications. As user-defined functions (UDFs) become widely used in these applications, an important task is to capture result uncertainty for queries that evaluate UDFs on uncertain data. In this work, we provide a general framework for supporting UDFs on uncertain data. Specifically, we propose a learning approach based on Gaussian processes (GPs) to compute approximate output distributions of a UDF when evaluated on uncertain input, with guaranteed error bounds. We also devise an online algorithm to compute such output distributions, which employs a suite of optimizations to improve accuracy and performance. Our evaluation using both real-world and synthetic functions shows that our proposed GP approach can outperform the state-of-the-art sampling approach with up to two orders of magnitude improvement for a variety of UDFs. Thanh T. L. Tran, Yanlei Diao, Charles Sutton, Anna Liu |
Proc. VLDB Endow. | 4 |
| 2012 | SLA-Based and Consumer-centric Dynamic Provisioning for Cloud DatabasesabstractOne of the main advantages of the cloud computing paradigm is that it simplifies the time-consuming processes of hardware provisioning, hardware purchasing and software deployment. Currently, we are witnessing a proliferation in the number of cloud-hosted applications with a tremendous increase in the scale of the data generated as well as being consumed by such applications. Cloud-hosted database systems powering these applications form a critical component in the software stack of these applications. Service Level Agreements (SLA) represent the contract which captures the agreed upon guarantees between a service provider and its customers. The specifications of existing service level agreement (SLA) for cloud services are not designed for flexibly handling even relatively straightforward performance and technical requirements of consumer applications. The concerns of consumers for cloud services regarding the SLA management of their hosted applications within the cloud environments will gain increasing importance as cloud computing becomes more pervasive. This paper introduces the notion, challenges and the importance of SLA-based provisioning and cost management for cloud-hosted databases from the consumer perspective. We present an end-to-end framework that acts as a middleware which resides between the consumer applications and the cloud-hosted databases. The aim of the framework is to facilitate adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. The experimental results demonstrate that SLA-based provisioning is more adequate for providing consumer applications the required flexibility in achieving their goals. Sherif Sakr, Anna Liu |
IEEE CLOUD | 2 |
| 2012 | Application-Managed Replication Controller for Cloud-Hosted DatabasesabstractData replication is a well-known strategy to achieve the availability, scalability and performance improvement goals in the data management world. However, the cost of maintaining several database replicas always strongly consistent is very high. The CAP theorem shows that a shared-data system can choose at most two out of three properties: consistency, availability, and tolerance to partitions. In practice, most of the cloud-based data management systems tend to overcome the difficulties of distributed replication by relaxing the consistency guarantees of the system. In particular, they implement various forms of weaker consistency models such as eventual consistency. This solution is accepted by many new Web 2.0 applications (e.g. social networks) which could be more tolerant with a wider window of data staleness (replication delay).However, unfortunately, there are no generic application-independent and consumer-centric mechanisms by which software applications can specify and manage to what extent inconsistencies can be tolerated. We introduce an adaptive framework for database replication at the middleware layer of cloud environments. The framework provides flexible mechanisms to enable software applications of keeping several database replicas (that can be hosted in different data centers) with different levels of service level agreements (SLA) for their data freshness. The experimental evaluation demonstrates the effectiveness of our framework in providing the software applications with the required flexibility to achieve and optimize their requirements in terms of overall system throughput, data freshness and invested monetary cost. Liang Zhao 0009, Sherif Sakr, Anna Liu |
IEEE CLOUD | 3 |
| 2012 | Event Aware Workload Prediction: A Study Using Auction Events
Matthew Sladescu, Alan D. Fekete, Anna Liu |
WISE | 4 |
| 2012 | How a consumer can measure elasticity for cloud platformsabstractOne major benefit claimed for cloud computing is elasticity: the cost to a consumer of computation can grow or shrink with the workload. This paper offers improved ways to quantify the elasticity concept, using data available to the consumer. We define a measure that reflects the financial penalty to a particular consumer, from under-provisioning (leading to unacceptable latency or unmet demand) or over-provisioning (paying more than necessary for the resources needed to support a workload). We have applied several workloads to a public cloud; from our experiments we extract insights into the characteristics of a platform that influence its elasticity. We explore the impact of the rules used to increase or decrease capacity. Sadeka Islam, Alan D. Fekete, Anna Liu |
ICPE | 4 |
| 2012 | Empirical prediction models for adaptive resource provisioning in the cloud
Sadeka Islam, Jacky W. Keung, Anna Liu |
Future Gener. Comput. Syst. | 4 |
| 2012 | CLARO: modeling and processing uncertain data streams
Thanh T. L. Tran, Liping Peng, Yanlei Diao, Andrew McGregor 0001, Anna Liu |
VLDB J. | 5 |
| 2011 | Data Consistency Properties and the Trade-offs in Commercial Cloud Storage: the Consumers' Perspective
Hiroshi Wada, Alan D. Fekete, Liang Zhao 0009, Anna Liu |
CIDR | 5 |
| 2011 | Size Estimation of Cloud Migration Projects with Cloud Migration Point (CMP)abstractOne major obstacle to enterprise adoption of cloud technologies has been the lack of visibility into migration effort and cost. In this paper, we present a methodology, called Cloud Migration Point (CMP), for estimating the size of cloud migration projects, by recasting a well-known software size estimation model called Function Point (FP) into the context of cloud migration. We empirically evaluate our CMP model by performing a cross-validation on six different small-scale cloud migration projects and show that our size estimation model can be used as a reliable predictor for effort estimation. Furthermore, we prove that our CMP model satisfies the fundamental properties of a software size measure. Van T. K. Tran, Alan D. Fekete, Anna Liu, Jacky W. Keung |
ESEM | 4 |
| 2011 | Third international workshop on principles of engineering service-oriented systems: (PESOS 2011)abstractService-oriented systems have attracted great interest from industry and research communities worldwide. Service integrators, developers, and providers are collaborating to address the various challenges in the field. PESOS 2011 is a forum for all these communities to present and discuss a wide range of topics related to service-oriented systems. The goal of PESOS is to bring together researchers from academia and industry, as well as practitioners working in the areas of software engineering and service-oriented systems to discuss research challenges, recent developments, novel applications, as well as methods, techniques, experiences, and tools to support the engineering of service-oriented systems. Manuel Carro, Dimka Karastoyanova, Grace A. Lewis, Anna Liu |
ICSE | 4 |
| 2011 | CloudDB AutoAdmin: Towards a Truly Elastic Cloud-Based Data StoreabstractIn this paper, we present the design and the architecture of the CloudDB AutoAdmin system which aims to fill the existing gaps between the provided cloud database services and the requirements of the consumer applications. In particular, it focuses on facilitating the job of the cloud database consumers in implementing database applications as distributed, scalable, and elastic services with a minimum effort on the side of the application developer and a limited footprint in the application code. Sherif Sakr, Liang Zhao 0009, Hiroshi Wada, Anna Liu |
ICWS | 4 |
| 2011 | Healing Web-Based Services on Cloud Using StrategiesabstractExceptions cause temporal failures of Web-based systems to deliver continuous and responsive end-to-end services. On cloud platforms, resource provisioning and administration become less explicit, which entails techniques that can act on exceptions according to known strategies. In this paper, we propose a middleware-based approach that encompasses a healing process that links available strategies to exceptions. Our solution is implemented as an embedded middleware called Exception Healing Manager (EHM). The EHM ensures non-intrusive service recovery and improves overall service response time. This approach is demonstrated by a loan broker service running on Amazon EC2. Markus Lachat, Anna Liu |
SERVICES | 3 |
| 2011 | Architecting Cloud Computing Applications and SystemsabstractIt is with great pleasure and privilege that we are organising the 1st Workshop at WICSA on Architecting Cloud Computing Applications and Systems. Cloud Computing is a popular topic in the IT industry right now, it is timely that we have the opportunity to share research results and practitioner experiences at WICSA, particularly focusing on the various issues related to the enterprise consumer perspective of cloud computing, where adoption is still nascent, and many software architecture research challenges remain. Our workshop agenda in this first year will address topics such as architectural styles for multi-tenant systems, assessing cloud platforms technical benefits for application development, modeling and reasoning design alternatives of software-as-a-service architectures, and engineering preprioception in SLA management for cloud architectures. We hope this will be the first of a successful series of workshops on Architecting Cloud Computing Applications at WICSA. Anna Liu, Rainbow Cai |
WICSA | 1 |
| 2011 | Optimizing Probabilistic Query Processing on Continuous Uncertain Data
Liping Peng, Yanlei Diao, Anna Liu |
Proc. VLDB Endow. | 3 |
| 2010 | Evaluating Cloud Platform Architecture with the CARE FrameworkabstractThere is an emergence of Cloud application platforms such as Microsoft's Azure, Google's App Engine and Amazon's EC2/SimpleDB/S3. Startups and Enterprise alike, lured by the promise of `infinite scalability', `ease of development', `low infrastructure setup cost' are increasingly using these Cloud service building blocks to develop and deploy their web based applications. However, the precise nature of these Cloud platforms and the resultant Cloud application runtime behavior is still largely an unknown. Given the black box nature of these platforms, and the novel programming and data models of Cloud, there is a dearth of tools and techniques for enabling the rigorously evaluation of Cloud platforms at runtime. This paper introduces the CARE (Cloud Architecture Runtime Evaluation) approach, a framework for evaluating Cloud application development and runtime platforms. CARE implements a unified interface with WSDL and REST in order to evaluate different Cloud platforms for Cloud application hosting servers and Cloud databases. With the unified interface, we are able to perform selective high stress and low stress evaluations corresponding to desired test scenarios. Result shows the effectiveness of CARE in the evaluation of Cloud variations in terms of scalability, availability and responsiveness, across both compute and storage capabilities. Thus placing CARE as an important tool in the path of Cloud computing research. Liang Zhao 0009, Anna Liu, Jacky W. Keung |
APSEC | 2 |
| 2010 | PODS: a new model and processing algorithms for uncertain data streamsabstractUncertain data streams, where data is incomplete, imprecise, and even misleading, have been observed in many environments. Feeding such data streams to existing stream systems produces results of unknown quality, which is of paramount concern to monitoring applications. In this paper, we present the PODS system that supports stream processing for uncertain data naturally captured using continuous random variables. PODS employs a unique data model that is flexible and allows efficient computation. Built on this model, we develop evaluation techniques for complex relational operators, i.e., aggregates and joins, by exploring advanced statistical theory and approximation. Evaluation results show that our techniques can achieve high performance while satisfying accuracy requirements, and significantly outperform a state-of-the-art sampling method. A case study further shows that our techniques can enable a tornado detection system (for the first time) to produce detection results at stream speed and with much improved quality. Thanh T. L. Tran, Liping Peng, Boduo Li, Yanlei Diao, Anna Liu |
SIGMOD Conference | 5 |
| 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond ExpectationsabstractUncertain data streams are increasingly common in real-world deployments and monitoring applications require the evaluation of complex queries on such streams. In this paper, we consider complex queries involving conditioning (e.g., selections and group by's) and aggregation operations on uncertain data streams. To characterize the uncertainty of answers to these queries, one generally has to compute the full probability distribution of each operation used in the query. Computing distributions of aggregates given conditioned tuple distributions is a hard, unsolved problem. Our work employs a new evaluation framework that includes a general data model, approximation metrics, and approximate representations. Within this framework we design fast data-stream algorithms, both deterministic and randomized, for returning approximate distributions with bounded errors as answers to those complex queries. Our experimental results demonstrate the accuracy and efficiency of our approximation techniques and offer insights into the strengths and limitations of deterministic and randomized algorithms. Thanh T. L. Tran, Andrew McGregor 0001, Yanlei Diao, Liping Peng, Anna Liu |
Proc. VLDB Endow. | 5 |
| 2009 | Capturing Data Uncertainty in High-Volume Stream Processing
Yanlei Diao, Boduo Li, Anna Liu, Liping Peng, Charles Sutton, Thanh T. L. Tran, Michael Zink |
CIDR | 3 |
| 2009 | COBRA - mining web for COrporate Brand and Reputation AnalysisabstractCorporations are extremely sensitive to issues such as brand stewardship and product reputation. Traditional brand image and reputation tracking is limited to news wires and contact centre analysis. However, with the emergence of web, Consumer Genera W. Scott Spangler, Ying Chen 0001, Larry Proctor, Ana Lelescu, Amit Behal, Bin He 0001, Thomas D. Griffin, Anna Liu, Brad Wade, Trevor Davis 0002 |
Web Intell. Agent Syst. | 8 |
| 2007 | COBRA - Mining Web for Corporate Brand and Reputation AnalysisabstractCorporations are extremely sensitive to issues such as brand stewardship and product reputation. Traditional brand image and reputation tracking is limited to news wires and contact centres analysis. However, with the emergence of Web, consumer generated media (COM), such as blogs, news forums, message boards, and Web pages/sites, is rapidly becoming the "voice of the people". This paper describes a COBRA (corporate brand and reputation analysis) solution that mines a wide range of COM contents for brand and reputation analysis. The solution contains a flexible ETL (Extract, Transform, and Load) engine that processes diverse sets of structured and unstructured information, a suite of analytical capabilities that mines COM content to extract semantic entities and insights out of the data, and an alerting mechanism that utilizes the analytics results to accurately generate brand and reputation alerts. We use a real-world case study to demonstrate the effectiveness of our approach. W. Scott Spangler, Ying Chen 0001, Larry Proctor, Ana Lelescu, Amit Behal, Bin He 0001, Thomas D. Griffin, Anna Liu, Brad Wade, Trevor Davis 0002 |
Web Intelligence | 8 |
| 2007 | Bayesian meta-analysis models for microarray data: a comparative studyabstractBACKGROUND: With the growing abundance of microarray data, statistical methods are increasingly needed to integrate results across studies. Two common approaches for meta-analysis of microarrays include either combining gene expression measures across studies or combining summaries such as p-values, probabilities or ranks. Here, we compare two Bayesian meta-analysis models that are analogous to these methods. RESULTS: Two Bayesian meta-analysis models for microarray data have recently been introduced. The first model combines standardized gene expression measures across studies into an overall mean, accounting for inter-study variability, while the second combines probabilities of differential expression without combining expression values. Both models produce the gene-specific posterior probability of differential expression, which is the basis for inference. Since the standardized expression integration model includes inter-study variability, it may improve accuracy of results versus the probability integration model. However, due to the small number of studies typical in microarray meta-analyses, the variability between studies is challenging to estimate. The probability integration model eliminates the need to model variability between studies, and thus its implementation is more straightforward. We found in simulations of two and five studies that combining probabilities outperformed combining standardized gene expression measures for three comparison values: the percent of true discovered genes in meta-analysis versus individual studies; the percent of true genes omitted in meta-analysis versus separate studies, and the number of true discovered genes for fixed levels of Bayesian false discovery. We identified similar results when pooling two independent studies of Bacillus subtilis. We assumed that each study was produced from the same microarray platform with only two conditions: a treatment and control, and that the data sets were pre-scaled. CONCLUSION: The Bayesian meta-analysis model that combines probabilities across studies does not aggregate gene expression measures, thus an inter-study variability parameter is not included in the model. This results in a simpler modeling approach than aggregating expression measures, which accounts for variability across studies. The probability integration model identified more true discovered genes and fewer true omitted genes than combining expression measures, for our data sets. Erin M. Conlon, Joon J. Song, Anna Liu |
BMC Bioinform. | 3 |
| 2005 | An architects guide to enterprise application integration with J2EE and .NETabstractArchitects are faced with the problem of building enterprise scale information systems, with streamlined, automated internal business processes and web-enabled business functions, all across multiple legacy applications. The underlying architectures for such systems are embodied in a range of diverse products known as Enterprise Application Integration (EAI) technologies. In this tutorial, we highlight some of the major problems, approaches and issues in designing EAI architectures and selecting appropriate supporting technology. An architect's perspective on designing large-scale integrated applications is taken, and we discuss requirements elicitation, architecture patterns, EAI technology and features, and risk mitigation. J2EE and .NET technologies are used to illustrate the capabilities of state-or-the-art integration technologies. Ian Gorton, Anna Liu |
ICSE | 2 |
| 2005 | Automatic Performance Tuning for J2EE Application Server Systems
Anna Liu |
WISE | 3 |
| 2005 | SoftArch/MTE: Generating Distributed System Test-Beds from High-Level Software Architecture Descriptions
John C. Grundy, Yuhong Cai, Anna Liu |
Autom. Softw. Eng. | 3 |
| 2005 | Performance prediction of component-based applications
Shiping Chen 0001, Yan Liu 0001, Ian Gorton, Anna Liu |
J. Syst. Softw. | 4 |
| 2004 | Architectures and Technologies for Enterprise Application IntegrationabstractArchitects are faced with the problem of building enterprise scale information systems, with streamlined, automated internal business processes and Web-enabled business functions, all across multiple legacy applications. The underlying architectures for such systems are embodied in a range of diverse products known as enterprise application integration (EAI) technologies. In this paper, we highlight some of the major problems, approaches and issues in designing EAI architectures and selecting appropriate supporting technology. The tutorial presents a range of the common architectural patterns frequently used for EAI applications. It also explains service-oriented architectures as the current best practice architectural framework for EAI. It then describes the state-or-the-art in EAI technologies that support these architectural styles, and discusses some of the key design trade-offs involved when selecting an appropriate integration technology (including buy versus build decisions). Ian Gorton, Anna Liu |
ICSE | 2 |
| 2003 | Workshop on Architectures for Complex Application Integration (WACAI 2003)abstractThe first Workshop on Architectures for Complex Application Integration (WACAI 2003) was held in Dallas, Texas on 5 November 2003. This was a joint workshop with COMPSAC 2003. The major aim of the workshop was to bring together key research ideas and practices in the field of evolving architectures for application integration, with the goal of assessing the state of the art and of identifying the fundamental open issues in the research field. Six papers are accepted for formal presentation. In the last section, a summary is outlined for each paper. Ian Gorton, Anna Liu |
COMPSAC | 3 |
| 2002 | Evaluating the Scalability of Enterprise JavaBeans TechnologyabstractOne of the major problems in building large-scale distributed systems is to anticipate the performance of the eventual solution before it has been built. This problem is especially germane to Internet-based e-business applications, where failure to provide high performance and scalability can lead to application and business failure. The fundamental software engineering problem is compounded by many factors, including individual application diversity, software architecture trade-offs, COTS component integration requirements, and differences in performance of various software and hardware infrastructures. We describe the results of an empirical investigation into the scalability of a widely used distributed component technology, Enterprise JavaBeans (EJB). A benchmark application is developed and tested to measure the performance of a system as both the client load and component infrastructure are scaled up. A scalability metric from the literature is then applied to analyze the scalability of the EJB component infrastructure under two different architectural solutions. Yan Jenny Liu, Ian Gorton, Anna Liu, Shiping Chen 0001 |
APSEC | 3 |
| 2002 | Software component quality assessment in practice: successes and practical impedimentsabstractThis paper describes the authors' experiences of initiating and sustaining a project at CSIRO aimed at accelerating the successful adoption of COTS middleware technologies in large business and scientific information systems. The projects aims are described, along with example outcomes and an assessment of what is needed for wide-scale software component quality assessments to succeed. Ian Gorton, Anna Liu |
ICSE | 2 |
| 2001 | Generation of Distributed System Test-Beds from High-Level Software Architecture DescriptionsabstractMost distributed system specifications have performance benchmark requirements. However, determining the likely performance of complex distributed system architectures during development is very challenging. We describe a system where software architects sketch an outline of their proposed system architecture at a high level of abstraction, including indicating client requests, server services, and choosing particular kinds of middleware and database technologies. A fully working implementation of this system is then automatically generated, allowing multiple clients and servers to be run. Performance tests are then automatically run for this generated code and results are displayed back in the original high-level architectural diagrams. Architects may change performance parameters and architecture characteristics, comparing multiple test run results to determine the most suitable abstractions to refine to detailed designs for actual system implementation. We demonstrate the utility of this approach and the accuracy of our generated performance test-beds for validating architectural choices during early system development. John C. Grundy, Yuhong Cai, Anna Liu |
ASE | 3 |
| 2000 | DeBOT - an approach for constructing high performance, scalable distributed object systems (poster)abstractThe Internet creates new opportunities for component distribution. Infrastructure for dynamic, Web-based composition of software components appears to be a very impelling need. This demonstration focuses on a Web-based system that supports dynamic component composition. Anna Liu |
ICSE | 1 |
| 2000 | Special issue on constructing software engineering tools
Jonathan Gray, Anna Liu, Louise Scott |
Inf. Softw. Technol. | 2 |
| 2000 | Issues in software engineering tool construction
Jonathan Gray, Anna Liu, Louise Scott |
Inf. Softw. Technol. | 2 |
| 1999 | The First International Symposium on Constructing Software Engineering Tools (CoSET'99)abstractNo abstract available. Jonathan Gray, Louise Scott, Anna Liu, Jennifer Harvey |
ICSE | 3 |
| 1999 | Guest Editorial
Anna Liu, Paddy Nixon |
Softw. Qual. J. | 1 |