Patrick Martin 0001

dblp:m/PatrickMartin · also T. Patrick Martin · DBLP profile ↗
← Back
66ranked-venue papers
16as first author
0since 2021 · last 2018
0000-0003-3210-9441ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 26 · 10 first-authorComputer networks · 11 · 2 first-authorArtificial intelligence and machine learning · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 9Applied, interdisciplinary, general and emerging computing · 9 · 2 first-authorSystems, architecture and hardware · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Database system architecture and tuning · 83% Data stream processing · 14% Query processing and optimization · 3%
Software engineering, system software, and programming languages
2 papers
Services computing and microservices · 60% Debugging and program repair · 40%
Artificial intelligence
2 papers
Kernel, tree and ensemble methods · 69% Multi-agent systems · 21% Time series and sequential data · 10%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 75% Parallel and multicore computing · 18% Performance modeling and evaluation · 5%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning
workload management
0.722018
Workload Management in Database Management Systems: A Taxonomy · IEEE Trans. Knowl. Data Eng. 2018
Workload Management in Database Management System: A Taxonomy (Extended Abstract) · ICDE 2018
Debugging and program repair
post-deployment debugging
0.212013
Assisting developers of big data analytics applications when deploying on hadoop clouds · ICSE 2013
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.112011
Classification Using Streaming Random Forests · IEEE Trans. Knowl. Data Eng. 2011
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.112011
Classification Using Streaming Random Forests · IEEE Trans. Knowl. Data Eng. 2011
Data stream processing › stream mining
stream classification
0.112011
Classification Using Streaming Random Forests · IEEE Trans. Knowl. Data Eng. 2011
Services computing and microservices
automated negotiation
0.112011
An Adaptive and Intelligent SLA Negotiation System for Web Services · IEEE Trans. Serv. Comput. 2011
Services computing and microservices
service level agreement
0.112011
An Adaptive and Intelligent SLA Negotiation System for Web Services · IEEE Trans. Serv. Comput. 2011
Database system architecture and tuning
commercial database systems
0.112018
Workload Management in Database Management System: A Taxonomy (Extended Abstract) · ICDE 2018
Cloud and datacenter computing
cloud deployment
0.012013
Assisting developers of big data analytics applications when deploying on hadoop clouds · ICSE 2013
Knowledge, reasoning and agents › Multi-agent systems
automated negotiation
0.012011
An Adaptive and Intelligent SLA Negotiation System for Web Services · IEEE Trans. Serv. Comput. 2011
Machine learning › Time series and sequential data › non-stationary environments
concept drift
0.012011
Classification Using Streaming Random Forests · IEEE Trans. Knowl. Data Eng. 2011
Knowledge, reasoning and agents › Multi-agent systems
intelligent agents
0.012011
An Adaptive and Intelligent SLA Negotiation System for Web Services · IEEE Trans. Serv. Comput. 2011
Query processing and optimization
join processing
0.011994
Parallel Hash-Based Join Algorithms for a Shared-Everything · IEEE Trans. Knowl. Data Eng. 1994
Query processing and optimization › join processing › parallel join
parallel hash join
0.011994
Parallel Hash-Based Join Algorithms for a Shared-Everything · IEEE Trans. Knowl. Data Eng. 1994
Parallel and multicore computing › parallel query processing
parallel join algorithms
0.011994
Parallel Hash-Based Join Algorithms for a Shared-Everything · IEEE Trans. Knowl. Data Eng. 1994
Information retrieval
distributed information retrieval
0.011986
A Design of a Distributed Full Text Retrieval System · SIGIR 1986
Performance modeling and evaluation
cost modeling
0.011994
Parallel Hash-Based Join Algorithms for a Shared-Everything · IEEE Trans. Knowl. Data Eng. 1994
Distributed systems
remote procedure call
0.011986
A Design of a Distributed Full Text Retrieval System · SIGIR 1986

Methods — techniques the papers use, named apart from their topics

taxonomy · 0.7log analysis · 0.3execution sequence comparison · 0.3time-based decision functions · 0.2streaming random forest · 0.2simulation · 0.2online learning · 0.2game-theoretic bargaining · 0.2concept drift adaptation · 0.2experimental validation · 0.0cost modeling · 0.0x.400 · 0.0remote procedure call · 0.0
YearPublicationVenuePosition
2018 Workload Management in Database Management System: A Taxonomy (Extended Abstract)
abstract
Workload management is the discipline of effectively monitoring, managing and controlling work flow across computing systems. In particular, workload management in database management systems (DBMSs) is the process or act of monitoring and controlling work (i.e., requests) executing on a database system in order to make efficient use of system resources in addition to achieving any performance objectives assigned to that work. In the past decade, workload management studies and practice have made considerable progress in both academia and industry. New techniques have been proposed by researchers, and new features of workload management facilities have been implemented in most commercial database products. In this paper, we provide a systematic study of workload management in today's DBMSs by developing a taxonomy of workload management techniques. We apply the taxonomy to evaluate and classify existing workload management techniques implemented in the commercial databases and available in the recent research literature. We also introduce the underlying principles of today's workload management technology for DBMSs, discuss open problems and outline some research opportunities in this research area.
Mingyi Zhang 0001, Patrick Martin 0001, Wendy Powley, Jianjun Chen 0001
ICDE2
2018 Workload Management in Database Management Systems: A Taxonomy
abstract
Workload management is the discipline of effectively monitoring, managing and controlling work flow across computing systems. In particular, workload management in database management systems (DBMSs) is the process or act of monitoring and controlling work (i.e., requests) executing on a database system in order to make efficient use of system resources in addition to achieving any performance objectives assigned to that work. In the past decade, workload management studies and practice have made considerable progress in both academia and industry. New techniques have been proposed by researchers, and new features of workload management facilities have been implemented in most commercial database products. In this paper, we provide a systematic study of workload management in today's DBMSs by developing a taxonomy of workload management techniques. We apply the taxonomy to evaluate and classify existing workload management techniques implemented in the commercial databases and available in the recent research literature. We also introduce the underlying principles of today's workload management technology for DBMSs, discuss open problems, and outline some research opportunities in this research area.
Mingyi Zhang 0001, Patrick Martin 0001, Wendy Powley, Jianjun Chen 0001
IEEE Trans. Knowl. Data Eng.2
2017 Leveraging distributed big data storage support in CLAaaS for WINGS workflow management system
abstract
Cloud-based Analytics-as-a-Service (CLAaaS) was developed by Zulkernine et al. with a goal to simplifying big data analytics users. It provides software-as-a-service access to a variety of back end analytics tools and data stores. One of the tools is the Workflow Instance Generation and Selection (WINGS). WINGS allows users to reuse predefined workflows and their components containing semantic meta-data to define new workflows; late binding of the workflows to data at the time of execution to enable the use of most recent data, and definition of domain specific software code as custom analytic components in workflows. How ever, the data used in WINGS for the workflows are mostly flat files that are stored on the WINGS server or shared directories. The goal of this project is to add support for big data storage systems to WINGS and validate the extensions using multiple data analytic workflows of different complexities with data residing in a variety of back end data sources. The extension allows the CLAaaS users to create, validate and execute analytic workflows in a distributed environment and use data from multiple big data storage systems. We validate our work using four big data storage systems in WINGS workflows namely, Apache HBase, MongoDB, MySQL with a front-end interface.
Hadeel Alghamdi, Farhana Zulkernine, Patrick Martin 0001
IEEE BigData3
2016 BINARY: A framework for big data integration for ad-hoc querying
abstract
Enormous amounts of data are generated everyday of both structured and unstructured nature. Regardless of their differences, data sources must be used in tandem in any effective big data operation. This paper proposes a Software as a Service (SaaS) framework called BINARY which provides a back-end infrastructure for ad-hoc querying, accessing, visualizing and joining data from different data sources such as Relational Database Management Systems like MySQL and big data storage systems like Apache Hive. BINARY is extendable and allows adding other storage engines (e.g. HBase) and analytics engines (e.g. R) as needed. A REST software architecture is used in the framework to enable loose connections between the engines and user interface programs to facilitate their independent updates without affecting the data infrastructure. Our approach is validated with a proof-of-concept prototype implemented on the OpenStack cloud system.
Azadeh Eftekhari, Farhana Zulkernine, Patrick Martin 0001
IEEE BigData3
2016 Personal mobile services
Khalid Elgazzar, Patrick Martin 0001, Hossam S. Hassanein
Serv. Oriented Comput. Appl.2
2016 Cloud-Assisted Computation Offloading to Support Mobile Services
abstract
The widespread use and increasing capabilities of mobiles devices are making them a viable platform for offering mobile services. However, the increasing resource demands of mobile services and the inherent constraints of mobile devices limit the quality and type of functionality that can be offered, preventing mobile devices from exploiting their full potential as reliable service providers. Computation offloading offers mobile devices the opportunity to transfer resource-intensive computations to more resourcefulcomputing infrastructures. We present a framework for cloud-assisted mobile service provisioning to assist mobile devices in delivering reliable services. The framework supports dynamic offloading based on the resource status of mobile systems and current network conditions, while satisfying the user-defined energy constraints. It also enables the mobile provider to delegate the cloud infrastructure to forward the service response directly to the user when no further processing is required by the provider. Performance evaluation shows up to 6x latency improvement for computation-intensive services that do not require large data transfer. Experiments show that the operation of the cloud-assisted service provisioning framework does not pose significant overhead on mobile resources, yet it offers robust and efficient computation offloading.
Khalid Elgazzar, Patrick Martin 0001, Hossam S. Hassanein
IEEE Trans. Cloud Comput.2
2015 goDiscovery: Web Service Discovery Made Efficient
abstract
The growing popularity of cloud computing has magnified the rise of software reuse by facilitating service provisioning over the Internet. At the same time, a new generation of mobile apps has emerged relying on backend services that expand the app functionally, while reducing the overhead on limited mobile resources. The Web service approach promises great flexibility in offering software functionality over the network, while maintaining interoperability between heterogeneous platforms. In addition, recent years have witnessed the rise of user-facing service developments that can be consumed on-the-go with a standard interface, such as Restful Web services. However, the discovery of such services does not match their growing popularity and remain challenging. Users cannot tolerate long latency in finding relevant services to their requests. In this paper, we propose a robust and efficient Web service discovery approach that uses statistical methods and indexing techniques to improve the precision and response time of the discovery process. Experimental results demonstrate that the proposed approach outperforms the state-of-the-art discovery mechanisms and significantly reduces the query response time by at least 77%, while maintaining comparable accuracy.
Yehia Elshater, Khalid Elgazzar, Patrick Martin 0001
ICWS3
2014 Secure and Efficient Data Placement in Mobile Healthcare Services
Anne V. D. M. Kayem, Khalid Elgazzar, Patrick Martin 0001
DEXA (1)3
2014 A mobile-based architecture for integrating personal health record data
abstract
Personal Health Record (PHR) systems provide patients with access to their own records, as well as control over who accesses their record. There are many PHR system providers available on the market. These PHR systems, however, have little means to integrate with healthcare facilities in the healthcare system network. This paper proposes a Personal Health Record (PHR) system solution which allows for exchange of patient data at the point-of-care using the patient's mobile device. The objective is to outline and address the issues that arise when adopting an hybrid PHR architecture that comprises a mobile component and an online remote server component. Preliminary tests are conducted in order to assess the system's usability.
Muhammad AboElFotoh, Patrick Martin 0001, Hossam S. Hassanein
Healthcom2
2014 Near-clouds: Bringing public clouds to users' doorsteps
abstract
The rate of data growth in many domains is straining our ability to manage and analyze it. Cloud computing appears as a promising platform for data-intensive computing because it offers “infinite” resources on demand, and on a pay-as-you-go basis. Surprisingly, we observe that public clouds are not being used for “serious” data-processing on a continuous basis except by the cloud vendors. We primarily attribute this observation to the long data transfer times over wide-area networks between the clients and the centralized public clouds. We introduce the idea of a near-cloud that brings a public cloud in the close proximity of a user to overcome the data transfer bottlenecks. The near-clouds also provide a unique opportunity to localize the privacy and security policies to boost confidence of a user in using shared resources for data processing. To the cloud providers, near-cloud ensures a minimum constant stream of revenue from dedicated clients.
Rizwan Mian, Khalid Elgazzar, Shady Khalifa, Patrick Martin 0001, Gabriel Silberman, Daniel Goldschmidt
ISCC4
2014 DaaS: Cloud-based mobile Web service discovery
Khalid Elgazzar, Hossam S. Hassanein, Patrick Martin 0001
Pervasive Mob. Comput.3
2014 Cloud Service Negotiation in Internet of Things Environment: A Mixed Approach
abstract
Internet of Things (IoT) allows connected objects to communicate via the Internet. IoT can benefit from the unlimited capabilities and resources of cloud computing. Also, when coupled with IoT, cloud computing can in turn deal with real world things in a more distributed and dynamic manner. As the cloud market becomes more open and competitive, Quality of Service (QoS) will be more important. However, cloud providers and cloud consumers have different, and sometimes opposite, preferences. If such a conflict occurs, a Service Level Agreement (SLA) cannot be reached without negotiation. A tradeoff negotiation approach can outperform a concession approach in terms of utility, but may incur more failures if information is incomplete. To balance utility and success rate, we propose a mixed approach for cloud service negotiation, which is based on the “game of chicken.” In particular, if one is uncertain about the strategy of its counterpart, it is best to mix concession and tradeoff strategies in negotiation. To evaluate the effectiveness of this approach, we conduct extensive simulations. Results show that a mixed negotiation approach can achieve a higher utility than a concession approach, while incurring fewer failures than a tradeoff approach.
Xianrong Zheng, Patrick Martin 0001, M. Kathryn Brohman
IEEE Trans. Ind. Informatics2
2014 CLOUDQUAL: A Quality Model for Cloud Services
abstract
Cloud computing is an important component of the backbone of the Internet of Things (IoT). Clouds will be required to support large numbers of interactions with varying quality requirements. Service quality will therefore be an important differentiator among cloud providers. In order to distinguish themselves from their competitors, cloud providers should offer superior services that meet customers' expectations. A quality model can be used to represent, measure, and compare the quality of the providers, such that a mutual understanding can be established among cloud stakeholders. In this paper, we take a service perspective and initiate a quality model named CLOUDQUAL for cloud services. It is a model with quality dimensions and metrics that targets general cloud services. CLOUDQUAL contains six quality dimensions, i.e., usability, availability, reliability, responsiveness, security, and elasticity, of which usability is subjective, whereas the others are objective. To demonstrate the effectiveness of CLOUDQUAL, we conduct empirical case studies on three storage clouds. Results show that CLOUDQUAL can evaluate their quality. To demonstrate its soundness, we validate CLOUDQUAL with standard criteria and show that it can differentiate service quality.
Xianrong Zheng, Patrick Martin 0001, M. Kathryn Brohman
IEEE Trans. Ind. Informatics2
2013 Assisting developers of big data analytics applications when deploying on hadoop clouds
abstract
Big data analytics is the process of examining large amounts of data (big data) in an effort to uncover hidden patterns or unknown correlations. Big Data Analytics Applications (BDA Apps) are a new type of software applications, which analyze big data using massive parallel processing frameworks (e.g., Hadoop). Developers of such applications typically develop them using a small sample of data in a pseudo-cloud environment. Afterwards, they deploy the applications in a large-scale cloud environment with considerably more processing power and larger input data (reminiscent of the mainframe days). Working with BDA App developers in industry over the past three years, we noticed that the runtime analysis and debugging of such applications in the deployment phase cannot be easily addressed by traditional monitoring and debugging approaches. In this paper, as a first step in assisting developers of BDA Apps for cloud deployments, we propose a lightweight approach for uncovering differences between pseudo and large-scale cloud deployments. Our approach makes use of the readily-available yet rarely used execution logs from these platforms. Our approach abstracts the execution logs, recovers the execution sequences, and compares the sequences between the pseudo and cloud deployments. Through a case study on three representative Hadoop-based BDA Apps, we show that our approach can rapidly direct the attention of BDA App developers to the major differences between the two deployments. Knowledge of such differences is essential in verifying BDA Apps when analyzing big data in the cloud. Using injected deployment faults, we show that our approach not only significantly reduces the deployment verification effort, but also provides very few false positives when identifying deployment failures.
Weiyi Shang, Zhen Ming (Jack) Jiang, Hadi Hemmati, Bram Adams, Ahmed E. Hassan, Patrick Martin 0001
ICSE6
2013 Personalized Mobile Web Service Discovery
abstract
Mobile devices with their various form factors have become the most convenient and pervasive computing platform, whether to carry out everyday business or to get online. Mobile users tend to adopt the fast food trend even in consuming online mobile services and functionalities. The Web service approach promises great flexibility in offering software functionality over the network, while maintaining interoperability between heterogeneous platforms. However, the diversity that exists in mobile devices and their platforms with variations in capabilities present unique challenges in developing services that can accommodate such diversity. Recent years have witnessed the rise of user-facing service developments that can be consumed on the go with standard interface, such as RESTful Web services. However, the discovery of such services does not match their growing popularity. In addition, existing discovery approaches lack supporting mechanisms that ensure the proper functioning of discovered services within the user context and failing to match personal preferences. This paper introduces personalized Web service discovery for mobile environments. Preliminary results show that incorporating user preferences and context significantly improves the overall precision of service discovery.
Khalid Elgazzar, Patrick Martin 0001, Hossam S. Hassanein
SERVICES2
2013 Towards personal mobile Web services
abstract
This paper introduces personal mobile Web services, a new user-centric architecture that enables service-oriented interactions among mobile devices that are controlled via user-specified authorization policies. Personal mobile Web services exploit the user's contact list (ranging from phonebook to social lists) in order to publish and discover Web services while placing users in full control of their own personal data and privacy. We present a proof-of-concept implementation of an example personal mobile Web service to demonstrate the usefulness and feasibility of the concept.
Khalid Elgazzar, Hossam S. Hassanein, Patrick Martin 0001
WCNC3
2013 Towards building performance models for data-intensive workloads in public clouds
abstract
The cloud computing paradigm provides the "illusion" of infinite resources and, therefore, becomes a promising candidate for large-scale data-intensive computing. In this paper, we explore experiment-driven performance models for data-intensive workloads executing in an infrastructure-as-a-service (IaaS) public cloud. The performance models help in predicting the workload behaviour, and serve as a key component of a larger framework for resource provisioning in the cloud. We determine a suitable prediction technique after comparing popular regression methods. We also enumerate the variables that impact variance in the workload performance in a public cloud. Finally, we build a performance model for a multi-tenant data service in the Amazon cloud. We find that a linear classifier is sufficient in most cases. On a few occasions, a linear classifier is unsuitable and non-linear modeling is required, which is time consuming. Consequently, we recommend that a linear classifier be used in training the performance model in the first instance. If the resulting model is unsatisfactory, then non-linear modeling can be carried out in the next step.
Rizwan Mian, Patrick Martin 0001, Farhana Zulkernine, José Luis Vázquez-Poletti
ICPE2
2013 Provisioning data analytic workloads in a cloud
Rizwan Mian, Patrick Martin 0001, José Luis Vázquez-Poletti
Future Gener. Comput. Syst.2
2012 IDSaaS: Intrusion Detection System as a Service in Public Clouds
abstract
In a public cloud computing environment, consumers cannot always just depend on the cloud provider's security infrastructure. They may need to monitor and protect their virtual existence by implementing their own intrusion detection capabilities along with other security technologies within the cloud fabric. Intrusion Detection as a Service (IDSaaS) targets security of the infrastructure level of a public cloud (IaaS) by providing intrusion detection technology that is highly elastic, portable and fully controlled by the cloud consumer. A prototype of IDSaaS is described.
Turki Alharkan, Patrick Martin 0001
CCGRID2
2012 Executing Data-Intensive Workloads in a Cloud
abstract
The promise of "infinite" resources given by the cloud computing paradigm has led to recent interest in exploiting clouds for large-scale data-intensive computing. Given this supposedly infinite resource set, we need a management function that regulates application workload on these resources. This doctoral research focuses on two aspects of workload management, namely scheduling and provisioning. We propose a novel framework for workload execution and resource provisioning, and associated models, algorithms, and protocols.
Rizwan Mian, Patrick Martin 0001
CCGRID2
2012 Cloud Service Negotiation: Concession vs. Tradeoff Approaches
abstract
For Cloud services, their non-functional properties like availability, reliability and security are important differentiators. However, service consumers and service providers may conflict over non-functional properties. In fact, the conflicts can be resolved via automated negotiation, which is considered as the most flexible approach to procure products and services. In this paper, we propose tradeoff approaches for Cloud service negotiation, and compare them with concession ones. As opposed to concession ones, tradeoff approaches do not reduce one's utility, but still can create a proposal attractive to its opponent. Indeed, simulation results show that tradeoff approaches outperform concession ones in terms of both individual utility and social benefit. However, simulation results also demonstrate that tradeoff approaches under perform concession ones in terms of success rate.
Xianrong Zheng, Patrick Martin 0001, M. Kathryn Brohman
CCGRID2
2011 Autonomic management of elastic services in the cloud
abstract
Cloud computing, with its support for elastic resources that are available on an on-demand, pay-as-you-go basis, is an attractive platform for hosting Web-based services that have variable demand, yet consistent performance requirements. Effective service management is mandatory in order for services running in the cloud, which we call elastic services, to be cost-effective. In this paper we describe a management framework to facilitate elasticity of resource consumption by services in the cloud. We extend our framework for services management with the necessary concepts and properties to support elastic services. A prototype implementation is described.
Patrick Martin 0001, Wendy Powley, José Luis Vázquez-Poletti
ISCC1
2011 Effective Web service discovery in mobile environments
abstract
Recent advancements in the design of mobile devices and wireless technologies have produced a successful coupling of mobile devices and Web services, where mobile devices can be a service provider or a consumer. However, finding relevant Web services that match requests remain a major hindrance to its booming. The challenges facing Web service discovery are further magnified by the stringent constraints of mobile devices, and the inherit complexity of wireless heterogeneous networks. While significant research has focused on service discovery protocols in isolation, they mostly lack a holistic capacity to address the different limitations collectively. We introduce a novel discovery framework that addresses all aspects of mobile Web service discovery, yet does not jeopardize the efficiency requirement for this discovery; especially as an application run in resource- constrained environments.
Khalid Elgazzar, Hossam S. Hassanein, Patrick Martin 0001
LCN3
2011 Enhancing identity trust in cryptographic key management systems for dynamic environments
abstract
Abstract Cryptographic key management (CKM) schemes can be used to support identity management (IM) systems where linking users securely to data objects is important. CKM schemes enforce data security by encrypting data granting access only to authorized users and security compromises are prevented by updating any keys that are held by users from whom access rights have been revoked. Handling key updates efficiently and providing security against collusion attacks is challenging in dynamic environments like the Internet where manual Security management increases the likelihood of delayed responses. Delay increases the system's vulnerability to security attacks and the potential of the system's violating its service level agreements. Adaptive CKM has emerged as a possibility of addressing this problem but needs to be designed in a way that justifies the cost/benefit tradeoff. In this paper, we show that the key update and collusion avoidance problems are NP‐complete and need heuristic algorithms to prevent performance degradations in comparison to standard CKM schemes. As an example of the benefits of a good heuristic, we present a collusion detection and resolution algorithm whose running time is polynomial in the number of keys. The algorithm operates by mapping the generated key set onto a key graph whose independent set is computed. In the key graph, the vertices represent the keys and the edges the probability that their endpoints can be combined to provoke a collusion attack. Collusion possibilities are resolved by applying a heuristic that resets the probability to zero. The performance of our algorithm is analyzed in comparison to the Akl and Taylor scheme that is secure against collusion attack, and the experimental results indicate that collusion prevention can be done dynamically without affecting performance. Copyright © 2010 John Wiley & Sons, Ltd.
Anne V. D. M. Kayem, Patrick Martin 0001, Selim G. Akl
Secur. Commun. Networks2
2011 Classification Using Streaming Random Forests
abstract
We consider the problem of data stream classification, where the data arrive in a conceptually infinite stream, and the opportunity to examine each record is brief. We introduce a stream classification algorithm that is online, running in amortized O(1) time, able to handle intermittent arrival of labeled records, and able to adjust its parameters to respond to changing class boundaries (“concept drift”) in the data stream. In addition, when blocks of labeled data are short, the algorithm is able to judge internally whether the quality of models updated from them is good enough for deployment on unlabeled records, or whether further labeled records are required. Unlike most proposed stream-classification algorithms, multiple target classes can be handled. Experimental results on real and synthetic data show that accuracy is comparable to a conventional classification algorithm that sees all of the data at once and is able to make multiple passes over it.
Hanady M. Abdulsalam, David B. Skillicorn, Patrick Martin 0001
IEEE Trans. Knowl. Data Eng.3
2011 An Adaptive and Intelligent SLA Negotiation System for Web Services
abstract
The effective use of services to compose business processes in services computing demands that the Quality of Services (QoS) meet consumers' expectations. Automated web-based negotiation of Service Level Agreements (SLA) can help define the QoS requirements of critical service-based processes. We propose a novel trusted Negotiation Broker (NB) framework that performs adaptive and intelligent bilateral bargaining of SLAs between a service provider and a service consumer based on each party's high-level business requirements. We define mathematical models to map business-level requirements to low-level parameters of the decision function, which obscures the complexity of the system from the parties. We also define an algorithm for adapting the decision functions during an ongoing negotiation to comply with an opponent's offers or with updated consumer preferences. The NB uses intelligent agents to conduct the negotiation locally by selecting the most appropriate time-based decision functions. The negotiation outcomes are validated by extensive experimental study for Exponential, Polynomial, and Sigmoid time-based decision functions using simulations on our prototype framework. Results are compared in terms of a total utility value of the negotiating parties to demonstrate the efficiency of our proposed approach.
Farhana Zulkernine, Patrick Martin 0001
IEEE Trans. Serv. Comput.2
2010 Clustering WSDL Documents to Bootstrap the Discovery of Web Services
abstract
The increasing use of the Web for everyday tasks is making Web services an essential part of the Internet customer's daily life. Users query the Internet for a required Web service and get back a set of Web services that may or may not satisfy their request. To get the most relevant Web services that fulfill the user's request, the user has to construct the request using the keywords that best describe the user's objective and match correctly with the Web Service name or location. Clustering Web services based on function similarities would greatly boost the ability of Web services search engines to retrieve the most relevant Web services. This paper proposes a novel technique to mine Web Service Description Language (WSDL) documents and cluster them into functionally similar Web service groups. The application of our approach to real Web services description files has shown good performance for clustering Web services based on function similarity, as a predecessor step to retrieving the relevant Web services for a user request by search engines.
Khalid Elgazzar, Ahmed E. Hassan, Patrick Martin 0001
ICWS3
2010 LTIX: a compact level-based tree to index XML databases
abstract
Indexing XML data is essential for XML query optimization. Most of the existing approaches that combine a labeling scheme with a path index use labeling schemes that reflect the structure of the indexed data. In addition, the labeling rules do not depend on the combined path indexes. By designing a labeling scheme that does not reflect the structure of XML data, since it is available in the accompanied path index; and by aligning the data nodes' labels with the path index nodes' labels, we can support the join process more efficiently. We propose a novel index structure called LTIX (Level-based Tree Index for XML databases). This index structure is based on Level-based Labeling Scheme (LLS) that not only minimizes the number of joins and matches required to evaluate twig queries, if it is used with path indexes, but also facilitates effective query optimization through early pruning of the space search. Experimental tests show the performance benefits of our proposed approach.
Samir Mohammad, Patrick Martin 0001
IDEAS2
2009 A Policy-Based Middleware for Web Services SLA Negotiation
abstract
Negotiation of service level agreements (SLAs) is very important for maintaining quality of service (QoS) of composite Web services-based business processes. The process of negotiation involves specification of negotiation parameters, exchanging offers to conduct the actual negotiation process, and then finally generating the formal SLA if the negotiating parties come to a consensus. We propose a negotiation broker (NB) middleware framework to facilitate automated negotiations of SLAs for Web services in a service oriented architecture (SOA). High level business goals, contexts, preferences, constraints, and values of the negotiation issues are expressed as a policy specification by each of the negotiating parties. The NB maps the policy specifications to low level negotiation strategy models and parameters in order to conduct the negotiation locally as a trusted broker. We present a model and an example of the high level negotiation policy specification. We also present our NB framework including a prototype implementation to illustrate the mapping of the policy to a time-dependent negotiation strategy model.
Farhana Zulkernine, Patrick Martin 0001, Chris Craddock, Kirk Wilson
ICWS2
2009 The Psychic-Skeptic Prediction framework for effective monitoring of DBMS workloads
Said Elnaffar, Patrick Martin 0001
Data Knowl. Eng.2
2009 Towards Autonomic Workload Management in DBMSs
abstract
Workload management is the discipline of effectively managing, controlling, and monitoring work flow across computing systems. It is an increasingly important requirement of database management systems (DBMSs) in view of the trends towards server consolidation and more diverse workloads. Workload management is necessary so the DBMS can be business-objective oriented, can provide efficient differentiated service at fine granularity, and can maintain high utilization of resources with low management costs. The authors see that workload management is shifting from offline planning to online adaptation. In this article, the authors discuss the objectives of workload management in autonomic DBMSs and provide a framework for examining how current workload management mechanisms match up with these objectives. They then use the framework to study several mechanisms from both DBMS products and research efforts. They also propose directions for future work in the area of workload management for autonomic DBMSs.
Baoning Niu, Patrick Martin 0001, Wendy Powley
J. Database Manag.2
2008 Classifying Evolving Data Streams Using Dynamic Streaming Random Forests
Hanady M. Abdulsalam, David B. Skillicorn, Patrick Martin 0001
DEXA3
2008 On replacing cryptographic keys in hierarchical key management systems
abstract
Shared data access maximizes resource utilization on the Internet but raises the issue of data security. We consider a method of shared data access control whereby the data is sub-divided into categories and each encrypted with a unique cryptographic key that is distributed to the user group requir ing access. Key management can be simplified by classifying every user into exactly one of a number of disjoint groups that are partially ordered such that lower level keys are mathematically derivable from higher level keys, but not the reverse. The drawback in this approach is that changes in group membership imply updating both the affected group key and those that are derivable from it. Moreover, the data encrypted with the affected keys must be re-encrypted with the new keys to preserve data security. In the worst case, when the affected group is at the highest level of the hierarchy, the entire hierarchy is affected. This paper presents an algorithm that minimizes the cost of key replacement (rekeying) by associating a timestamp to each key. The timestamp and key are used to compute a verification signature that is used to authenticate users before data access is granted. Thus, whenever group membership changes, instead of rekeying and re-encrypting the affected data, only the timestamp is updated and a new verification signature computed. The new scheme is analyzed using both a time complexity and experimental analysis.
Anne V. D. M. Kayem, Selim G. Akl, Patrick Martin 0001
J. Comput. Secur.3
2008 Is it DSS or OLTP: automatically identifying DBMS workloads
Said Elnaffar, Patrick Martin 0001, Berni Schiefer, Sam Lightstone
J. Intell. Inf. Syst.2
2007 Reputation-Enhanced QoS-based Web Services Discovery
abstract
With an increasing number of Web services providing similar functionalities, quality of service (QoS) is becoming an important criterion for selection of the best available service. Currently the problem is twofold. The Universal Description, Discovery and Integration (UDDI) registries do not have the ability to publish the QoS information, and the authenticity of the advertised QoS information available elsewhere may be questionable. We propose a model of reputation-enhanced QoS- based Web services discovery that combines an augmented UDDI registry to publish the QoS information and a reputation manager to assign reputation scores to the services based on customer feedback of their performance. A discovery agent facilitates QoS-based service discovery using the reputation scores in a service matching, ranking and selection algorithm. The novelty of our model lies in its simplicity and in its coordination of the above mentioned components. We present experiments to evaluate the effectiveness of our approach using a prototype implementation of the model.
Patrick Martin 0001, Wendy Powley, Farhana Zulkernine
ICWS2
2007 Streaming Random Forests
abstract
Many recent applications deal with data streams, conceptually endless sequences of data records, often arriving at high flow rates. Standard data-mining techniques typically assume that records can be accessed multiple times and so do not naturally extend to streaming data. Algorithms for mining streams must be able to extract all necessary information from records with only one, or perhaps a few, passes over the data. We present the streaming random forests algorithm, an online and incremental stream classification algorithm that extends Breiman's random forests algorithm. The streaming random forests algorithm grows multiple decision trees, and classifies unlabeled records based on the plurality of tree votes. We evaluate the classification accuracy of the streaming random forests algorithm on several datasets, and show that its accuracy is comparable to the standard random forest algorithm.
Hanady M. Abdulsalam, David B. Skillicorn, Patrick Martin 0001
IDEAS3
2006 AmbiTalk: Enhancing Wireless Communication Services through the Automatic Adaptation of Mobile Communication
abstract
AmbiTalk is a ubiquitous system based on the session initiation protocol (SIP) and Bluetooth that allows a mobile device to automatically adapt the behavior of its communication services as its user moves from one location to another. The adaptation is policy-based and occurs both pre-call and mid-call. We present the AmbiTalk architecture and discuss how it allows a device to automatically adapt its communication properties to its current environment and other devices. We demonstrate the viability of the AmbiTalk approach with the implementation of a prototype
Eric Karmouch, Patrick Martin 0001, Hossam S. Hassanein
WiMob2
2006 Automated Configuration of Multiple Buffer Pools
abstract
Database management systems (DBMSs) use a main memory area as a buffer to reduce the number of disk accesses performed by a transaction. Some DBMSs divide the buffer area into a number of independent buffer pools and each database object (table or index) is assigned to a specific buffer pool. The tasks of configuring the buffer pools, which define the mapping of database objects to buffer pools and setting a size for each of the buffer pools, are crucial for achieving optimal performance. In this paper we describe an automated approach to multiple buffer pool configuration. Our approach, called BPCluster, analyses the characteristics of a given workload and partitions objects into buffer pools according to their access patterns and inherent characteristics. Similar objects are grouped into the same buffer pool, thus separating those objects that may conflict. A size configuration for the multiple buffer pools is determined using a greedy algorithm that attempts to minimize the cost of a logical read. A set of experimental results validate the approach and show that the configurations suggested by BPCluster outperform naïve configurations and, in most cases, perform as well as configurations suggested by an experienced database administrator.
Patrick Martin 0001, Wendy Powley, Xiaoyi Xu, Wenhu Tian
Comput. J.1
2005 Proactive control of distributed denial of service attacks with source router preferential dropping
abstract
Summary form only given. A distributed denial of service (DDoS) attack is an explicit attempt to interrupt an online service by generating a high volume of malicious traffic. These attacks consume all available network resources, thus rendering legitimate users unable to access the services. Most existing solutions propose to detect and drop attack packets at or near the destination network where the attack packets have already traversed the network and consumed considerable bandwidth. The aggregate traffic at the destination router may consist of hundreds of thousands of flows making it hard for the router to distinguish between legitimate and malicious packets. So, collateral damage is unavoidable. In this paper, we present a source router preferential dropping (SRPD) scheme to detect possible DDoS attacks and defeat them at their sources. SRPD monitors only high-rate outgoing flows at source networks and preferentially drops the packets belonging to these flows when it senses the existence of an attack. A simulation model is constructed and a number of simulation experiments have been conducted to evaluate the performance of the proposed scheme. Simulation results show that SRPD effectively controls DDoS attacks at their sources and reduces collateral damage to a minimum level.
Yinghong Fan, Hossam S. Hassanein, Patrick Martin 0001
AICCSA3
2005 Autonomic buffer pool configuration in PostgreSQL
abstract
As database management systems (DBMSs) continue to expand into new application areas, the complexity of the systems and the diversity of database workloads are increasing. Managing the performance of DBMSs via manual adjustment of resource allocations in this new environment has become impractical. Autonomic DBMSs shift the responsibility for performance management onto the systems themselves. This paper serves as a proof of concept, illustrating how autonomic principles can be applied to a DBMS to provide automatic sizing of buffer pools, a key resource in a DBMS. We describe an implementation of our autonomic system in PostgreSQL, an open source database management system, and provide a set of experiments that verify our approach.
Wendy Powley, Patrick Martin 0001, Nailah Ogeer, Wenhu Tian
SMC2
2005 Experimental Study of a Self-Tuning Algorithm for DBMS Buffer Pools
abstract
The tasks of configuring and tuning large database management systems (DBMSs) have always been both complex and time-consuming. They require knowledge of the characteristics of the system, the data, and the workload, and of the interrelationships between them. The increasing diversity of the data and the workloads handled by today’s systems is making manual tuning by database administrators almost impossible. Self-tuning DBMSs, which dynamically reallocate resources in response to changes in their workload in order to maintain predefined levels of performance, are one approach to handling the tuning problem. In this paper, we apply self-tuning technology to managing the buffer pools, which are a key resource in a DBMS. Tuning the size of the buffer pools to a workload is crucial to achieving good performance. We describe a Buffer Pool Tuning Wizard that can be used by database administrators to determine effective buffer pool sizes. The wizard is based on a self-tuning algorithm called the Dynamic Reconfiguration algorithm (DRF), which uses the principle of goal-oriented resource management. It is an iterative algorithm that uses greedy heuristics to find a reallocation that benefits a target transaction class. We define and motivate the cost estimate equations used in the algorithm. We present the results of a set of experiments to investigate the performance of the algorithm.
Patrick Martin 0001, Wendy Powley
J. Database Manag.1
2004 Developing a characterization of business intelligence workloads for sizing new database systems
abstract
{tedjw, martin, skill} @ cs.queensu.ca Computer system sizing involves estimating the amount of hardware resources needed to support a new workload not yet deployed in a production environment. In order to determine the type and quantity of resources required, a methodology is required for describing the new workload. In this paper, we discuss the sizing process for database management systems and describe an analysis for characterizing business intelligence (BI) workloads, using the TPC-H benchmark as our workload basis. The characterization yields four general classes of queries, each with different characteristics. Our approach for sizing a BI application’s database tier quantifies a new BI workload in terms of the response time goals and mix of the different query classes obtained from the characterization analysis.
Ted J. Wasserman, Patrick Martin 0001, David B. Skillicorn, Haider Rizvi
DOLAP2
2004 Using Reflection to Introduce Self-Tuning Technology into DBMSs
Patrick Martin 0001, Wendy Powley, Darcy G. Benoit
IDEAS1
2004 Differentiated caching of dynamic content using effective page classification
abstract
As the use of dynamic documents increases, caching dynamic content is becoming an important issue for the usability and scalability of the Web. Dynamic content, which is not retained by current Web caching schemes, is adding significant load to Web servers and network links and hence increasing request response times. This paper proposes a scheme, called eager page dynamic caching (EPDC), to effectively cache dynamic content at proxy servers. The scheme identifies two kinds of dynamic pages, called eager-update pages and lazy-update pages, and uses different strategies to deal with each type. For eager-update pages, the Web server pushes the newest data to the proxy server after updates to the dynamic page content. For lazy-update pages, proxy servers pull the newest data from the Web server when clients request it. We use delta-encoding to decrease the amount of data transferred from the Web server to the cache server. We describe a set of simulation experiments we conducted to evaluate our scheme. We show that our scheme can achieve higher hit ratios and lower network latencies, under a variety of conditions, than both simple delta-encoding and traditional Web caching with the least recently used (LRU) scheme.
Wenzhong Chen, Patrick Martin 0001, Hossam S. Hassanein
IPCCC2
2004 QoS differentiation in switching-based Web caching
abstract
Differentiated services (DiffServ) [1998] are being adopted for various Internet applications, including Web services. In the Web-caching field, researchers have proposed to realize DiffServ on Web servers, cache servers, and the client. We argue that there are significant advantages of implementing DiffServ on edge routers in a distributed Web caching system. Edge routers can perform request classification, and assign the type of service, hence the per-hop behavior of the classified requests. If the edge router has knowledge of each cache server, then the edge router is able to provide quality of service to different requests by forwarding the requests to the most appropriate cache server. We propose a switching-based differentiated service aching scheme that provides different types of service to three classes of requests, namely streaming class, real-time assured class and best-effort class. A detailed simulation model is described and then used to examine the conditions under which our scheme is able to satisfy the service requirements of the three classes.
Patrick Martin 0001, Hossam S. Hassanein
IPCCC2
2003 Transparent distributed Web caching with minimum expected response time
abstract
Web caching is a standard approach to improving the performance and quality of Web services. The effectiveness of a single cache in this environment, however, is relatively low. Cache hit rates of 40% or lower are typical in the Web. Distributed caching seeks to improve the effectiveness of Web caching by supporting the sharing of data across multiple cache servers. We describe the minimum expected response time (MRT) distributed Web caching scheme. MRT uses a layer 5 switch to redirect cacheable HTTP requests transparently to the cache server with the minimum expected response time. The response time estimate produced is based on information about cache server content, cache server workload, Web server workload and network latency. We present simulation experiments to show that MRT outperforms existing distributed Web caching schemes in terms of average HTTP request response times.
Patrick Martin 0001, Hossam S. Hassanein
IPCCC2
2002 Automatically classifying database workloads
abstract
The type of the workload on a database management system (DBMS) is a key consideration in tuning the system. Allocations for resources such as main memory can be very different depending on whether the workload type is Online Transaction Processing (OLTP) or Decision Support System (DSS). In this paper, we present an approach to automatically identifying a DBMS workload as either OLTP or DSS. We build a classification model based on the most significant workload characteristics that differentiate OLTP from DSS, and then use the model to identify any change in the workload type. We construct a workload classifier from the Browsing and Ordering profiles of the TPC-W benchmark. Experiments with an industry-supplied workload show that our classifier accurately identifies the mix of OLTP and DSS work within an application workload.
Said Elnaffar, Patrick Martin 0001, Randy Horman
CIKM2
2002 Performance comparison of alternative Web caching techniques
abstract
Web caching is a popular technique to improve the performance and scalability of the Web by increasing document availability and enabling download sharing. Distributed cache cooperation, a mechanism for sharing documents between caches, can further improve performance by providing a shared cache to a large user population. Layer 5 switching-based transparent Web caching schemes intercept HTTP requests and redirect requests according to their contents. This technique not only makes the deployment and configuration of the caching system easier, but also improves its performance by redirecting non-cacheable HTTP requests to bypass cache servers. In this paper, we compare the performance of a number of cooperative (ICP and Cache Digest) and transparent (L5 transparent Web caching and LB-L5) Web caching techniques. We conduct a number of simulation experiments under different HTTP request intensities, network link delays and populations of cooperating cache servers. The relative merits of the different schemes are reported.
Hossam S. Hassanein, Zhengang Liang, Patrick Martin 0001
ISCC3
2002 QRTP: A Middleware for Broadband Networks
Patrick Martin 0001, Michel Sim, Zhenjun Zhu, Hussein T. Mouftah
Multim. Tools Appl.1
2001 Transparent Distributed Web Caching
abstract
Layer 5 switching-based transparent Web caching intercepts HTTP requests and redirects requests according to their contents. This technique makes the deployment and configuration of a caching system easier and improves its performance by ensuring that non-cacheable HTTP requests bypass the cache servers. We propose a Load Balancing Layer 5 switching-based (LB-L5) Web caching scheme that uses the Layer 5 switching-based technique to support distributed Web caching. We present simulation results that show that LB-L5 outperforms existing Web caching schemes, namely ICP, Cache Digest, and basic L5 transparent Web caching, in terms of cache server workload balancing and response time. LB-L5 is also shown to be more adaptable to high HTTP request intensity than the other schemes.
Zhengang Liang, Hossam S. Hassanein, Patrick Martin 0001
LCN3
2000 Dynamic Reconfiguration Algorithm: Dynamically Tuning Multiple Buffer Pools
Patrick Martin 0001, Hoi-Ying Li, Keri Romanufa, Wendy Powley
DEXA1
2000 Using Metadata to Query Passive Data Sources
abstract
In the not too distant past, the amount of online data available to general users was relatively small. Most of the online data was maintained in organizations' database management systems and accessible only through the interfaces provided by those systems. The popularity of the Internet, in particular, has meant that there is now an abundance of online data available to users in the form of Web pages and files. This data, however, is maintained in passive data sources, that is sources that do not provide facilities to search or query their data. The data must be queried and examined using applications such as browsers and search engines. In this paper, we explore an approach to querying passive data sources based on the extraction, and subsequent exploitation, of metadata from the data sources. We describe two situations in which this approach has been used, evaluate the approach and draw some general conclusions.
Patrick Martin 0001, Wendy Powley, Andrew Weston, Peter Zion
Int. J. Cooperative Inf. Syst.1
1998 The performance of SQL queries to an X.500 directory system
David Barrowman, Patrick Martin 0001
Comput. Commun.2
1996 A management information repository for distributed applications management
abstract
The management of distributed applications and systems (MANDAS) project addresses problems arising in the management applications. The MANDAS information repository (MIR) provides database support for the management applications and supports their integration into a single management environment. We examine the problem of distributed applications management to extract the requirements for an MIR. Based on the requirements, we present an information model for distributed applications management and outline a prototype MIR developed for the MANDAS project.
Patrick Martin 0001
ICPADS1
1994 Managing Global Information in the CORDS Multidatabase System
Michael A. Bauer 0001, Neil Coburn, Per-Åke Larson, Patrick Martin 0001
CoopIS4
1994 Parallel Hash-Based Join Algorithms for a Shared-Everything
abstract
Analyzes the costs, and describes the implementation, of three hash-based join algorithms for a general purpose shared-memory multiprocessor. The three algorithms considered are the hashed loops, GRACE and hybrid algorithms. We also describe the results of a set of experiments that validate the cost models presented and demonstrate the relative performance of the three algorithms.>
Patrick Martin 0001, Per-Åke Larson, Vinay Deshpande
IEEE Trans. Knowl. Data Eng.1
1993 Querying and Exploring Large Knowledge Bases
Hing-Kai Hung, Patrick Martin 0001, Janice I. Glasgow, Chris Walmsley, Michael A. Jenkins
DEXA2
1992 Supporting Browsing of Large Knowledge Bases
Patrick Martin 0001, Hing-Kai Hung, Chris Walmsley
DEXA1
1991 A Knowledge-Based System for Fault Diagnosis in Real-Time Engineering Applications
Patrick Martin 0001, Janice I. Glasgow, Michel P. Féret, Todd Kelley
DEXA1
1991 Data caching strategies for distributed full text retrieval systems
Patrick Martin 0001, Judy I. Russell
Inf. Syst.1
1990 An Evaluation of Site Selection Algorithms for Distributed Query Processing
abstract
Site selection in distributed query processing is a computationally intractable problem. We report on experiments comparing solutions from four algorithms — branch-and-bound, greedy, local search and simulated annealing. The algorithms are evaluated with respect to the total query costs obtained for a range of queries for both partially-replicated and fully-replicated databases. We demonstrate the cost-effectiveness of sophisticated algorithms for site selection during the optimization of compiled queries in a large, replicated, distributed database system.
Patrick Martin 0001, K. H. Lam, Judy I. Russell
Comput. J.1
1990 A case study of caching strategies for a distributed full text retrieval system
Patrick Martin 0001, Ian A. Macleod, Judy I. Russell, Ken Leese, Brett Foster
Inf. Process. Manag.1
1989 Remote procedure call facility for a PC environment
Patrick Martin 0001, Ian A. Macleod, Brent Nordin
Comput. Commun.1
1987 Strategies for building distributed information retrieval systems
Ian A. Macleod, Patrick Martin 0001, Brent Nordin, John R. Phillips
Inf. Process. Manag.2
1986 A Design of a Distributed Full Text Retrieval System
abstract
This paper describes the design of a distributed information system for full text retrieval. The system is similar in functionality to STAIRS and is being developed on a network of PC's interconnected by PC Network. The implementation is built on a generalisation of the remote procedure call concept. Communications are based upon the recent CCITT X.400 standard. Examples are given of the design strategy for a subset of the STAIRS system.
Patrick Martin 0001, Ian A. Macleod, Brent Nordin
SIGIR1
1981 Abstraction hierarchies in top-down design
Glenn H. MacEwen, Patrick Martin 0001
J. Syst. Softw.2