Paul Greenfield

dblp:78/5798 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
0since 2021 · last 2014
0000-0003-4028-9243ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis
sequencing error correction
0.212014
Blue: correcting sequencing errors using consensus and context · Bioinform. 2014
Bioinformatics and computational biology › sequence analysis › sequencing data processing
sequencing quality control
0.212014
Blue: correcting sequencing errors using consensus and context · Bioinform. 2014
Services computing and microservices
web services
0.112005
Consistency for Web Services Applications · VLDB 2005
Distributed systems
distributed coordination
0.012005
Consistency for Web Services Applications · VLDB 2005

Methods — techniques the papers use, named apart from their topics

k-mer consensus · 0.2context-based correction · 0.2
YearPublicationVenuePosition
2014 Blue: correcting sequencing errors using consensus and context
abstract
MOTIVATION: Bioinformatics tools, such as assemblers and aligners, are expected to produce more accurate results when given better quality sequence data as their starting point. This expectation has led to the development of stand-alone tools whose sole purpose is to detect and remove sequencing errors. A good error-correcting tool would be a transparent component in a bioinformatics pipeline, simply taking sequence data in any of the standard formats and producing a higher quality version of the same data containing far fewer errors. It should not only be able to correct all of the types of errors found in real sequence data (substitutions, insertions, deletions and uncalled bases), but it has to be both fast enough and scalable enough to be usable on the large datasets being produced by current sequencing technologies, and work on data derived from both haploid and diploid organisms. RESULTS: This article presents Blue, an error-correction algorithm based on k-mer consensus and context. Blue can correct substitution, deletion and insertion errors, as well as uncalled bases. It accepts both FASTQ and FASTA formats, and corrects quality scores for corrected bases. Blue also maintains the pairing of reads, both within a file and between pairs of files, making it compatible with downstream tools that depend on read pairing. Blue is memory efficient, scalable and faster than other published tools, and usable on large sequencing datasets. On the tests undertaken, Blue also proved to be generally more accurate than other published algorithms, resulting in more accurately aligned reads and the assembly of longer contigs containing fewer errors. One significant feature of Blue is that its k-mer consensus table does not have to be derived from the set of reads being corrected. This decoupling makes it possible to correct one dataset, such as small set of 454 mate-pair reads, with the consensus derived from another dataset, such as Illumina reads derived from the same DNA sample. Such cross-correction can greatly improve the quality of small (and expensive) sets of long reads, leading to even better assemblies and higher quality finished genomes. AVAILABILITY AND IMPLEMENTATION: The code for Blue and its related tools are available from http://www.bioinformatics.csiro.au/Blue. These programs are written in C# and run natively under Windows and under Mono on Linux.
Paul Greenfield, Konsta Duesing, Alexie Papanicolaou, Denis C. Bauer
Bioinform.1
2013 A performance evaluation of distributed database architectures
abstract
SUMMARY The globally integrated contemporary business environment has prompted new challenges to database architectures in order to enable organizations to improve database applications performance, scalability, reliability and data privacy in adapting to the evolving nature of business. Although a number of distributed database architectures are available for choice, there is a lack of an in‐depth understanding of the performance characteristics of these database architectures in a comparison way. In this paper, we report a performance study of three typical (centralized, partitioned and replicated) database architectures. We used the TPC‐C as the evaluation benchmark to simulate a contemporary business environment, and a commercially available database management system that supports the three architectures. We compared the performance of the partitioned and replicated architectures against the centralized database, which results in some interesting observations and practical experience. The findings and the practice presented in this paper provide useful information and experience for the enterprise architects and database administrators in determining the appropriate database architecture in moving from centralized to distributed environments. Copyright © 2012 John Wiley & Sons, Ltd.
Shiping Chen 0001, Alex Ng, Paul Greenfield
Concurr. Comput. Pract. Exp.3
2013 Answering biological questions by querying k-mer databases
abstract
SUMMARY This paper describes a k‐mer approach to analysing DNA data and quickly answering certain types of ad hoc biological questions. These k‐mers (short DNA strings) are stored in a conventional relational database and indexed to support efficient exact match operations. We show that k‐mers around 20–25 bases long have interesting and useful uniqueness properties that can be used to compute a ‘relatedness’ metric and also allow k‐mers to be used as ‘unique enough’ tags to identify organisms and genes. This relatedness metric is used in SQL queries that can directly answer questions such as how two related species differ, and what genes are unique to an organism. The k‐mer tags have proven useful in applications, largely metagenomic ones that can quickly process large volumes of sequencing data to say something about what organisms and genes might be present in an environmental sample. All of this work is based on simple and fast exact matches of k‐mer strings using a database, rather than conventional alignment based on inexact matches of much longer strings. These k‐mer tools provide ways of rapidly exploring large genome spaces and handling large volumes of sequence data, and complement rather than replace existing alignment and assembly tools. Copyright © 2012 John Wiley & Sons, Ltd.
Paul Greenfield, Uwe Röhm
Concurr. Comput. Pract. Exp.1
2007 Isolation Support for Service-based Applications: A Position Paper
Paul Greenfield, Alan D. Fekete, Julian Jang, Dean Kuo, Surya Nepal
CIDR1
2007 Delivering Promises for Web Services Applications
abstract
Among the problems facing designers of complex multi-participant Web services-based applications is dealing with the consequences of the lack of suitable isolation mechanisms. This deficiency means that concurrent applications can interfere with each other, resulting in race conditions and lost updates. This paper considers a proposed solution to this problem based on 'promises' and shows that this model can be implemented in practice. We consider implementation issues that need to be handled in promise-based systems and discuss a proof of concept prototype that supports promise-based isolation without requiring changes to existing applications and resources.
Julian Jang, Alan D. Fekete, Paul Greenfield
ICWS3
2006 An Event-Driven Workflow Engine for Service-based Business Systems
abstract
This paper discusses a novel implementation of a workflow engine that supports service-based applications. The applications are defined according to the GAT model, which is an event-based programming model using conditional guards to determine when both normal and exception-handling activities are to be executed. We propose implementation techniques for key features of GAT. These include implementing control flow based on the evaluation of guards, the management and distribution of events, and enforcing atomicity across the evaluation of guards and the execution of the corresponding activities. We have built an engine following this approach which uses available technologies to support translating GAT models into executable applications
Julian Jang, Alan D. Fekete, Paul Greenfield, Surya Nepal
EDOC3
2006 Expressing and Reasoning about Service Contracts in Service-Oriented Computing
abstract
The Web services and service-oriented architectures (SOA) vision by Helland, P. (2005) is about building large-scale distributed applications by composing coarse-grained autonomous services in a flexible architecture that can adapt to changing business requirements. These services interact by exchanging one-way messages through standardized message processing and transport protocols. This vision is being driven by economic imperatives for integration and automation across administrative and organizational boundaries. This paper presents a concise yet expressive model for service contracts to describe messaging behavior. The idea is simple: we use Boolean conditions to specify when a message can be sent and received, where the conditions refer only to other messages in the service contract - that is, conditions only refer to a service's externalized messaging state and not to internal state
Dean Kuo, Alan D. Fekete, Paul Greenfield, Surya Nepal, John Zic, Savas Parastatidis, Jim Webber
ICWS3
2005 Consistency for Web Services Applications
Paul Greenfield, Dean Kuo, Surya Nepal, Alan D. Fekete
VLDB1
2003 Just What Could Possibly Go Wrong In B2B Integration?
abstract
One important trend in enterprise-scale IT has been the increasing use of business-business integration (B2Bi) technologies to automate business processes that cross organizational boundaries, such as the interactions between partner companies along a supply chain. It is relatively easy to describe a pattern of interaction, or choreography, in the case where everything proceeds smoothly. However, the abnormal cases, such as where a process fails or a message is lost, are much more complicated, and risk introducing data and process inconsistencies into computer-based systems. Current B2Bi technologies do not supply an infrastructure that can provide reliability without considerable sophistication from the architects and developers. As a first step towards guiding architects to the design of B2Bi systems that maintain consistency despite failures, this paper describes a variety of types of failure that can arise in practice, based on a realistic e-procurement scenario. We describe these failures in terms of the different types of state that naturally occur within the distributed system. Understanding the types of failure that need to be handled, or prevented, is essential to an architect or developer who must design and write handlers for all the exceptions that can occur in their workflows.
Dean Kuo, Alan D. Fekete, Paul Greenfield, Julian Jang, Doug Palmer
COMPSAC3
2003 Compensation is Not Enough
abstract
An important problem in designing infrastructure to support business-to-business integration (B2Bi) is how to cancel a long-running interaction (either because the user has changed their mind, or in response to an unrecoverable failure). We review the fault-handling and compensation mechanism that is now used in most workflow products and business process modeling standards. We then use an e-procurement case-study to extract a set of requirements for an effective cancellation mechanism, and we show that the standard approach using fault-handling, and compensation transactions is not adequate to meet these requirements.
Paul Greenfield, Alan D. Fekete, Julian Jang, Dean Kuo
EDOC1
2003 Expressiveness of Workflow Description Languages
Julian Jang, Alan D. Fekete, Paul Greenfield, Dean Kuo
ICWS3
2003 Role and Expression of Consent in Web Services
Christine M. O'Keefe, Paul Greenfield
ICWS2