Paul Brown

dblp:04/5664 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Data models and query languages · 36% Database system architecture and tuning · 33% Data integration and cleaning · 23%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100% Environmental and earth informatics · 0%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 85% Performance modeling and evaluation · 15%

Topics — the 12 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data models and query languages › multidimensional database
array DBMS
0.312018
ArrayBridge: Interweaving Declarative Array Processing in SciDB with Imperative HDF5-Based Programs · ICDE 2018
Bioinformatics and computational biology
gene expression analysis
0.212013
MEME-LaB: motif analysis in clusters · Bioinform. 2013
Bioinformatics and computational biology › sequence analysis › motif discovery
motif analysis
0.212013
MEME-LaB: motif analysis in clusters · Bioinform. 2013
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
0.212013
MEME-LaB: motif analysis in clusters · Bioinform. 2013
High-performance computing
scientific computing systems
0.112018
ArrayBridge: Interweaving Declarative Array Processing in SciDB with Imperative HDF5-Based Programs · ICDE 2018
Data integration and cleaning › web data integration
data mashup
0.112007
DAMIA - A Data Mashup Fabric for Intranet Applications · VLDB 2007
Data integration and cleaning
dependency discovery
0.122006
CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies · SIGMOD Conference 2004
GORDIAN: Efficient and Scalable Discovery of Composite Keys · VLDB 2006
Query processing and optimization
selectivity estimation
0.012004
CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies · SIGMOD Conference 2004
Data integration and cleaning
data profiling
0.012003
BHUNT: Automatic Discovery of Fuzzy Algebraic Constraints in Relational Data · VLDB 2003
Data models and query languages
object-relational database
0.011999
Implementing the Spirit of SQL-99 · SIGMOD Conference 1999
Performance modeling and evaluation
benchmarking
0.011997
The BUCKY Object-Relational Benchmark (Experience Paper) · SIGMOD Conference 1997
Database system architecture and tuning
scientific data management
0.011995
BigSur: A System For the Management of Earth Science Data · VLDB 1995

Methods — techniques the papers use, named apart from their topics

array view mechanism · 0.7web interface · 0.2ab initio motif finding · 0.2sampling · 0.0chi-squared analysis · 0.0
YearPublicationVenuePosition
2018 ArrayBridge: Interweaving Declarative Array Processing in SciDB with Imperative HDF5-Based Programs
abstract
Scientists are increasingly turning to datacenter-scale computers to analyze massive arrays. Despite decades of database research that extols the virtues of declarative query processing, scientists still write, debug and parallelize imperative HPC programs even for the most mundane queries. This impedance mismatch is due to the cumbersome and costly data format conversions that are needed to use scientific data management tools, such as SciDB, in an HPC setting. Our goal is to make declarative array manipulations from SciDB interoperable with imperative, file-centric analyses from HDF5-based programs. This paper describes ArrayBridge, a bi-directional array view mechanism for the HDF5 file format, that allows scientists to use SciDB, TensorFlow and HDF5-based analysis code in the same file-centric pipeline without converting between file formats. In addition to fast querying over HDF5 array objects, ArrayBridge produces arrays in the HDF5 file format as easily as it can read from it. ArrayBridge also supports time travel queries from imperative codes through the unmodified HDF5 API, and automatically deduplicates between versions for space efficiency. Our performance evaluation in a large scientific computing facility shows that ArrayBridge exhibits statistically indistinguishable performance and I/O scalability to the native SciDB storage engine and is 3× faster than TileDB.
Haoyuan Xing, Sofoklis Floratos, Spyros Blanas, Surendra Byna, Prabhat, Kesheng Wu, Paul Brown
ICDE7
2018 Bringing numerous methods for expression and promoter analysis to a public cloud computing service
abstract
Summary: Every year, a large number of novel algorithms are introduced to the scientific community for a myriad of applications, but using these across different research groups is often troublesome, due to suboptimal implementations and specific dependency requirements. This does not have to be the case, as public cloud computing services can easily house tractable implementations within self-contained dependency environments, making the methods easily accessible to a wider public. We have taken 14 popular methods, the majority related to expression data or promoter analysis, developed these up to a good implementation standard and housed the tools in isolated Docker containers which we integrated into the CyVerse Discovery Environment, making these easily usable for a wide community as part of the CyVerse UK project. Availability and implementation: The integrated apps can be found at http://www.cyverse.org/discovery-environment, while the raw code is available at https://github.com/cyversewarwick and the corresponding Docker images are housed at https://hub.docker.com/r/cyversewarwick/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Krzysztof Polanski, Bo Gao 0001, Sam A. Mason, Paul Brown, Sascha Ott, Katherine J. Denby, David L. Wild
Bioinform.4
2016 Implementing connected component labeling as a user defined operator for SciDB
abstract
We have implemented a flexible User Defined Operator (UDO) for labeling connected components of a binary mask expressed as an array in SciDB, a parallel distributed database management system based on the array data model. This UDO is able to process very large multidimensional arrays by exploiting SciDB's memory management mechanism that efficiently manipulates arrays whose memory requirements far exceed available physical memory. The UDO takes as primary inputs a binary mask array and a binary stencil array that specifies the connectivity of a given cell to its neighbors. The UDO returns an array of the same shape as the input mask array with each foreground cell containing the label of the component it belongs to. By default, dimensions are treated as non-periodic, but the UDO also accepts optional input parameters to specify periodicity in any of the array dimensions. The UDO requires four stages to completely label connected components. In the first stage, labels are computed for each subarray or chunk of the mask array in parallel across SciDB instances using the weighted quick union (WQU) with half-path compression algorithm. In the second stage, labels around chunk boundaries from the first stage are stored in a temporary SciDB array that is then replicated across all SciDB instances. Equivalences are resolved by again applying the WQU algorithm to these boundary labels. In the third stage, relabeling is done for each chunk using the resolved equivalences. In the fourth stage, the resolved labels, which so far are “flattened” coordinates of the original binary mask array, are renamed with sequential integers for legibility. The UDO is demonstrated on a 3-D mask of 0(10n) elements, with 0(108) foreground cells and o(106) connected components. The operator completes in 19 minutes using 84 SciDB instances.
Amidu Oloso, Kwo-Sen Kuo, Thomas L. Clune, Paul Brown, Alex Poliakov, Hongfeng Yu 0001
IEEE BigData4
2013 MEME-LaB: motif analysis in clusters
abstract
SUMMARY: Genome-wide expression analysis can result in large numbers of clusters of co-expressed genes. Although there are tools for ab initio discovery of transcription factor-binding sites, most do not provide a quick and easy way to study large numbers of clusters. To address this, we introduce a web tool called MEME-LaB. The tool wraps MEME (an ab initio motif finder), providing an interface for users to input multiple gene clusters, retrieve promoter sequences, run motif finding and then easily browse and condense the results, facilitating better interpretation of the results from large-scale datasets. AVAILABILITY: MEME-LaB is freely accessible at: http://wsbc.warwick.ac.uk/wsbcToolsWebpage/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Paul Brown, Laura Baxter, Richard Hickman, Jim Beynon, Jonathan D. Moore, Sascha Ott
Bioinform.1
2011 The Architecture of SciDB
Michael Stonebraker, Paul Brown, Alex Poliakov, Suchi Raman
SSDBM2
2007 DAMIA - A Data Mashup Fabric for Intranet Applications
Mehmet Altinel, Paul Brown, Susan Cline, Rajesh Kartha, Eric Louie, Volker Markl, Louis Mau, Yip-Hing Ng, David E. Simmen
VLDB2
2006 GORDIAN: Efficient and Scalable Discovery of Composite Keys
Yannis Sismanis, Paul Brown, Peter J. Haas, Berthold Reinwald
VLDB2
2006 Assimilation patterns in the use of electronic procurement innovations: A cluster analysis
Arun Rai, Xinlin Tang, Paul Brown, Mark Keil 0001
Inf. Manag.3
2005 Creating histories
abstract
This panel invites four speakers to discuss the history of creativity and cognition. Phil Husbands talks about the pioneering group that played an important role in the emergence of cybernetics in the UK - The Ratio Club; Margaret Boden describes the history of creativity research in AI; Catherine Mason presents her research into the role that institutions played in the development of the computer arts in the UK and; Alan Sutcliffe describes his experiences as a co-founder of the influential Computer Arts Society.
Paul Brown, Phil Husbands, Margaret A. Boden, Catherine Mason, Alan Sutcliffe
Creativity & Cognition1
2004 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies
abstract
The rich dependency structure found in the columns of real-world relational databases can be exploited to great advantage, but can also cause query optimizers---which usually assume that columns are statistically independent---to underestimate the selectivities of conjunctive predicates by orders of magnitude. We introduce CORDS, an efficient and scalable tool for automatic discovery of correlations and soft functional dependencies between columns. CORDS searches for column pairs that might have interesting and useful dependency relations by systematically enumerating candidate pairs and simultaneously pruning unpromising candidates using a flexible set of heuristics. A robust chi-squared analysis is applied to a sample of column values in order to identify correlations, and the number of distinct values in the sampled columns is analyzed to detect soft functional dependencies. CORDS can be used as a data mining tool, producing dependency graphs that are of intrinsic interest. We focus primarily on the use of CORDS in query optimization. Specifically, CORDS recommends groups of columns on which to maintain certain simple joint statistics. These "column-group" statistics are then used by the optimizer to avoid naive selectivity estimates based on inappropriate independence assumptions. This approach, because of its simplicity and judicious use of sampling, is relatively easy to implement in existing commercial systems, has very low overhead, and scales well to the large numbers of columns and large table sizes found in real-world databases. Experiments with a prototype implementation show that the use of CORDS in query optimization can speed up query execution times by an order of magnitude. CORDS can be used in tandem with query feedback systems such as the LEO learning optimizer, leveraging the infrastructure of such systems to correct bad selectivity estimates and ameliorating the poor performance of feedback systems during slow learning phases.
Ihab F. Ilyas, Volker Markl, Peter J. Haas, Paul Brown, Ashraf Aboulnaga
SIGMOD Conference4
2003 BHUNT: Automatic Discovery of Fuzzy Algebraic Constraints in Relational Data
Paul Brown, Peter J. Haas
VLDB1
1999 Implementing the Spirit of SQL-99
abstract
This paper describes the current INFORMIX IDS/UD release (9.2 or Centaur) and compares and contrasts its functionality with the features of the SQL-99 language standard. INFORMIX and Illustra have been shipping DBMSs implementing the spirit of the SQL-99 standard for five years. In this paper, we review our experience working with ORDBMS technology, and argue that while SQL-99 is a huge improvement over SQL-92, substantial further work is necessary to make object-relational DBMSs truly useful. Specifically, we describe several interesting pieces of functionality unique to IDS/UD, and several dilemmas our customers have encountered that the standard does not address.
Paul Brown
SIGMOD Conference1
1997 The BUCKY Object-Relational Benchmark (Experience Paper)
abstract
According to various trade journals and corporate marketing machines, we are now on the verge of a revolution—the object-relational database revolution. Since we believe that no one should face a revolution without appropriate armaments, this paper presents BUCKY, a new benchmark for object-relational database systems. BUCKY is a query-oriented benchmark that tests many of the key features offered by object-relational systems, including row types and inheritance, references and path expressions, sets of atomic values and of references, methods and late binding, and user-defined abstract data types and their methods. To test the maturity of object-relational technology relative to relational technology, we provide both an object-relational version of BUCKY and a relational equivalent thereof (i.e., a relational BUCKY simulation). Finally, we briefly discuss the initial performance results and lessons that resulted from applying BUCKY to one of the early object-relational database system products.
Michael J. Carey 0001, David J. DeWitt, Jeffrey F. Naughton, Mohammad Asgarian, Paul Brown, Johannes Gehrke, Dhaval Shah
SIGMOD Conference5
1995 BigSur: A System For the Management of Earth Science Data
Paul Brown, Michael Stonebraker
VLDB1