Paul G. Brown

dblp:41/6335 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data stream processing · 36% Distributed and cloud data management · 32% Query processing and optimization · 17%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 61% Storage systems · 39%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management
distributed analytics
0.612022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Data stream processing
streaming analytics
0.612022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Query processing and optimization
analytical query processing
0.212022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Cloud and datacenter computing
cloud platform
0.212022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Data models and query languages
array data management
0.112010
Overview of sciDB: large scale array storage, processing and analysis · SIGMOD Conference 2010
Data models and query languages › multidimensional data model
array data model
0.112010
Overview of sciDB: large scale array storage, processing and analysis · SIGMOD Conference 2010
Query processing and optimization
approximate query processing
0.112006
Techniques for Warehousing of Sample Data · ICDE 2006
Data integration and cleaning
data warehouse
0.112006
Techniques for Warehousing of Sample Data · ICDE 2006
Data stream processing
random sampling
0.112006
Techniques for Warehousing of Sample Data · ICDE 2006
Query processing and optimization
query optimization
0.012004
CORDS: Automatic Generation of Correlation Statistics in DB2 · VLDB 2004
Query processing and optimization
cardinality estimation
0.012004
CORDS: Automatic Generation of Correlation Statistics in DB2 · VLDB 2004

Methods — techniques the papers use, named apart from their topics

distributed computing · 1.1uniform random sampling · 0.1stream sampling · 0.1sample merging · 0.1
YearPublicationVenuePosition
2022 Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels
abstract
Data Analytics is at the core of almost all modern ap-plications ranging from science and finance to healthcare and web applications. The evolution of data analytics over the last decade has been dramatic - new methods, new tools and new platforms - with no slowdown in sight. This rapid evolution has pushed the boundaries of data analytics along several axis including scalability especially with the rise of distributed infrastructures and the Big Data era, and interoperability with diverse data management systems such as relational databases, Hadoop and Spark. However, many analytic application developers struggle with the challenge of production deployment. Recent experience suggests that it is difficult to deliver modern data analytics with the level of reliability, security and manageability that has been a feature of traditional SQL DBMSs. In this tutorial, we discuss the advances and innovations introduced at both the infrastructure and algorithmic levels, directed at making analytic workloads scale, while paying close attention to the kind of quality of service guarantees different technology provide. We start with an overview of the classical centralized analytical techniques, describing the shift towards distributed analytics over non-SQL infrastructures. We contrast such approaches with systems that integrate analytic functionality inside, above or adjacent to SQL engines. We also explore how Cloud platforms' virtualization capabilities make it easier - and cheaper - for end users to apply these new analytic techniques to their data. Finally, we conclude with the learned lessons and a vision for the near future.
Mohammed Al-Kateb, Mohamed Y. Eltabakh, Awny Al-Omari, Paul G. Brown
ICDE4
2016 Addressing the big-earth-data variety challenge with the hierarchical triangular mesh
abstract
We have implemented an updated Hierarchical Triangular Mesh (HTM) as the basis for a unified data model and an indexing scheme for geoscience data to address the variety challenge of Big Earth Data. In the absence of variety, the volume challenge of Big Data is relatively easily addressable with parallel processing. The more important challenge in achieving optimal value with a Big Data solution for Earth Science (ES) data analysis, however, is being able to achieve good scalability with variety. With HTM unifying at least the three popular data models, i.e. Grid, Swath, and Point, used by current ES data products, data preparation time for integrative analysis of diverse datasets can be drastically reduced and better variety scaling can be achieved. HTM is also an indexing scheme, and when applied to all ES datasets, data placement alignment (or co-location) on the shared nothing architecture, which most Big Data systems are based on, is guaranteed and better performance is ensured. With HTM most geospatial set operations become integer interval operations with further performance advantages.
Mike Rilee, Kwo-Sen Kuo, Thomas L. Clune, Amidu Oloso, Paul G. Brown, Hongfeng Yu 0001
IEEE BigData5
2010 Overview of sciDB: large scale array storage, processing and analysis
abstract
SciDB [4, 3] is a new open-source data management system intended primarily for use in application domains that involve very large (petabyte) scale array data; for example, scientific applications such as astronomy, remote sensing and climate modeling, bio-science information management, risk management systems in financial applications, and the analysis of web log data. In this talk we will describe our set of motivating examples and use them to explain the features of SciDB. We then briefly give an overview of the project 'in flight', explaining our novel storage manager, array data model, query language, and extensibility frameworks.
Paul G. Brown
SIGMOD Conference1
2006 Techniques for Warehousing of Sample Data
abstract
We consider the problem of maintaining a warehouse of sampled data that "shadows" a full-scale data warehouse, in order to support quick approximate analytics and metadata discovery. The full-scale warehouse comprises many "data sets," where a data set is a bag of values; the data sets can vary enormously in size. The values constituting a data set can arrive in batch or stream form. We provide and compare several new algorithms for independent and parallel uniform random sampling of data-set partitions, where the partitions are created by dividing the batch or splitting the stream. We also provide novel methods for merging samples to create a uniform sample from an arbitrary union of data-set partitions. Our sampling/merge methods are the first to simultaneously support statistical uniformity, a priori bounds on the sample footprint, and concise sample storage. As partitions are rolled in and out of the warehouse, the corresponding samples are rolled in and out of the sample warehouse. In this manner our sampling methods approximate the behavior of more sophisticated stream-sampling methods, while also supporting parallel processing. Experiments indicate that our methods are efficient and scalable, and provide guidance for their application.
Paul G. Brown, Peter J. Haas
ICDE1
2004 CORDS: Automatic Generation of Correlation Statistics in DB2
Ihab F. Ilyas, Volker Markl, Peter J. Haas, Paul G. Brown, Ashraf Aboulnaga
VLDB4