Bill Blake

dblp:61/1378 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
0since 2021 · last 2006
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Database system architecture and tuning · 33% Information retrieval · 33% Query processing and optimization · 33%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 62% Cloud and datacenter computing · 19% Storage systems · 19%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning
parallel database system
0.112006
Storage solutions II - Analyzing all the data all the time · SC 2006
Query processing and optimization › query optimization
parallel query optimization
0.112006
Storage solutions II - Analyzing all the data all the time · SC 2006
Information retrieval
query processing
0.112006
Storage solutions II - Analyzing all the data all the time · SC 2006
Hardware accelerators and domain-specific architectures › database accelerator
FPGA-based query processing
0.112006
Storage solutions II - Analyzing all the data all the time · SC 2006
Storage systems › computational storage
active disks
0.012006
Storage solutions II - Analyzing all the data all the time · SC 2006
Cloud and datacenter computing › datacenter architecture
storage-compute integration
0.012006
Storage solutions II - Analyzing all the data all the time · SC 2006

Methods — techniques the papers use, named apart from their topics

massively parallel processing · 0.1FPGA acceleration · 0.1
YearPublicationVenuePosition
2006 Storage solutions II - Analyzing all the data all the time
abstract
A growing number of on-line merchants, wireless phone operators, internet search companies and companies mining credit card transactions to discover customer buying pattern insights have reached the multi-Petabyte scale of data in their operations. It has been recently estimated that all credit card transaction information for all card holders in North America for a ten year period amounts to slightly less than 200 Terabytes, and many Internet businesses have amassed Petabytes of clickstreams. All of these companies are pushing the limits of relational database technology to perform "business intelligence" analysis of the hundreds of billions of records created by their business processes to better understand their customers and gain business advantage. A further challenge to this analysis is the need to perform very complex ad hoc queries, often in the span of hours in the case of fraud detection, on multi-terabyte fact files without the ability to perform the indexing and normalization of data needed to speed up relational data bases used in on-line transaction processing.This presentation will explain the work at Netezza Corporation developing massively parallel systems; purpose built for tera-scale database analysis that can deliver dramatically higher rates of I/O bandwidth to a single table--arguably the single most important metric in data warehouses--than large SMP or cluster of SMP systems. In a nutshell, Netezza is developing systems with a high degree of processor and storage integration that is reminiscent of an active disk approach to bringing processing power as physically close as possible to where the database tables reside. This is a significant departure from the current orthodoxy of large scale systems built with highly virtualized network connected storage. Additional performance gains are achieved by performing many of the key algebraic set operations of the relational database, such as the projection of columns and restriction of rows, in FPGA hardware as the data records stream off the disk drives. Next, the important role of the database optimizer and planner in supporting effective parallelization of queries written in the declarative query language ANSI standard SQL will be outlined. Finally, a number of actual results, many showing a 10 to 50 times speedup over comparable systems, achieved in the customer deployment of systems will be demonstrated.
Bill Blake
SC1