Steve Herbert

dblp:47/2556 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 61% Electronic design automation · 30% Cloud and datacenter computing · 9%
Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 78% Data integration and cleaning · 12% Database system architecture and tuning · 10%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
anomaly detection
0.212015
Learning a Hierarchical Monitoring System for Detecting and Diagnosing Service Issues · KDD 2015
Electronic design automation › hardware verification and test
diagnosis
0.212015
Learning a Hierarchical Monitoring System for Detecting and Diagnosing Service Issues · KDD 2015
Distributed systems
fault tolerance
0.212015
Learning a Hierarchical Monitoring System for Detecting and Diagnosing Service Issues · KDD 2015
Query processing and optimization
query execution
0.112008
Optimizing Star Join Queries for Data Warehousing in Microsoft SQL Server · ICDE 2008
Query processing and optimization
query optimization
0.112008
Optimizing Star Join Queries for Data Warehousing in Microsoft SQL Server · ICDE 2008
Software testing
random testing
0.112007
A genetic approach for random testing of database systems · VLDB 2007
Data integration and cleaning
data warehouse
0.012008
Optimizing Star Join Queries for Data Warehousing in Microsoft SQL Server · ICDE 2008
Database system architecture and tuning
database engine testing
0.012007
A genetic approach for random testing of database systems · VLDB 2007

Methods — techniques the papers use, named apart from their topics

machine learning · 0.2hierarchical monitoring · 0.2genetic algorithm · 0.1right-deep hash joins · 0.1nested loops joins · 0.1cost-based optimization · 0.1
YearPublicationVenuePosition
2015 Learning a Hierarchical Monitoring System for Detecting and Diagnosing Service Issues
abstract
We propose a machine learning based framework for building a hierarchical monitoring system to detect and diagnose service issues. We demonstrate its use for building a monitoring system for a distributed data storage and computing service consisting of tens of thousands of machines. Our solution has been deployed in production as an end-to-end system, starting from telemetry data collection from individual machines, to a visualization tool for service operators to examine the detection outputs. Evaluation results are presented on detecting 19 customer impacting issues in the past three months.
Vinod Nair, Ameya Raul, Shwetabh Khanduja, Vikas Bahirwani, Sundararajan Sellamanickam, S. Sathiya Keerthi, Steve Herbert, Sudheer Dhulipalla
KDD7
2008 Optimizing Star Join Queries for Data Warehousing in Microsoft SQL Server
abstract
As mainstream data warehouses are growing into the multi-terabyte range, adequate performance for decision support queries remains challenging for database query processors. Proper choice of query plan is essential in data warehouses where fact tables often store billions of rows. This paper discusses query optimization and execution strategies that Microsoft SQL Server employs for decision support queries in dimensionally modeled relational data warehouses. Our approach is based on pattern matching to detect typical star query patterns. When matching the pattern, the optimizer generates additional query plan alternatives specifically optimized for data warehouse performance. For high selectivity queries, the plans use nested loops joins and seeks. Medium selectivity queries in turn rely on right-deep hash joins with bitmap filters. Bitmap filters perform semi-join reductions to efficiently prune out non-qualifying rows early. Final plan choice is left for cost-based optimization which also compares the data warehouse specific plans against conventional query plans. We conducted an extensive experimental investigation using both synthetic workloads and several customer workloads. As our results show, the new plan shapes and execution strategies yield significant performance improvements across the targeted workloads as compared to earlier versions of Microsoft SQL Server.
César A. Galindo-Legaria, Torsten Grabs, Sreenivas Gukal, Steve Herbert, Aleksandras Surna, Shirley Wang, Peter Zabback, Shin Zhang
ICDE4
2007 A genetic approach for random testing of database systems
Hardik Bati, Leo Giakoumakis, Steve Herbert, Aleksandras Surna
VLDB3