Dina Bitton

dblp:03/2311 · DBLP profile ↗
← Back
17ranked-venue papers
11as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 15 · 11 first-authorSystems, architecture and hardware · 2Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
11 papers
Query processing and optimization · 50% Data integration and cleaning · 20% Database system architecture and tuning · 17%
Computer architecture, parallel and distributed computing, and storage systems
8 papers
Parallel and multicore computing · 92% Storage systems · 4% Performance modeling and evaluation · 3%

Topics — the 29 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query execution
pipelining
0.212016
Elastic Pipelining in an In-Memory Database Cluster · SIGMOD Conference 2016
Parallel and multicore computing › task scheduling
dynamic scheduling
0.212016
Elastic Pipelining in an In-Memory Database Cluster · SIGMOD Conference 2016
Data integration and cleaning
enterprise information integration
0.112005
Enterprise information integration: successes, challenges and controversies · SIGMOD Conference 2005
Data integration and cleaning
data fusion
0.012006
Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality? · VLDB 2006
Data integration and cleaning
data warehouse
0.012005
Enterprise information integration: successes, challenges and controversies · SIGMOD Conference 2005
Data integration and cleaning › data integration system
virtual data integration
0.012005
Enterprise information integration: successes, challenges and controversies · SIGMOD Conference 2005
Database system architecture and tuning
main-memory database
0.021987
Performance of Complex Queries in Main Memory Database Systems · ICDE 1987
The Effect of Large Main memory on Database Systems · SIGMOD Conference 1986
Database system architecture and tuning › database design
database design tools
0.011989
A Feasibility and Performance Study of Dependency Inference · ICDE 1989
Database theory
dependency theory
0.011989
A Feasibility and Performance Study of Dependency Inference · ICDE 1989
Data integration and cleaning › dependency discovery
functional dependency discovery
0.011989
A Feasibility and Performance Study of Dependency Inference · ICDE 1989
Parallel and multicore computing
parallel algorithms
0.011988
Sorting Large Files on a Backend Multiprocessor · IEEE Trans. Computers 1988
Parallel and multicore computing › parallel algorithms › sorting
parallel sorting
0.011988
Sorting Large Files on a Backend Multiprocessor · IEEE Trans. Computers 1988
Storage systems
storage reliability
0.011988
Disk Shadowing · VLDB 1988
Query processing and optimization
complex query processing
0.011987
Performance of Complex Queries in Main Memory Database Systems · ICDE 1987
Query processing and optimization › query execution
in-memory query processing
0.011987
Performance of Complex Queries in Main Memory Database Systems · ICDE 1987
Storage systems › storage architecture
mass storage system
0.011987
Technology Trends in Mass-Storage Systems · SIGMOD Conference 1987
Query processing and optimization
cardinality estimation
0.011986
Estimating Block Accessses when Attributes are Correlated · VLDB 1986
Query processing and optimization › query execution › relational operators
duplicate elimination
0.011983
Duplicate Record Elimination in Large Data Files · ACM Trans. Database Syst. 1983
Query processing and optimization
parallel query processing
0.011983
Parallel Algorithms for the Execution of Relational Database Operations · ACM Trans. Database Syst. 1983
Performance modeling and evaluation
benchmarking
0.011983
Benchmarking Database Systems A Systematic Approach · VLDB 1983
Performance modeling and evaluation › benchmarking
database system benchmarking
0.011983
Benchmarking Database Systems A Systematic Approach · VLDB 1983
Indexing and storage engines › storage management
storage architecture
0.011988
Disk Shadowing · VLDB 1988
Distributed systems › distributed system architecture
distributed operating systems
0.011988
Sorting Large Files on a Backend Multiprocessor · IEEE Trans. Computers 1988
Query processing and optimization › aggregate query processing
join-aggregate query
0.011987
Performance of Complex Queries in Main Memory Database Systems · ICDE 1987
Query processing and optimization
join processing
0.011987
Performance of Complex Queries in Main Memory Database Systems · ICDE 1987
Performance modeling and evaluation › workload characterization
database workload characterization
0.011986
Estimating Block Accessses when Attributes are Correlated · VLDB 1986
Memory systems
main memory
0.011986
The Effect of Large Main memory on Database Systems · SIGMOD Conference 1986
Query processing and optimization › cost estimation
query optimizer cost estimation
0.011983
Duplicate Record Elimination in Large Data Files · ACM Trans. Database Syst. 1983
Parallel and multicore computing
parallel query processing
0.011983
Parallel Algorithms for the Execution of Relational Database Operations · ACM Trans. Database Syst. 1983

Methods — techniques the papers use, named apart from their topics

elastic iterator model · 0.5dynamic scheduler · 0.5experiment · 0.0cost analysis · 0.0complexity analysis · 0.0performance modeling · 0.0parallel sort-merge · 0.0benchmarking · 0.0combinatorial cost model · 0.0
YearPublicationVenuePosition
2016 Elastic Pipelining in an In-Memory Database Cluster
abstract
An in-memory database cluster consists of multiple interconnected nodes with a large capacity of RAM and modern multi-core CPUs. As a conventional query processing strategy, pipelining remains a promising solution for in-memory parallel database systems, as it avoids expensive intermediate result materialization and parallelizes the data processing among nodes. However, to fully unleash the power of pipelining in a cluster with multi-core nodes, it is crucial for the query optimizer to generate good query plans with appropriate intra-node parallelism, in order to maximize CPU and network bandwidth utilization. A suboptimal plan, on the contrary, causes load imbalance in the pipelines and consequently degrades the query performance. Parallelism assignment optimization at compile time is nearly impossible, as the workload in each node is affected by numerous factors and is highly dynamic during query evaluation. To tackle this problem, we propose elastic pipelining, which makes it possible to optimize intra-node parallelism assignments in the pipelines based on the actual workload at runtime. It is achieved with the adoption of new elastic iterator model and a fully optimized dynamic scheduler. The elastic iterator model generally upgrades traditional iterator model with new dynamic multi-core execution adjustment capability. And the dynamic scheduler efficiently provisions CPU cores to query execution segments in the pipelines based on the light-weight measurements on the operators. Extensive experiments on real and synthetic (TPC-H) data show that our proposal achieves almost full CPU utilization on typical decision-making analytical queries, outperforming state-of-the-art open-source systems by a huge margin.
Minqi Zhou, Yin Yang 0001, Aoying Zhou, Dina Bitton
SIGMOD Conference6
2006 Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality?
Dina Bitton, Franz Färber, Laura M. Haas, Jayavel Shanmugasundaram
VLDB1
2005 Enterprise information integration: successes, challenges and controversies
abstract
The goal of EII systems is to provide uniform access to multiple data sources without having to first load them into a data warehouse. Since the late 1990's, several EII products have appeared in the marketplace and significant experience has been accumulated from fielding such systems. This collection of articles, by individuals who were involved in this industry in various ways, describes some of these experiences and points to the challenges ahead.
Alon Y. Halevy, Naveen Ashish, Dina Bitton, Michael J. Carey 0001, Denise Draper, Jeff Pollock, Arnon Rosenthal, Vishal Sikka
SIGMOD Conference3
1991 DBE: An Expert Tool for Database Design
Dina Bitton, Jeffrey Millman, Solveig Torgersen
CAiSE1
1989 A Feasibility and Performance Study of Dependency Inference
abstract
The feasibility of inferring functional dependencies from an example relation is investigated. The problem occurs in the context of automatic database design, when a tool is needed to assist the database designer in the process of specifying logical dependencies. The complexity of the dependency inference problem is inherently exponential. However, algorithms could be developed that perform well when the input relation has certain characteristics. Two such algorithms for dependency inference are implemented and optimized. An extensive set of experiments is presented, in which dependencies were inferred from example relations with different cardinalities, number of attributes, and degree of normalization. It is concluded that for practical example relations, an adequate implementation of a dependence inference function leads to acceptable interactive response times.>
Dina Bitton, Jeffrey Millman, Solveig Torgersen
ICDE1
1989 Database Tools and Interfaces
Dina Bitton
VLDB1
1988 Disk Shadowing
Dina Bitton, Jim Gray 0001
VLDB1
1988 Sorting Large Files on a Backend Multiprocessor
abstract
The authors investigate the feasibility and efficiency of a parallel sort-merge algorithm by considering its implementation of the JASMIN prototype, a backend multiprocessor built around a fast packet bus. They describe the design and implementation of a parallel sort utility and present and analyze the results of measurements corresponding to a range of file sizes and processor configurations. The results show that using current, off-the-shelf technology coupled with a streamlined distributed operating system, three- and five-microprocessor configurations, provide a very cost-effective sort of large files. The three-processor configuration sorts a 100-Mb file in 1 hr which compares well to commercial sort packages available on high-performance mainframes. In additional experiments, the authors investigate a model to tune their sort software and scale their results to higher processor and network capabilities.>
Micah D. Beck, Dina Bitton, Kevin Wilkinson
IEEE Trans. Computers2
1987 Performance of Complex Queries in Main Memory Database Systems
abstract
Memory residence can buy both functionality and performance for a database management system. In this paper, we present a description and a benchmark of an experimental implementation of a Main Memory Database System (MMDBS) that was designed to support complex interactive queries. We describe and evaluate the main memory database structures and query processing algorithms implemented in this prototype. Our measurements and analysis, focused on aggregates and joins, include both memory requirements and response time, since there is a clear trade-off between space and time in the design of a MMDBS. In contrast to conventional Disk-based Database Systems (DDBS's), we found that an MMDBS can efficiently execute complex relational queries. We identify strategies that exploit memory residence effectively. We also identified a number of performance problems related to query optimization in main memory and memory management for MMDBS's.
Dina Bitton, Maria Hanrahan, Carolyn Turbyfill
ICDE1
1987 Technology Trends in Mass-Storage Systems
Dina Bitton
SIGMOD Conference1
1987 A general framework for computing block accesses
Bradley T. Vander Zanden, Howard M. Taylor, Dina Bitton
Inf. Syst.3
1986 Design and Evaluation of a Parallel Sort Utility
Micah D. Beck, Dina Bitton, Kevin Wilkinson
ICPP2
1986 The Effect of Large Main memory on Database Systems
Dina Bitton
SIGMOD Conference1
1986 Estimating Block Accessses when Attributes are Correlated
Bradley T. Vander Zanden, Howard M. Taylor, Dina Bitton
VLDB3
1983 Benchmarking Database Systems A Systematic Approach
Dina Bitton, David J. DeWitt, Carolyn Turbyfill
VLDB1
1983 Parallel Algorithms for the Execution of Relational Database Operations
abstract
This paper presents and analyzes algorithms for parallel processing of relational database operations in a general multiprocessor framework. To analyze alternative algorithms, we introduce an analysis methodology which incorporates I/O, CPU, and message costs and which can be adjusted to fit different multiprocessor architectures. Algorithms are presented and analyzed for sorting, projection, and join operations. While some of these algorithms have been presented and analyzed previously, we have generalized each in order to handle the case where the number of pages is significantly larger than the number of processors. In addition, we present and analyze algorithms for the parallel execution of update and aggregate operations.
Dina Bitton, Haran Boral, David J. DeWitt, Kevin Wilkinson
ACM Trans. Database Syst.1
1983 Duplicate Record Elimination in Large Data Files
abstract
The issue of duplicate elimination for large data files in which many occurrences of the same record may appear is addressed. A comprehensive cost analysis of the duplicate elimination operation is presented. This analysis is based on a combinatorial model developed for estimating the size of intermediate runs produced by a modified merge-sort procedure. The performance of this modified merge-sort procedure is demonstrated to be significantly superior to the standard duplicate elimination technique of sorting followed by a sequential pass to locate duplicate records. The results can also be used to provide critical input to a query optimizer in a relational database system.
Dina Bitton, David J. DeWitt
ACM Trans. Database Syst.1