EDBT 2026 Demo / reviewers in the wild / expert
Laura M. Haas
dblp:h/LauraMHaas
· DBLP profile ↗
46ranked-venue papers
15as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 40 · 13 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 2Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
32 papers |
Data integration and cleaning · 42% Data stream processing · 19% Data mining · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 92% Computational social science and digital humanities · 8% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 30 heaviest of 55, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computing education › STEM education
data science education |
0.3 | 1 | 2017 | Data Science Education: We're Missing the Boat, Again · ICDE 2017 |
Data integration and cleaning
schema mapping |
0.3 | 5 | 2010 | TRAMP: Understanding the Behavior of Schema Mappings through Provenance · Proc. VLDB Endow. 2010 Clio grows up: from research prototype to industrial tool · SIGMOD Conference 2005 Data-Driven Understanding and Refinement of Schema Mappings · SIGMOD Conference 2001 |
Data integration and cleaning
data exchange |
0.2 | 2 | 2011 | Debugging Data Exchange with Vagabond · Proc. VLDB Endow. 2011 Just-in-time Data Integration in Action · Proc. VLDB Endow. 2010 |
Data stream processing › continuous query processing
window semantics |
0.1 | 1 | 2010 | SECRET: A Model for Analysis of the Execution Semantics of Stream Processing Systems · Proc. VLDB Endow. 2010 |
Information retrieval
contextual search |
0.1 | 1 | 2017 | Leveraging Data and People to Accelerate Data Science · ICDE 2017 |
Data integration and cleaning
data provenance |
0.1 | 1 | 2017 | Leveraging Data and People to Accelerate Data Science · ICDE 2017 |
Data integration and cleaning
data quality |
0.1 | 1 | 2017 | Data Science Education: We're Missing the Boat, Again · ICDE 2017 |
Empirical software engineering
data science workflows |
0.1 | 1 | 2017 | Leveraging Data and People to Accelerate Data Science · ICDE 2017 |
Empirical software engineering › developer studies
user study |
0.1 | 1 | 2017 | Leveraging Data and People to Accelerate Data Science · ICDE 2017 |
Data integration and cleaning
heterogeneous data source integration |
0.0 | 2 | 2010 | A demonstration of the MaxStream federated stream processing system · ICDE 2010 The Garlic Project · SIGMOD Conference 1996 |
Distributed and cloud data management
federated database |
0.0 | 2 | 2002 | Garlic: a new flavor of federated query processing for DB2 · SIGMOD Conference 2002 Optimizing Queries Across Diverse Data Sources · VLDB 1997 |
Data integration and cleaning › schema matching
attribute correspondence |
0.0 | 1 | 2002 | Attribute Classification Using Feature Analysis · ICDE 2002 |
Query processing and optimization
query planning |
0.0 | 1 | 2002 | Garlic: a new flavor of federated query processing for DB2 · SIGMOD Conference 2002 |
Data integration and cleaning
schema matching |
0.0 | 1 | 2002 | Attribute Classification Using Feature Analysis · ICDE 2002 |
Data integration and cleaning
data transformation |
0.0 | 1 | 2001 | Data-Driven Understanding and Refinement of Schema Mappings · SIGMOD Conference 2001 |
Data models and query languages › query language
declarative query language |
0.0 | 1 | 2001 | Data-Driven Understanding and Refinement of Schema Mappings · SIGMOD Conference 2001 |
Data integration and cleaning
metadata management |
0.0 | 1 | 2000 | Panel: Is Generic Metadata Management Feasible? · VLDB 2000 |
Query processing and optimization
cost model |
0.0 | 1 | 1999 | Cost Models DO Matter: Providing Cost Information for Diverse Data Sources in a Federated System · VLDB 1999 |
Query processing and optimization
query result caching |
0.0 | 1 | 1999 | Loading a Cache with Query Results · VLDB 1999 |
Query processing and optimization
query execution |
0.0 | 2 | 2002 | Garlic: a new flavor of federated query processing for DB2 · SIGMOD Conference 2002 Tapes Hold Data, Too: Challenges of Tuples on Tertiary Store · SIGMOD Conference 1993 |
Data integration and cleaning
data fusion |
0.0 | 1 | 2006 | Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality? · VLDB 2006 |
Information retrieval
distributed information retrieval |
0.0 | 1 | 1997 | Data Structures for Efficient Broker Implementation · ACM Trans. Inf. Syst. 1997 |
Indexing and storage engines › multidimensional indexing
grid file |
0.0 | 1 | 1997 | Data Structures for Efficient Broker Implementation · ACM Trans. Inf. Syst. 1997 |
Query processing and optimization › cost estimation
join cost estimation |
0.0 | 1 | 1997 | Seeking the Truth About ad hoc Join Costs · VLDB J. 1997 |
Query processing and optimization
join processing |
0.0 | 1 | 1997 | Seeking the Truth About ad hoc Join Costs · VLDB J. 1997 |
Data models and query languages
object-oriented database |
0.0 | 1 | 1996 | PESTO : An Integrated Query/Browser for Object Databases · VLDB 1996 |
Database system architecture and tuning
extensible database system |
0.0 | 2 | 1990 | Starburst Mid-Flight: As the Dust Clears · IEEE Trans. Knowl. Data Eng. 1990 Extensible Query Processing in Starburst · SIGMOD Conference 1989 |
Storage systems › magnetic storage
tape storage |
0.0 | 1 | 1993 | Tapes Hold Data, Too: Challenges of Tuples on Tertiary Store · SIGMOD Conference 1993 |
Storage systems › storage hierarchy
tertiary storage |
0.0 | 1 | 1993 | Tapes Hold Data, Too: Challenges of Tuples on Tertiary Store · SIGMOD Conference 1993 |
Memory systems
cache |
0.0 | 1 | 1999 | Loading a Cache with Query Results · VLDB 1999 |
Methods — techniques the papers use, named apart from their topics
user study · 0.6provenance queries · 0.6query translation · 0.1persistency management · 0.1experimental analysis · 0.1descriptive modeling · 0.1speculative research synthesis · 0.1query graph · 0.1XQuery · 0.1SQL/XML · 0.1simulation · 0.0distributed detection algorithm · 0.0virtual circuit communication · 0.0process-per-user computation · 0.0process tree computation model · 0.0communication protocol · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Leveraging Data and People to Accelerate Data ScienceabstractDoing data science - extracting insight by analyzing data - is not easy. Data science is used to answer interesting questions that typically involve multiple diverse data sources, many different types of analysis, and often, large and messy data volumes. To answer one of these questions, several types of expertise may be needed to understand the context and domain being served, to import and transform individual data sets, to implement effective machine learning and/or statistical methods, to design and program applications and interfaces to extract and share data and insights, and to manage the data and systems used for analysis and storage. In the IBM Research Accelerated Discovery Lab, we are studying how data scientists work, and using what we learn to help them gain insights faster. In this talk, we will look at what we have learned to date, through user studies and experience with tens of analytics projects, and the environment that we've built as a result. In particular, I will describe how we capture information to enable contextual search, provenance queries, and other functionality to afford teams faster progress in data-intensive investigations. I will also touch on our efforts to leverage data and people to explain what happens during an investigation, with an ultimate goal of moving from descriptive to prescriptive analytics in order to accelerate data science and the analytic process. I will illustrate these various efforts using an ambitious current project on applying metagenomics to food safety, and will conclude with a discussion of where more work is needed and our future directions. Laura M. Haas |
ICDE | 1 |
| 2017 | Data Science Education: We're Missing the Boat, AgainabstractIn the first wave of data science education programs, data engineering topics (systems, scalable algorithms, data management, integration) tended to be de-emphasized in favor of machine learning and statistical modeling. The anecdotal evidence suggests this was a mistake: data scientists report spending most of their time grappling with data far upstream of modeling activities. A second wave of data science education is emerging, one with increased emphasis on practical issues in ethics, legal compliance, scientific reproducibility, data quality, and algorithmic bias. The data engineering community has a second chance to influence these programs beyond just providing a set of tools. In this panel, we'll discuss the role of data engineering in data science education programs, and how best to capitalize on emerging opportunities in this space. Bill Howe, Michael J. Franklin, Laura M. Haas, Tim Kraska, Jeffrey D. Ullman |
ICDE | 3 |
| 2015 | The Power Behind the Throne: Information Integration in the Age of Data-Driven DiscoveryabstractIntegrating data has always been a challenge. The information management community has made great progress in tackling this challenge, both on the theory and the practice. But in the last ten years, the world has changed dramatically. New platforms, devices and applications have made huge volumes of heterogeneous data available at speeds never contemplated before, while the quality of the available data has if anything degraded. Unstructured and semi-structured formats and no-sql data stores undercut the old reliable tools of schema, forcing applications to deal with data at the instance level. Deep expertise in the data and domain, in the tools and systems for integration and analysis, in mathematics, computer science, and business are needed to discover insights from data, but rarely are all of these skills found in a single individual or even team. Meanwhile, the availability of all these data has raised expectations for rapid breakthroughs in many sciences, for quick solutions to business problems, and for ever more sophisticated applications that combine and analyze information to solve our daily needs. These expectations raise the bar for integration technology, while opening the door for it to play a broader role. Integration has always been a key player in handling data variety, for example, but now more than ever must deal with scale (in the number of types as well as in the volume and speed of data). While data cleansing has been one step of an integration pipeline, this technology must be leveraged throughout data integration, so that the integration process is better able to deal with the uncertainty in data, offering means to eliminate or reduce it, or, to elucidate it by linking important contextual information, such as provenance and usage. The complexity of today's data-driven challenges in fact suggests that the integration process should be context-aware, so that data sets may be combined differently depending on the proposed usage. Laura M. Haas |
SIGMOD Conference | 1 |
| 2013 | Modeling the execution semantics of stream processing engines with SECRET
Nihal Dindar, Nesime Tatbul, Renée J. Miller, Laura M. Haas, Irina Botan |
VLDB J. | 4 |
| 2011 | Debugging Data Exchange with Vagabond
Boris Glavic, Jiang Du 0001, Renée J. Miller, Gustavo Alonso, Laura M. Haas |
Proc. VLDB Endow. | 5 |
| 2010 | A demonstration of the MaxStream federated stream processing systemabstractMaxStream is a federated stream processing system that seamlessly integrates multiple autonomous and heterogeneous Stream Processing Engines (SPEs) and databases. In this paper, we propose to demonstrate the key features of MaxStream using two application scenarios, namely the Sales Map & Spikes business monitoring scenario and the Linear Road Benchmark, each with a different set of requirements. More specifically, we will show how the MaxStream Federator can translate and forward the application queries to two different commercial SPEs (Coral8 and StreamBase), as well as how it does so under various persistency requirements. Irina Botan, Younggoo Cho, Roozbeh Derakhshan, Nihal Dindar, Laura M. Haas, Chulwon Lee, Girish Mundada, Ming-Chien Shan, Nesime Tatbul, Beomjin Yun |
ICDE | 6 |
| 2010 | Foreword
Manish Bhide, Laura M. Haas, Zachary G. Ives, Mukesh K. Mohania |
Inf. Syst. | 2 |
| 2010 | Time for Our Field to Grow UpabstractCompared to centuries of physics and millennia of mathematics, the 50-year-history of computer science and information management research makes us the toddlers of the scientific community. Yet during our brief existence, we've revolutionized the world and, not content with that, gone on to build and study virtual worlds. We have justly taken pride in our accomplishments, and developed our own unique way of conducting research, unlike other scientific and engineering fields. But cracks have appeared in this edifice we have built. The conference system that served us so well for our first 50 years is falling apart. Our ever-increasing population competes ever more energetically for a finite set of resources. Other scientific and engineering disciplines still think that our field equates to programming, and look down on us. While we may also look down on them, it is undeniably true that high-energy physicists get many more research dollars per capita than we do, and our computer science colleagues wonder whether all the data management problems haven't already been solved. Other departments have started to teach courses that overlap our turf. Are we our own worst enemies? Why doesn't everyone understand how important our research is? Do we have to abandon the conference system? Must we become more like the stodgy old fields of science and engineering? Or can we find our own way? Anastasia Ailamaki, Laura M. Haas, H. V. Jagadish, David Maier 0001, M. Tamer Özsu, Marianne Winslett |
Proc. VLDB Endow. | 2 |
| 2010 | SECRET: A Model for Analysis of the Execution Semantics of Stream Processing SystemsabstractThere are many academic and commercial stream processing engines (SPEs) today, each of them with its own execution semantics. This variation may lead to seemingly inexplicable differences in query results. In this paper, we present SECRET, a model of the behavior of SPEs. SECRET is a descriptive model that allows users to analyze the behavior of systems and understand the results of window-based queries for a broad range of heterogeneous SPEs. The model is the result of extensive analysis and experimentation with several commercial and academic engines. In the paper, we describe the types of heterogeneity found in existing engines, and show with experiments on real systems that our model can explain the key differences in windowing behavior. Irina Botan, Roozbeh Derakhshan, Nihal Dindar, Laura M. Haas, Renée J. Miller, Nesime Tatbul |
Proc. VLDB Endow. | 4 |
| 2010 | TRAMP: Understanding the Behavior of Schema Mappings through ProvenanceabstractThough partially automated, developing schema mappings remains a complex and potentially error-prone task. In this paper, we present TRAMP (TRAnsformation Mapping Provenance), an extensive suite of tools supporting the debugging and tracing of schema mappings and transformation queries. TRAMP combines and extends data provenance with two novel notions, transformation provenance and mapping provenance, to explain the relationship between transformed data and those transformations and mappings that produced that data. In addition we provide query support for transformations, data, and all forms of provenance. We formally define transformation and mapping provenance, present an efficient implementation of both forms of provenance, and evaluate the resulting system through extensive experiments. Boris Glavic, Gustavo Alonso, Renée J. Miller, Laura M. Haas |
Proc. VLDB Endow. | 4 |
| 2010 | Just-in-time Data Integration in ActionabstractToday's data integration systems must be flexible enough to support the typical iterative and incremental process of integration, and may need to scale to hundreds of data sources. In this work we present a novel data integration system that offers great flexibility and scalability. Our approach to data integration is unique in that it executes mapping rules at query runtime using annotations. On top, we have built the People People People application. It allows users to search for people, display information about people, and browse through a network of related people, where the data is integrated from local and remote data sources. The demo presents all features of our underlying data integration engine through a set of motivating scenarios. Martin Hentschel 0001, Laura M. Haas, Renée J. Miller |
Proc. VLDB Endow. | 2 |
| 2009 | New Challenges in Information Integration
Laura M. Haas, Aya Soffer |
DaWaK | 1 |
| 2009 | Schema AND Data: A Holistic Approach to Mapping, Resolution and Fusion in Information Integration
Laura M. Haas, Martin Hentschel 0001, Donald Kossmann, Renée J. Miller |
ER | 1 |
| 2008 | Impact! The Challenge of Industrial Research in Computer Science in a web 2.0 world
Laura M. Haas |
SEKE | 1 |
| 2007 | Information for PeopleabstractOrdinary people have access to unprecedented volumes of information today. Researchers in the fields of information management (IM) and human-computer interaction (HCI) are reacting to this challenge from their own unique perspectives. Having access to a billion records is cool, but having access to a billion people is awesome. In this paper, we look at recent research from both communities, and speculate on how interactions between the communities could enhance the user experience of information. Laura M. Haas, Steve B. Cousins |
ICDE | 1 |
| 2007 | Beauty and the Beast: The Theory and Practice of Information Integration
Laura M. Haas |
ICDT | 1 |
| 2007 | Special issue: best papers of VLDB 2005
Laura M. Haas, Christian S. Jensen, Martin L. Kersten |
VLDB J. | 1 |
| 2006 | Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality?
Dina Bitton, Franz Färber, Laura M. Haas, Jayavel Shanmugasundaram |
VLDB | 3 |
| 2005 | Clio grows up: from research prototype to industrial toolabstractClio, the IBM Research system for expressing declarative schema mappings, has progressed in the past few years from a research prototype into a technology that is behind some of IBM's mapping technology. Clio provides a declarative way of specifying schema mappings between either XML or relational schemas. Mappings are compiled into an abstract query graph representation that captures the transformation semantics of the mappings. The query graph can then be serialized into different query languages, depending on the kind of schemas and systems involved in the mapping. Clio currently produces XQuery, XSLT, SQL, and SQL/XML queries. In this paper, we revisit the architecture and algorithms behind Clio. We then discuss some implementation issues, optimizations needed for scalability, and general lessons learned in the road towards creating an industrial-strength tool. Laura M. Haas, Mauricio A. Hernández, C. T. Howard Ho, Lucian Popa 0001, Mary Roth |
SIGMOD Conference | 1 |
| 2002 | Attribute Classification Using Feature AnalysisabstractThe basis of many systems that integrate data from multiple sources is a set of correspondences between source schemata and a target schema. Correspondences express a relationship between sets of source attributes, possibly from multiple sources, and a set of target attributes. Clio is an integration tool that assists users in defining value correspondences between attributes. In real life scenarios there may be many sources and the source relations may have many attributes. Users can get lost and might miss or be unable to find some correspondences. Also, in many real life schemata the attribute names reveal little or nothing about the semantics of the data values. Only the data values in the attribute columns can convey the semantic meaning of the attribute. Our work relieves users of the problems of too many attributes and meaningless attribute names, by automatically suggesting correspondences between source and target attributes. For each attribute, we analyze the data values and derive a set of features. Felix Naumann, C. T. Howard Ho, Xuqing Tian, Laura M. Haas, Nimrod Megiddo |
ICDE | 4 |
| 2002 | Garlic: a new flavor of federated query processing for DB2abstractIn a large modern enterprise, information is almost inevitably distributed among several database management systems. Despite considerable attention from the research community, relatively few commercial systems have attempted to address this issue. This paper describes new technology that enables clients of IBM's DB2 Universal Database to access the data and specialized computational capabilities of a wide range of non-relational data sources. This technology, based on the Garlic prototype developed at the Almaden Research Center, complements and extends DB2's existing ability to federate relational data sources.The paper focuses on three topics. Firstly, we show how the DB2 catalogs are used as an extensible repository for the metadata needed to access remotely-stored information. Secondly, we describe how the Garlic approach to query planning, in which source-specific modules and the federated server cooperate to develop an optimized execution plan, has been realized in DB2. Lastly, we describe how DB2's query execution engine has been extended to support queries and functions that are evaluated remotely. Vanja Josifovski, Peter M. Schwarz, Laura M. Haas, Eileen Tien Lin |
SIGMOD Conference | 3 |
| 2001 | Clio: A Semi-Automatic Tool For Schema MappingabstractWe consider the integration requirements of modern data intensive applications including data warehousing, global information systems and electronic commerce. At the heart of these requirements lies the schema mapping problem in which a source (legacy) database must be mapped into a different, but xed, target schema. The goal of schema mapping is the discovery of a query or set of queries to map source databases into the new structure. We demonstrate Clio, a new semi-automated tool for creating schema mappings. Clio employs a mapping-by-example paradigm that relies on the use of value correspondences describing how a value of a target attribute can be created from a set of values of source attributes. A typical session with Clio starts with the user loading a source and a target schema into the system. These schemas are read from either an underlying Object-Relational database or from an XML le with an associated XML Schema. Users can then draw value correspondences mapping source attributes into target attributes. Clio's mapping engine incrementally produces the SQL queries that realize the mappings implied by the correspondences. Clio provides schema and data browsers and other feedback to allow users to understand the mapping produced. Entering and manipulating value correspondences can be done in two modes. In the Schema View mode, users see a representation of the source and target schema and create value correspondences by selecting schema objects from the source and mapping them to a target attribute. The alternative Data View mode o ers a WYSIWYG interface for the mapping process that displays example data for both the source and target tables [3]. Users may add and delete value correspondences from this view and immediately see the changes re ected in the resulting target tuples. Also, the Data View mode helps users navigate through alternative mappings, understanding the often subtle di erences between them. For example, in some cases, changing a join from an inner join to an outer join may dramatically change the resulting table. In other cases, the same change may have no e ect due to constraints that hold on the source Mauricio A. Hernández, Renée J. Miller, Laura M. Haas |
SIGMOD Conference | 3 |
| 2001 | Data-Driven Understanding and Refinement of Schema MappingsabstractAt the heart of many data-intensive applications is the problem of quickly and accurately transforming data into a new form. Database researchers have long advocated the use of declarative queries for this process. Yet tools for creating, managing and understanding the complex queries necessary for data transformation are still too primitive to permit widespread adoption of this approach. We present a new framework that uses data examples as the basis for understanding and refining declarative schema mappings. We identify a small set of intuitive operators for manipulating examples. These operators permit a user to follow and refine an example by walking through a data source. We show that our operators are powerful enough both to identify a large class of schema mappings and to distinguish effectively between alternative schema mappings. These operators permit a user to quickly and intuitively build and refine complex data transformation queries that map one data source into another. Ling-Ling Yan, Renée J. Miller, Laura M. Haas, Ronald Fagin |
SIGMOD Conference | 3 |
| 2000 | Integrating Life Sciences Data - With a Little GarlicabstractVast amounts of life sciences data today reside in specialized data sources, with specialized query processing capabilities. Data from one source must often be combined with data from other sources to give users the information they desire. Database middleware systems such as Garlic allow users to combine data from multiple sources in a single query. Garlic provides the user with a virtual database to which they can pose arbitrarily complex queries, though the actual data needed to answer the query may be stored in several different sources, and those sources may not even possess all the functionality needed to answer such a query themselves. The Garlic technology, as incorporated in IBM's DB2 product, forms the basis of the DiscoveryLink service offering for the life sciences industry. We describe the DiscoveryLink offering, focusing on two key contributions of Garlic, the wrapper architecture and the query optimizer, and illustrate how it can be used to integrate life sciences data from heterogeneous data sources. Laura M. Haas, Prasad Kodali, Julia E. Rice, Peter M. Schwarz, William C. Swope |
BIBE | 1 |
| 2000 | Panel: Is Generic Metadata Management Feasible?
Philip A. Bernstein, Laura M. Haas, Matthias Jarke, Erhard Rahm, Gio Wiederhold |
VLDB | 2 |
| 2000 | Schema Mapping as Query Discovery
Renée J. Miller, Laura M. Haas, Mauricio A. Hernández |
VLDB | 2 |
| 1999 | Using Fagin's Algorithm for Merging Ranked Results in Multimedia MiddlewareabstractA distributed multimedia information system allows users to access data of different modalities, from different data sources, ranked by various combinations of criteria. Fagin (1996) gives an algorithm for efficiently merging multiple ordered streams of ranked results, to form a new stream ordered by a combination of those ranks. In this paper we describe the implementation of Fagin's algorithm in an actual multimedia middleware system, including a novel, incremental version of the algorithm that supports dynamic exploration of data. We show that the algorithm would perform well as part of a single multimedia server and can even be effective in the distributed environment (for a limited set of queries), but that the assumptions it makes about random access limit its applicability dramatically. Our experience provides a better understanding of an important algorithm, and exposes an open problem for distributed multimedia information systems. Edward L. Wimmers, Laura M. Haas, Mary Roth, Christoph Braendli |
CoopIS | 2 |
| 1999 | Loading a Cache with Query Results
Laura M. Haas, Donald Kossmann, Ioana Ursu |
VLDB | 1 |
| 1999 | Cost Models DO Matter: Providing Cost Information for Diverse Data Sources in a Federated System
Mary Roth, Fatma Özcan 0001, Laura M. Haas |
VLDB | 3 |
| 1998 | Capabilities-Based Query Rewriting in Mediator Systems
Yannis Papakonstantinou, Ashish Gupta 0001, Laura M. Haas |
Distributed Parallel Databases | 3 |
| 1997 | Optimizing Queries Across Diverse Data Sources
Laura M. Haas, Donald Kossmann, Edward L. Wimmers, Jun Yang 0001 |
VLDB | 1 |
| 1997 | Data Structures for Efficient Broker ImplementationabstractWith the profusion of text databases on the Internet, it is becoming increasingly hard to find the most useful databases for a given query. To attack this problem, several existing and proposed systems employ brokers to direct user queries, using a local database of summary information about the available databases. This summary information must effectively distinguish relevant databases and must be compact while allowing efficient access. We offer evidence that one broker, GlOSS , can be effective at locating databases of interest even in a system of hundreds of databased and can examine the performance of accessing the GlOSS summeries for two promising storage methods: the grid file and partitioned hashing. We show that both methods can be tuned to provide good performance for a particular workload (within a broad range of workloads), and we discuss the tradeoffs between the two data structures. As a side effect of our work, we show that grid files are more broadly applicable than previously thought; inparticular, we show that by varying the policies used to construct the grid file we can provide good performance for a wide range of workloads even when storing highly skewed data. Anthony Tomasic, Luis Gravano, Calvin Lue, Peter M. Schwarz, Laura M. Haas |
ACM Trans. Inf. Syst. | 5 |
| 1997 | Seeking the Truth About ad hoc Join Costs
Laura M. Haas, Michael J. Carey 0001, Miron Livny, Amit Shukla 0001 |
VLDB J. | 1 |
| 1996 | The Garlic ProjectabstractThe goal of the Garlic [1] project is to build a multimedia information system capable of integrating data that resides in different database systems as well as in a variety of non-database data servers. This integration must be enabled while maintaining the independence of the data servers, and without creating copies of their data. "Multimedia" should be interpreted broadly to mean not only images, video, and audio, but also text and application specific data types (e.g., CAD drawings, medical objects, …). Since much of this data is naturally modeled by objects, Garlic provides an object-oriented schema to applications, interprets object queries, creates execution plans for sending pieces of queries to the appropriate data servers, and assembles query results for delivery back to the applications. A significant focus of the project is support for "intelligent" data servers, i.e., servers that provide media-specific indexing and query capabilities [2]. Database optimization technology is being extended to deal with heterogeneous collections of data servers so that efficient data access plans can be employed for multi-repository queries.A prototype of the Garlic system has been operational since January 1995. Queries are expressed in an SQL-like query language that has been extended to include object-oriented features such as reference-valued attributes and nested sets. In addition to a C++ API, Garlic supports a novel query/browser interface called PESTO [3]. This component of Garlic provides end users of the system with a friendly, graphical interface that supports interactive browsing, navigation, and querying of the contents of Garlic databases. Unlike existing interfaces to databases, PESTO allows users to move back and forth seamlessly between querying and browsing activities, using queries to identify interesting subsets of the database, browsing the subset, querying the content of a set-valued attribute of a particularly interesting object in the subset, and so on. Mary Roth, Manish Arya, Laura M. Haas, Michael J. Carey 0001, William F. Cody, Ronald Fagin, Peter M. Schwarz, Joachim Thomas 0002, Edward L. Wimmers |
SIGMOD Conference | 3 |
| 1996 | PESTO : An Integrated Query/Browser for Object Databases
Michael J. Carey 0001, Laura M. Haas, Vivekananda Maganty, John H. Williams |
VLDB | 2 |
| 1993 | Tapes Hold Data, Too: Challenges of Tuples on Tertiary StoreabstractArticle Free Access Share on Tapes hold data, too: challenges of tuples on tertiary store Authors: Michael J. Carey Computer Science Dept., University of Wisconsin, Madison, WI Computer Science Dept., University of Wisconsin, Madison, WIView Profile , Laura M. Haas IBM Almaden Research Center, K55/801, San Jose, CA IBM Almaden Research Center, K55/801, San Jose, CAView Profile , Miron Livny Computer Science Dept., University of Wisconsin, Madison, WI Computer Science Dept., University of Wisconsin, Madison, WIView Profile Authors Info & Claims SIGMOD '93: Proceedings of the 1993 ACM SIGMOD international conference on Management of dataJune 1993Pages 413–417https://doi.org/10.1145/170035.170103Published:01 June 1993Publication History 30citation211DownloadsMetricsTotal Citations30Total Downloads211Last 12 Months55Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Michael J. Carey 0001, Laura M. Haas, Miron Livny |
SIGMOD Conference | 2 |
| 1990 | Starburst Mid-Flight: As the Dust ClearsabstractThe purpose of the Starburst project is to improve the design of relational database management systems and enhance their performance, while building an extensible system to better support nontraditional applications and to serve as a testbed for future improvements in database technology. The design and implementation of the Starburst system to date are considered. Some key design decisions and how they affect the goal of improved structure and performance are examined. How well the goal of extensibility has been met is examined: what aspects of the system are extensible, how extensions can be done, and how easy it is to add extensions. Some actual extensions to the system, including the experiences of the first real customizers, are discussed.> Laura M. Haas, Walter Chang, Guy M. Lohman, John McPherson, Paul F. Wilms, George Lapis, Bruce G. Lindsay 0001, Hamid Pirahesh, Michael J. Carey 0001, Eugene J. Shekita |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1989 | Extensible Query Processing in StarburstabstractToday's DBMSs are unable to support the increasing demands of the various applications that would like to use a DBMS. Each kind of application poses new requirements for the DBMS. The Starburst project at IBM's Almaden Research Center aims to extend relational DBMS technology to bridge this gap between applications and the DBMS. While providing a full function relational system to enable sharing across applications, Starburst will also allow (sophisticated) programmers to add many kinds of extensions to the base system's capabilities, including language extensions (e.g., new datatypes and operations), data management extensions (e.g., new access and storage methods) and internal processing extensions (e.g., new join methods and new query transformations). To support these features, the database query language processor must be very powerful and highly extensible. Starburst's language processor features a powerful query language, rule-based optimization and query rewrite, and an execution system based on an extended relational algebra. In this paper, we describe the design of Starburst's query language processor and discuss the ways in which the language processor can be extended to achieve Starburst's goals. Laura M. Haas, Johann-Christoph Freytag, Guy M. Lohman, Hamid Pirahesh |
SIGMOD Conference | 1 |
| 1988 | Views and Security in Distributed Database Management Systems
Elisa Bertino, Laura M. Haas |
EDBT | 2 |
| 1986 | A Snapshot Differential Refresh AlgorithmabstractThis article presents an algorithm to refresh the contents of database snapshots. A database snapshot is a read-only table whose contents are extracted from other tables in the database. The snapshot contents can be periodically refreshed to reflect the current state of the database. Snapshots are useful in many applications as a cost effective substitute for replicated data in a distributed database system. Bruce G. Lindsay 0001, Laura M. Haas, C. Mohan 0001, Hamid Pirahesh, Paul F. Wilms |
SIGMOD Conference | 2 |
| 1984 | Optimization of Nested Queries in a Distributed Relational Database
Guy M. Lohman, Dean Daniels, Laura M. Haas, Ruth Kistler, Patricia G. Selinger |
VLDB | 3 |
| 1984 | Computation and Communication in R*: A Distributed Database ManagerabstractThis article presents and discusses the computation and communication model used by R*, a prototype distributed database management system.An R* computation consists of a tree of processes connected by virtual circuit communication paths.The process management and communication protocols used by R* enable the system to provide reliable, distributed transactions while maintaining adequate levels of performance.Of particular interest is the use of processes in R* to retain user context from one transaction to another, in order to improve the system performance and recovery characteristics. Bruce G. Lindsay 0001, Laura M. Haas, C. Mohan 0001, Paul F. Wilms, Robert A. Yost |
ACM Trans. Comput. Syst. | 2 |
| 1983 | Computation & Communication in R*: A Distributed Database Manager (Extended Abstract)abstractR* is an experimental prototype distributed database management system. The computation needed to perform a sequence of multisite user transactions in R* is structured as a tree of processes communicating over virtual circuit communication links. Distributed computation can be supported by providing a server process per site which performs requests on behalf of remote users. Alternatively, a new process could be created to service each incoming request. Instead of using a shared server process or using the process per request approach, R* creates a process associated with the computation of the user on the first request to the remote site. This process is incorporated into the tree of processes serving a single user and is retained for the duration of the user computation. This approach allows R* to factor some of the request execution overhead into the process creation phase, and simplifies the retention of user and transaction context at the multiple sites of the distributed computation. Bruce G. Lindsay 0001, Laura M. Haas, C. Mohan 0001, Paul F. Wilms, Robert A. Yost |
SOSP | 2 |
| 1983 | View Management in Distributed Data Base Systems
Elisa Bertino, Laura M. Haas, Bruce G. Lindsay 0001 |
VLDB | 2 |
| 1983 | Site autonomy issues in R*: A distributed database management system
Patricia G. Selinger, Dean Daniels, Laura M. Haas, Bruce G. Lindsay 0001, Pui Ng, Paul F. Wilms, Robert A. Yost |
Inf. Sci. | 3 |
| 1983 | Distributed Deadlock DetectionabstractDistributed deadlock models are presented for resource and communication deadlocks.Simple distributed algorithms for detection of these deadlocks are given.We show that all true deadlocks are detected and that no false deadlocks are reported.In our algorithms, no process maintains global information; all messages have an identical short length.The algorithms can be applied in distributed database and other message communication systems. K. Mani Chandy, Jayadev Misra, Laura M. Haas |
ACM Trans. Comput. Syst. | 3 |