Laura M. Haas

dblp:h/LauraMHaas · DBLP profile ↗
← Back
46ranked-venue papers
15as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 40 · 13 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 2Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
32 papers
Data integration and cleaning · 42% Data stream processing · 19% Data mining · 17%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computing education · 92% Computational social science and digital humanities · 8%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 30 heaviest of 55, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computing education › STEM education
data science education
0.312017
Data Science Education: We're Missing the Boat, Again · ICDE 2017
Data integration and cleaning
schema mapping
0.352010
TRAMP: Understanding the Behavior of Schema Mappings through Provenance · Proc. VLDB Endow. 2010
Clio grows up: from research prototype to industrial tool · SIGMOD Conference 2005
Data-Driven Understanding and Refinement of Schema Mappings · SIGMOD Conference 2001
Data integration and cleaning
data exchange
0.222011
Debugging Data Exchange with Vagabond · Proc. VLDB Endow. 2011
Just-in-time Data Integration in Action · Proc. VLDB Endow. 2010
Data stream processing › continuous query processing
window semantics
0.112010
SECRET: A Model for Analysis of the Execution Semantics of Stream Processing Systems · Proc. VLDB Endow. 2010
Information retrieval
contextual search
0.112017
Leveraging Data and People to Accelerate Data Science · ICDE 2017
Data integration and cleaning
data provenance
0.112017
Leveraging Data and People to Accelerate Data Science · ICDE 2017
Data integration and cleaning
data quality
0.112017
Data Science Education: We're Missing the Boat, Again · ICDE 2017
Empirical software engineering
data science workflows
0.112017
Leveraging Data and People to Accelerate Data Science · ICDE 2017
Empirical software engineering › developer studies
user study
0.112017
Leveraging Data and People to Accelerate Data Science · ICDE 2017
Data integration and cleaning
heterogeneous data source integration
0.022010
A demonstration of the MaxStream federated stream processing system · ICDE 2010
The Garlic Project · SIGMOD Conference 1996
Distributed and cloud data management
federated database
0.022002
Garlic: a new flavor of federated query processing for DB2 · SIGMOD Conference 2002
Optimizing Queries Across Diverse Data Sources · VLDB 1997
Data integration and cleaning › schema matching
attribute correspondence
0.012002
Attribute Classification Using Feature Analysis · ICDE 2002
Query processing and optimization
query planning
0.012002
Garlic: a new flavor of federated query processing for DB2 · SIGMOD Conference 2002
Data integration and cleaning
schema matching
0.012002
Attribute Classification Using Feature Analysis · ICDE 2002
Data integration and cleaning
data transformation
0.012001
Data-Driven Understanding and Refinement of Schema Mappings · SIGMOD Conference 2001
Data models and query languages › query language
declarative query language
0.012001
Data-Driven Understanding and Refinement of Schema Mappings · SIGMOD Conference 2001
Data integration and cleaning
metadata management
0.012000
Panel: Is Generic Metadata Management Feasible? · VLDB 2000
Query processing and optimization
cost model
0.011999
Cost Models DO Matter: Providing Cost Information for Diverse Data Sources in a Federated System · VLDB 1999
Query processing and optimization
query result caching
0.011999
Loading a Cache with Query Results · VLDB 1999
Query processing and optimization
query execution
0.022002
Garlic: a new flavor of federated query processing for DB2 · SIGMOD Conference 2002
Tapes Hold Data, Too: Challenges of Tuples on Tertiary Store · SIGMOD Conference 1993
Data integration and cleaning
data fusion
0.012006
Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality? · VLDB 2006
Information retrieval
distributed information retrieval
0.011997
Data Structures for Efficient Broker Implementation · ACM Trans. Inf. Syst. 1997
Indexing and storage engines › multidimensional indexing
grid file
0.011997
Data Structures for Efficient Broker Implementation · ACM Trans. Inf. Syst. 1997
Query processing and optimization › cost estimation
join cost estimation
0.011997
Seeking the Truth About ad hoc Join Costs · VLDB J. 1997
Query processing and optimization
join processing
0.011997
Seeking the Truth About ad hoc Join Costs · VLDB J. 1997
Data models and query languages
object-oriented database
0.011996
PESTO : An Integrated Query/Browser for Object Databases · VLDB 1996
Database system architecture and tuning
extensible database system
0.021990
Starburst Mid-Flight: As the Dust Clears · IEEE Trans. Knowl. Data Eng. 1990
Extensible Query Processing in Starburst · SIGMOD Conference 1989
Storage systems › magnetic storage
tape storage
0.011993
Tapes Hold Data, Too: Challenges of Tuples on Tertiary Store · SIGMOD Conference 1993
Storage systems › storage hierarchy
tertiary storage
0.011993
Tapes Hold Data, Too: Challenges of Tuples on Tertiary Store · SIGMOD Conference 1993
Memory systems
cache
0.011999
Loading a Cache with Query Results · VLDB 1999

Methods — techniques the papers use, named apart from their topics

user study · 0.6provenance queries · 0.6query translation · 0.1persistency management · 0.1experimental analysis · 0.1descriptive modeling · 0.1speculative research synthesis · 0.1query graph · 0.1XQuery · 0.1SQL/XML · 0.1simulation · 0.0distributed detection algorithm · 0.0virtual circuit communication · 0.0process-per-user computation · 0.0process tree computation model · 0.0communication protocol · 0.0
YearPublicationVenuePosition
2017 Leveraging Data and People to Accelerate Data Science
abstract
Doing data science - extracting insight by analyzing data - is not easy. Data science is used to answer interesting questions that typically involve multiple diverse data sources, many different types of analysis, and often, large and messy data volumes. To answer one of these questions, several types of expertise may be needed to understand the context and domain being served, to import and transform individual data sets, to implement effective machine learning and/or statistical methods, to design and program applications and interfaces to extract and share data and insights, and to manage the data and systems used for analysis and storage. In the IBM Research Accelerated Discovery Lab, we are studying how data scientists work, and using what we learn to help them gain insights faster. In this talk, we will look at what we have learned to date, through user studies and experience with tens of analytics projects, and the environment that we've built as a result. In particular, I will describe how we capture information to enable contextual search, provenance queries, and other functionality to afford teams faster progress in data-intensive investigations. I will also touch on our efforts to leverage data and people to explain what happens during an investigation, with an ultimate goal of moving from descriptive to prescriptive analytics in order to accelerate data science and the analytic process. I will illustrate these various efforts using an ambitious current project on applying metagenomics to food safety, and will conclude with a discussion of where more work is needed and our future directions.
Laura M. Haas
ICDE1
2017 Data Science Education: We're Missing the Boat, Again
abstract
In the first wave of data science education programs, data engineering topics (systems, scalable algorithms, data management, integration) tended to be de-emphasized in favor of machine learning and statistical modeling. The anecdotal evidence suggests this was a mistake: data scientists report spending most of their time grappling with data far upstream of modeling activities. A second wave of data science education is emerging, one with increased emphasis on practical issues in ethics, legal compliance, scientific reproducibility, data quality, and algorithmic bias. The data engineering community has a second chance to influence these programs beyond just providing a set of tools. In this panel, we'll discuss the role of data engineering in data science education programs, and how best to capitalize on emerging opportunities in this space.
Bill Howe, Michael J. Franklin, Laura M. Haas, Tim Kraska, Jeffrey D. Ullman
ICDE3
2015 The Power Behind the Throne: Information Integration in the Age of Data-Driven Discovery
abstract
Integrating data has always been a challenge. The information management community has made great progress in tackling this challenge, both on the theory and the practice. But in the last ten years, the world has changed dramatically. New platforms, devices and applications have made huge volumes of heterogeneous data available at speeds never contemplated before, while the quality of the available data has if anything degraded. Unstructured and semi-structured formats and no-sql data stores undercut the old reliable tools of schema, forcing applications to deal with data at the instance level. Deep expertise in the data and domain, in the tools and systems for integration and analysis, in mathematics, computer science, and business are needed to discover insights from data, but rarely are all of these skills found in a single individual or even team. Meanwhile, the availability of all these data has raised expectations for rapid breakthroughs in many sciences, for quick solutions to business problems, and for ever more sophisticated applications that combine and analyze information to solve our daily needs. These expectations raise the bar for integration technology, while opening the door for it to play a broader role. Integration has always been a key player in handling data variety, for example, but now more than ever must deal with scale (in the number of types as well as in the volume and speed of data). While data cleansing has been one step of an integration pipeline, this technology must be leveraged throughout data integration, so that the integration process is better able to deal with the uncertainty in data, offering means to eliminate or reduce it, or, to elucidate it by linking important contextual information, such as provenance and usage. The complexity of today's data-driven challenges in fact suggests that the integration process should be context-aware, so that data sets may be combined differently depending on the proposed usage.
Laura M. Haas
SIGMOD Conference1
2013 Modeling the execution semantics of stream processing engines with SECRET
Nihal Dindar, Nesime Tatbul, Renée J. Miller, Laura M. Haas, Irina Botan
VLDB J.4
2011 Debugging Data Exchange with Vagabond
Boris Glavic, Jiang Du 0001, Renée J. Miller, Gustavo Alonso, Laura M. Haas
Proc. VLDB Endow.5
2010 A demonstration of the MaxStream federated stream processing system
abstract
MaxStream is a federated stream processing system that seamlessly integrates multiple autonomous and heterogeneous Stream Processing Engines (SPEs) and databases. In this paper, we propose to demonstrate the key features of MaxStream using two application scenarios, namely the Sales Map & Spikes business monitoring scenario and the Linear Road Benchmark, each with a different set of requirements. More specifically, we will show how the MaxStream Federator can translate and forward the application queries to two different commercial SPEs (Coral8 and StreamBase), as well as how it does so under various persistency requirements.
Irina Botan, Younggoo Cho, Roozbeh Derakhshan, Nihal Dindar, Laura M. Haas, Chulwon Lee, Girish Mundada, Ming-Chien Shan, Nesime Tatbul, Beomjin Yun
ICDE6
2010 Foreword
Manish Bhide, Laura M. Haas, Zachary G. Ives, Mukesh K. Mohania
Inf. Syst.2
2010 Time for Our Field to Grow Up
abstract
Compared to centuries of physics and millennia of mathematics, the 50-year-history of computer science and information management research makes us the toddlers of the scientific community. Yet during our brief existence, we've revolutionized the world and, not content with that, gone on to build and study virtual worlds. We have justly taken pride in our accomplishments, and developed our own unique way of conducting research, unlike other scientific and engineering fields. But cracks have appeared in this edifice we have built. The conference system that served us so well for our first 50 years is falling apart. Our ever-increasing population competes ever more energetically for a finite set of resources. Other scientific and engineering disciplines still think that our field equates to programming, and look down on us. While we may also look down on them, it is undeniably true that high-energy physicists get many more research dollars per capita than we do, and our computer science colleagues wonder whether all the data management problems haven't already been solved. Other departments have started to teach courses that overlap our turf. Are we our own worst enemies? Why doesn't everyone understand how important our research is? Do we have to abandon the conference system? Must we become more like the stodgy old fields of science and engineering? Or can we find our own way?
Anastasia Ailamaki, Laura M. Haas, H. V. Jagadish, David Maier 0001, M. Tamer Özsu, Marianne Winslett
Proc. VLDB Endow.2
2010 SECRET: A Model for Analysis of the Execution Semantics of Stream Processing Systems
abstract
There are many academic and commercial stream processing engines (SPEs) today, each of them with its own execution semantics. This variation may lead to seemingly inexplicable differences in query results. In this paper, we present SECRET, a model of the behavior of SPEs. SECRET is a descriptive model that allows users to analyze the behavior of systems and understand the results of window-based queries for a broad range of heterogeneous SPEs. The model is the result of extensive analysis and experimentation with several commercial and academic engines. In the paper, we describe the types of heterogeneity found in existing engines, and show with experiments on real systems that our model can explain the key differences in windowing behavior.
Irina Botan, Roozbeh Derakhshan, Nihal Dindar, Laura M. Haas, Renée J. Miller, Nesime Tatbul
Proc. VLDB Endow.4
2010 TRAMP: Understanding the Behavior of Schema Mappings through Provenance
abstract
Though partially automated, developing schema mappings remains a complex and potentially error-prone task. In this paper, we present TRAMP (TRAnsformation Mapping Provenance), an extensive suite of tools supporting the debugging and tracing of schema mappings and transformation queries. TRAMP combines and extends data provenance with two novel notions, transformation provenance and mapping provenance, to explain the relationship between transformed data and those transformations and mappings that produced that data. In addition we provide query support for transformations, data, and all forms of provenance. We formally define transformation and mapping provenance, present an efficient implementation of both forms of provenance, and evaluate the resulting system through extensive experiments.
Boris Glavic, Gustavo Alonso, Renée J. Miller, Laura M. Haas
Proc. VLDB Endow.4
2010 Just-in-time Data Integration in Action
abstract
Today's data integration systems must be flexible enough to support the typical iterative and incremental process of integration, and may need to scale to hundreds of data sources. In this work we present a novel data integration system that offers great flexibility and scalability. Our approach to data integration is unique in that it executes mapping rules at query runtime using annotations. On top, we have built the People People People application. It allows users to search for people, display information about people, and browse through a network of related people, where the data is integrated from local and remote data sources. The demo presents all features of our underlying data integration engine through a set of motivating scenarios.
Martin Hentschel 0001, Laura M. Haas, Renée J. Miller
Proc. VLDB Endow.2
2009 New Challenges in Information Integration
Laura M. Haas, Aya Soffer
DaWaK1
2009 Schema AND Data: A Holistic Approach to Mapping, Resolution and Fusion in Information Integration
Laura M. Haas, Martin Hentschel 0001, Donald Kossmann, Renée J. Miller
ER1
2008 Impact! The Challenge of Industrial Research in Computer Science in a web 2.0 world
Laura M. Haas
SEKE1
2007 Information for People
abstract
Ordinary people have access to unprecedented volumes of information today. Researchers in the fields of information management (IM) and human-computer interaction (HCI) are reacting to this challenge from their own unique perspectives. Having access to a billion records is cool, but having access to a billion people is awesome. In this paper, we look at recent research from both communities, and speculate on how interactions between the communities could enhance the user experience of information.
Laura M. Haas, Steve B. Cousins
ICDE1
2007 Beauty and the Beast: The Theory and Practice of Information Integration
Laura M. Haas
ICDT1
2007 Special issue: best papers of VLDB 2005
Laura M. Haas, Christian S. Jensen, Martin L. Kersten
VLDB J.1
2006 Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality?
Dina Bitton, Franz Färber, Laura M. Haas, Jayavel Shanmugasundaram
VLDB3
2005 Clio grows up: from research prototype to industrial tool
abstract
Clio, the IBM Research system for expressing declarative schema mappings, has progressed in the past few years from a research prototype into a technology that is behind some of IBM's mapping technology. Clio provides a declarative way of specifying schema mappings between either XML or relational schemas. Mappings are compiled into an abstract query graph representation that captures the transformation semantics of the mappings. The query graph can then be serialized into different query languages, depending on the kind of schemas and systems involved in the mapping. Clio currently produces XQuery, XSLT, SQL, and SQL/XML queries. In this paper, we revisit the architecture and algorithms behind Clio. We then discuss some implementation issues, optimizations needed for scalability, and general lessons learned in the road towards creating an industrial-strength tool.
Laura M. Haas, Mauricio A. Hernández, C. T. Howard Ho, Lucian Popa 0001, Mary Roth
SIGMOD Conference1
2002 Attribute Classification Using Feature Analysis
abstract
The basis of many systems that integrate data from multiple sources is a set of correspondences between source schemata and a target schema. Correspondences express a relationship between sets of source attributes, possibly from multiple sources, and a set of target attributes. Clio is an integration tool that assists users in defining value correspondences between attributes. In real life scenarios there may be many sources and the source relations may have many attributes. Users can get lost and might miss or be unable to find some correspondences. Also, in many real life schemata the attribute names reveal little or nothing about the semantics of the data values. Only the data values in the attribute columns can convey the semantic meaning of the attribute. Our work relieves users of the problems of too many attributes and meaningless attribute names, by automatically suggesting correspondences between source and target attributes. For each attribute, we analyze the data values and derive a set of features.
Felix Naumann, C. T. Howard Ho, Xuqing Tian, Laura M. Haas, Nimrod Megiddo
ICDE4
2002 Garlic: a new flavor of federated query processing for DB2
abstract
In a large modern enterprise, information is almost inevitably distributed among several database management systems. Despite considerable attention from the research community, relatively few commercial systems have attempted to address this issue. This paper describes new technology that enables clients of IBM's DB2 Universal Database to access the data and specialized computational capabilities of a wide range of non-relational data sources. This technology, based on the Garlic prototype developed at the Almaden Research Center, complements and extends DB2's existing ability to federate relational data sources.The paper focuses on three topics. Firstly, we show how the DB2 catalogs are used as an extensible repository for the metadata needed to access remotely-stored information. Secondly, we describe how the Garlic approach to query planning, in which source-specific modules and the federated server cooperate to develop an optimized execution plan, has been realized in DB2. Lastly, we describe how DB2's query execution engine has been extended to support queries and functions that are evaluated remotely.
Vanja Josifovski, Peter M. Schwarz, Laura M. Haas, Eileen Tien Lin
SIGMOD Conference3
2001 Clio: A Semi-Automatic Tool For Schema Mapping
abstract
We consider the integration requirements of modern data intensive applications including data warehousing, global information systems and electronic commerce. At the heart of these requirements lies the schema mapping problem in which a source (legacy) database must be mapped into a different, but xed, target schema. The goal of schema mapping is the discovery of a query or set of queries to map source databases into the new structure. We demonstrate Clio, a new semi-automated tool for creating schema mappings. Clio employs a mapping-by-example paradigm that relies on the use of value correspondences describing how a value of a target attribute can be created from a set of values of source attributes. A typical session with Clio starts with the user loading a source and a target schema into the system. These schemas are read from either an underlying Object-Relational database or from an XML le with an associated XML Schema. Users can then draw value correspondences mapping source attributes into target attributes. Clio's mapping engine incrementally produces the SQL queries that realize the mappings implied by the correspondences. Clio provides schema and data browsers and other feedback to allow users to understand the mapping produced. Entering and manipulating value correspondences can be done in two modes. In the Schema View mode, users see a representation of the source and target schema and create value correspondences by selecting schema objects from the source and mapping them to a target attribute. The alternative Data View mode o ers a WYSIWYG interface for the mapping process that displays example data for both the source and target tables [3]. Users may add and delete value correspondences from this view and immediately see the changes re ected in the resulting target tuples. Also, the Data View mode helps users navigate through alternative mappings, understanding the often subtle di erences between them. For example, in some cases, changing a join from an inner join to an outer join may dramatically change the resulting table. In other cases, the same change may have no e ect due to constraints that hold on the source
Mauricio A. Hernández, Renée J. Miller, Laura M. Haas
SIGMOD Conference3
2001 Data-Driven Understanding and Refinement of Schema Mappings
abstract
At the heart of many data-intensive applications is the problem of quickly and accurately transforming data into a new form. Database researchers have long advocated the use of declarative queries for this process. Yet tools for creating, managing and understanding the complex queries necessary for data transformation are still too primitive to permit widespread adoption of this approach. We present a new framework that uses data examples as the basis for understanding and refining declarative schema mappings. We identify a small set of intuitive operators for manipulating examples. These operators permit a user to follow and refine an example by walking through a data source. We show that our operators are powerful enough both to identify a large class of schema mappings and to distinguish effectively between alternative schema mappings. These operators permit a user to quickly and intuitively build and refine complex data transformation queries that map one data source into another.
Ling-Ling Yan, Renée J. Miller, Laura M. Haas, Ronald Fagin
SIGMOD Conference3
2000 Integrating Life Sciences Data - With a Little Garlic
abstract
Vast amounts of life sciences data today reside in specialized data sources, with specialized query processing capabilities. Data from one source must often be combined with data from other sources to give users the information they desire. Database middleware systems such as Garlic allow users to combine data from multiple sources in a single query. Garlic provides the user with a virtual database to which they can pose arbitrarily complex queries, though the actual data needed to answer the query may be stored in several different sources, and those sources may not even possess all the functionality needed to answer such a query themselves. The Garlic technology, as incorporated in IBM's DB2 product, forms the basis of the DiscoveryLink service offering for the life sciences industry. We describe the DiscoveryLink offering, focusing on two key contributions of Garlic, the wrapper architecture and the query optimizer, and illustrate how it can be used to integrate life sciences data from heterogeneous data sources.
Laura M. Haas, Prasad Kodali, Julia E. Rice, Peter M. Schwarz, William C. Swope
BIBE1
2000 Panel: Is Generic Metadata Management Feasible?
Philip A. Bernstein, Laura M. Haas, Matthias Jarke, Erhard Rahm, Gio Wiederhold
VLDB2
2000 Schema Mapping as Query Discovery
Renée J. Miller, Laura M. Haas, Mauricio A. Hernández
VLDB2
1999 Using Fagin's Algorithm for Merging Ranked Results in Multimedia Middleware
abstract
A distributed multimedia information system allows users to access data of different modalities, from different data sources, ranked by various combinations of criteria. Fagin (1996) gives an algorithm for efficiently merging multiple ordered streams of ranked results, to form a new stream ordered by a combination of those ranks. In this paper we describe the implementation of Fagin's algorithm in an actual multimedia middleware system, including a novel, incremental version of the algorithm that supports dynamic exploration of data. We show that the algorithm would perform well as part of a single multimedia server and can even be effective in the distributed environment (for a limited set of queries), but that the assumptions it makes about random access limit its applicability dramatically. Our experience provides a better understanding of an important algorithm, and exposes an open problem for distributed multimedia information systems.
Edward L. Wimmers, Laura M. Haas, Mary Roth, Christoph Braendli
CoopIS2
1999 Loading a Cache with Query Results
Laura M. Haas, Donald Kossmann, Ioana Ursu
VLDB1
1999 Cost Models DO Matter: Providing Cost Information for Diverse Data Sources in a Federated System
Mary Roth, Fatma Özcan 0001, Laura M. Haas
VLDB3
1998 Capabilities-Based Query Rewriting in Mediator Systems
Yannis Papakonstantinou, Ashish Gupta 0001, Laura M. Haas
Distributed Parallel Databases3
1997 Optimizing Queries Across Diverse Data Sources
Laura M. Haas, Donald Kossmann, Edward L. Wimmers, Jun Yang 0001
VLDB1
1997 Data Structures for Efficient Broker Implementation
abstract
With the profusion of text databases on the Internet, it is becoming increasingly hard to find the most useful databases for a given query. To attack this problem, several existing and proposed systems employ brokers to direct user queries, using a local database of summary information about the available databases. This summary information must effectively distinguish relevant databases and must be compact while allowing efficient access. We offer evidence that one broker, GlOSS , can be effective at locating databases of interest even in a system of hundreds of databased and can examine the performance of accessing the GlOSS summeries for two promising storage methods: the grid file and partitioned hashing. We show that both methods can be tuned to provide good performance for a particular workload (within a broad range of workloads), and we discuss the tradeoffs between the two data structures. As a side effect of our work, we show that grid files are more broadly applicable than previously thought; inparticular, we show that by varying the policies used to construct the grid file we can provide good performance for a wide range of workloads even when storing highly skewed data.
Anthony Tomasic, Luis Gravano, Calvin Lue, Peter M. Schwarz, Laura M. Haas
ACM Trans. Inf. Syst.5
1997 Seeking the Truth About ad hoc Join Costs
Laura M. Haas, Michael J. Carey 0001, Miron Livny, Amit Shukla 0001
VLDB J.1
1996 The Garlic Project
abstract
The goal of the Garlic [1] project is to build a multimedia information system capable of integrating data that resides in different database systems as well as in a variety of non-database data servers. This integration must be enabled while maintaining the independence of the data servers, and without creating copies of their data. "Multimedia" should be interpreted broadly to mean not only images, video, and audio, but also text and application specific data types (e.g., CAD drawings, medical objects, …). Since much of this data is naturally modeled by objects, Garlic provides an object-oriented schema to applications, interprets object queries, creates execution plans for sending pieces of queries to the appropriate data servers, and assembles query results for delivery back to the applications. A significant focus of the project is support for "intelligent" data servers, i.e., servers that provide media-specific indexing and query capabilities [2]. Database optimization technology is being extended to deal with heterogeneous collections of data servers so that efficient data access plans can be employed for multi-repository queries.A prototype of the Garlic system has been operational since January 1995. Queries are expressed in an SQL-like query language that has been extended to include object-oriented features such as reference-valued attributes and nested sets. In addition to a C++ API, Garlic supports a novel query/browser interface called PESTO [3]. This component of Garlic provides end users of the system with a friendly, graphical interface that supports interactive browsing, navigation, and querying of the contents of Garlic databases. Unlike existing interfaces to databases, PESTO allows users to move back and forth seamlessly between querying and browsing activities, using queries to identify interesting subsets of the database, browsing the subset, querying the content of a set-valued attribute of a particularly interesting object in the subset, and so on.
Mary Roth, Manish Arya, Laura M. Haas, Michael J. Carey 0001, William F. Cody, Ronald Fagin, Peter M. Schwarz, Joachim Thomas 0002, Edward L. Wimmers
SIGMOD Conference3
1996 PESTO : An Integrated Query/Browser for Object Databases
Michael J. Carey 0001, Laura M. Haas, Vivekananda Maganty, John H. Williams
VLDB2
1993 Tapes Hold Data, Too: Challenges of Tuples on Tertiary Store
abstract
Article Free Access Share on Tapes hold data, too: challenges of tuples on tertiary store Authors: Michael J. Carey Computer Science Dept., University of Wisconsin, Madison, WI Computer Science Dept., University of Wisconsin, Madison, WIView Profile , Laura M. Haas IBM Almaden Research Center, K55/801, San Jose, CA IBM Almaden Research Center, K55/801, San Jose, CAView Profile , Miron Livny Computer Science Dept., University of Wisconsin, Madison, WI Computer Science Dept., University of Wisconsin, Madison, WIView Profile Authors Info & Claims SIGMOD '93: Proceedings of the 1993 ACM SIGMOD international conference on Management of dataJune 1993Pages 413–417https://doi.org/10.1145/170035.170103Published:01 June 1993Publication History 30citation211DownloadsMetricsTotal Citations30Total Downloads211Last 12 Months55Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Michael J. Carey 0001, Laura M. Haas, Miron Livny
SIGMOD Conference2
1990 Starburst Mid-Flight: As the Dust Clears
abstract
The purpose of the Starburst project is to improve the design of relational database management systems and enhance their performance, while building an extensible system to better support nontraditional applications and to serve as a testbed for future improvements in database technology. The design and implementation of the Starburst system to date are considered. Some key design decisions and how they affect the goal of improved structure and performance are examined. How well the goal of extensibility has been met is examined: what aspects of the system are extensible, how extensions can be done, and how easy it is to add extensions. Some actual extensions to the system, including the experiences of the first real customizers, are discussed.>
Laura M. Haas, Walter Chang, Guy M. Lohman, John McPherson, Paul F. Wilms, George Lapis, Bruce G. Lindsay 0001, Hamid Pirahesh, Michael J. Carey 0001, Eugene J. Shekita
IEEE Trans. Knowl. Data Eng.1
1989 Extensible Query Processing in Starburst
abstract
Today's DBMSs are unable to support the increasing demands of the various applications that would like to use a DBMS. Each kind of application poses new requirements for the DBMS. The Starburst project at IBM's Almaden Research Center aims to extend relational DBMS technology to bridge this gap between applications and the DBMS. While providing a full function relational system to enable sharing across applications, Starburst will also allow (sophisticated) programmers to add many kinds of extensions to the base system's capabilities, including language extensions (e.g., new datatypes and operations), data management extensions (e.g., new access and storage methods) and internal processing extensions (e.g., new join methods and new query transformations). To support these features, the database query language processor must be very powerful and highly extensible. Starburst's language processor features a powerful query language, rule-based optimization and query rewrite, and an execution system based on an extended relational algebra. In this paper, we describe the design of Starburst's query language processor and discuss the ways in which the language processor can be extended to achieve Starburst's goals.
Laura M. Haas, Johann-Christoph Freytag, Guy M. Lohman, Hamid Pirahesh
SIGMOD Conference1
1988 Views and Security in Distributed Database Management Systems
Elisa Bertino, Laura M. Haas
EDBT2
1986 A Snapshot Differential Refresh Algorithm
abstract
This article presents an algorithm to refresh the contents of database snapshots. A database snapshot is a read-only table whose contents are extracted from other tables in the database. The snapshot contents can be periodically refreshed to reflect the current state of the database. Snapshots are useful in many applications as a cost effective substitute for replicated data in a distributed database system.
Bruce G. Lindsay 0001, Laura M. Haas, C. Mohan 0001, Hamid Pirahesh, Paul F. Wilms
SIGMOD Conference2
1984 Optimization of Nested Queries in a Distributed Relational Database
Guy M. Lohman, Dean Daniels, Laura M. Haas, Ruth Kistler, Patricia G. Selinger
VLDB3
1984 Computation and Communication in R*: A Distributed Database Manager
abstract
This article presents and discusses the computation and communication model used by R*, a prototype distributed database management system.An R* computation consists of a tree of processes connected by virtual circuit communication paths.The process management and communication protocols used by R* enable the system to provide reliable, distributed transactions while maintaining adequate levels of performance.Of particular interest is the use of processes in R* to retain user context from one transaction to another, in order to improve the system performance and recovery characteristics.
Bruce G. Lindsay 0001, Laura M. Haas, C. Mohan 0001, Paul F. Wilms, Robert A. Yost
ACM Trans. Comput. Syst.2
1983 Computation & Communication in R*: A Distributed Database Manager (Extended Abstract)
abstract
R* is an experimental prototype distributed database management system. The computation needed to perform a sequence of multisite user transactions in R* is structured as a tree of processes communicating over virtual circuit communication links. Distributed computation can be supported by providing a server process per site which performs requests on behalf of remote users. Alternatively, a new process could be created to service each incoming request. Instead of using a shared server process or using the process per request approach, R* creates a process associated with the computation of the user on the first request to the remote site. This process is incorporated into the tree of processes serving a single user and is retained for the duration of the user computation. This approach allows R* to factor some of the request execution overhead into the process creation phase, and simplifies the retention of user and transaction context at the multiple sites of the distributed computation.
Bruce G. Lindsay 0001, Laura M. Haas, C. Mohan 0001, Paul F. Wilms, Robert A. Yost
SOSP2
1983 View Management in Distributed Data Base Systems
Elisa Bertino, Laura M. Haas, Bruce G. Lindsay 0001
VLDB2
1983 Site autonomy issues in R*: A distributed database management system
Patricia G. Selinger, Dean Daniels, Laura M. Haas, Bruce G. Lindsay 0001, Pui Ng, Paul F. Wilms, Robert A. Yost
Inf. Sci.3
1983 Distributed Deadlock Detection
abstract
Distributed deadlock models are presented for resource and communication deadlocks.Simple distributed algorithms for detection of these deadlocks are given.We show that all true deadlocks are detected and that no false deadlocks are reported.In our algorithms, no process maintains global information; all messages have an identical short length.The algorithms can be applied in distributed database and other message communication systems.
K. Mani Chandy, Jayadev Misra, Laura M. Haas
ACM Trans. Comput. Syst.3