Hans-Jörg Schek

dblp:s/HansJorgSchek · DBLP profile ↗
← Back
64ranked-venue papers
13as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 55 · 12 first-authorArtificial intelligence and machine learning · 4Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
24 papers
Transaction processing and concurrency control · 45% Information retrieval · 21% Distributed and cloud data management · 12%
Computer architecture, parallel and distributed computing, and storage systems
9 papers
Distributed systems · 47% Parallel and multicore computing · 24% Electronic design automation · 19%
Software engineering, system software, and programming languages
2 papers
Services computing and microservices · 100%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transaction processing and concurrency control
correctness criteria
0.131999
Concurrency Control and Recovery in Transactional Process Management · PODS 1999
Correctness in General Configurations of Transactional Components · PODS 1999
Correctness and Parallelism of Composite Systems · PODS 1997
Distributed and cloud data management › distributed database architecture
database cluster
0.122001
Cache-Aware Query Routing in a Cluster of Databases · ICDE 2001
High-level Parallelism in a Database Cluster: A Feasibility Study Using Document Services · ICDE 2001
Transaction processing and concurrency control
concurrency control and recovery
0.122002
Atomicity and isolation for transactional processes · ACM Trans. Database Syst. 2002
Concurrency Control and Recovery in Transactional Process Management · PODS 1999
Information retrieval
similarity search
0.132001
Interactive-Time Similarity Search for Large Image Collections Using Parallel VA-Files · ICDE 2000
A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces · VLDB 1998
Fast Evaluation Techniques for Complex Similarity Queries · VLDB 2001
Parallel and multicore computing › task scheduling › task graph scheduling
duplication-based scheduling
0.112005
Fine-Grained Replication and Scheduling with Freshness and Correctness Guarantees · VLDB 2005
Distributed systems
fault tolerance
0.112005
Fine-Grained Replication and Scheduling with Freshness and Correctness Guarantees · VLDB 2005
Distributed systems
replication
0.112005
Fine-Grained Replication and Scheduling with Freshness and Correctness Guarantees · VLDB 2005
Electronic design automation › high-level synthesis
scheduling
0.112005
Fine-Grained Replication and Scheduling with Freshness and Correctness Guarantees · VLDB 2005
Services computing and microservices › service composition
web service composition
0.012003
WebService Composition with O'GRAPE and OSIRIS · VLDB 2003
Information retrieval › distributed information retrieval
query routing
0.012001
Cache-Aware Query Routing in a Cluster of Databases · ICDE 2001
Query processing and optimization
similarity query processing
0.012001
Fast Evaluation Techniques for Complex Similarity Queries · VLDB 2001
Information retrieval
indexing
0.012000
Interactive-Time Similarity Search for Large Image Collections Using Parallel VA-Files · ICDE 2000
Indexing and storage engines › vector index
VA-file
0.012000
Interactive-Time Similarity Search for Large Image Collections Using Parallel VA-Files · ICDE 2000
Transaction processing and concurrency control › nested transactions
multilevel transactions
0.031994
Semantics-Based Multilevel Transaction Management in Federated Systems · ICDE 1994
The DASDBS Project: Objectives, Experiences, and Future Prospects · IEEE Trans. Knowl. Data Eng. 1990
Architecture and Implementation of the Darmstadt Database Kernel System · SIGMOD Conference 1987
Information retrieval
retrieval models
0.021998
A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces · VLDB 1998
Data Structures for an Integrated Data Base Management and Information Retrieval System · VLDB 1982
Transaction processing and concurrency control
concurrency control
0.031999
Towards a Unified Theory of Concurrency Control and Recovery · PODS 1993
Correctness in General Configurations of Transactional Components · PODS 1999
Semantics-Based Multilevel Transaction Management in Federated Systems · ICDE 1994
Information retrieval › similarity search
high-dimensional similarity search
0.011998
A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces · VLDB 1998
Distributed and cloud data management
distributed query processing
0.011997
Distributed Processing over Stand-alone Systems and Applications · VLDB 1997
Distributed systems
distributed coordination
0.022002
FAS - A Freshness-Sensitive Coordination Middleware for a Cluster of OLAP Components · VLDB 2002
Distributed Processing over Stand-alone Systems and Applications · VLDB 1997
Transaction processing and concurrency control › transaction execution
intratransaction parallelism
0.011996
Intra-Transaction Parallelism in the Mapping of an Object Model to a Relational Multi-Processor System · VLDB 1996
Data models and query languages › relational model
nested relational model
0.031990
The DASDBS Project: Objectives, Experiences, and Future Prospects · IEEE Trans. Knowl. Data Eng. 1990
Supporting Flat Relations by a Nested Relational Kernel · VLDB 1987
Remarks on the Algebra of Non First Normal Form Relations · PODS 1982
Distributed and cloud data management
federated database
0.011994
Semantics-Based Multilevel Transaction Management in Federated Systems · ICDE 1994
Transaction processing and concurrency control › distributed transaction processing
global transaction scheduling
0.011994
Semantics-Based Multilevel Transaction Management in Federated Systems · ICDE 1994
Transaction processing and concurrency control
recovery
0.011993
Towards a Unified Theory of Concurrency Control and Recovery · PODS 1993
Transaction processing and concurrency control
serializability
0.011993
Towards a Unified Theory of Concurrency Control and Recovery · PODS 1993
Parallel and multicore computing › parallel algorithms
parallel search
0.012000
Interactive-Time Similarity Search for Large Image Collections Using Parallel VA-Files · ICDE 2000
Performance modeling and evaluation
benchmarking
0.011998
A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces · VLDB 1998
Performance modeling and evaluation
quantitative evaluation
0.011998
A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces · VLDB 1998
Data integration and cleaning › schema mapping
object-relational mapping
0.011996
Intra-Transaction Parallelism in the Mapping of an Object Model to a Relational Multi-Processor System · VLDB 1996
Data models and query languages › semistructured data
semi-structured data management
0.011996
An Open Storage System for Abstract Objects · SIGMOD Conference 1996

Methods — techniques the papers use, named apart from their topics

middleware · 0.1transaction processing monitor · 0.1experiment · 0.1parallel VA-files · 0.1proof technique for arbitrary configurations · 0.0unified process model · 0.0transaction scheduling model · 0.0signature approximations · 0.0unified theory of concurrency control and recovery · 0.0multilevel nested transaction model · 0.0
YearPublicationVenuePosition
2010 Flexible Data Access in a Cloud Based on Freshness Requirements
abstract
Data clouds are newly emerging environments in which commercial providers manage large volumes of data with individual quality of service (QoS) guarantees per customer. These guarantees mainly include keeping several replicas of each data item in different distributed data centers for availability purposes. However, as the cost of maintaining several updateable replicas per data object is very high, cloud providers rather offer only a limited number of synchronously updated replicas (i.e., replicas that are always up-to-date) together with several read-only replicas that are updated in a lazy way and thus might hold stale data. QoS agreements may also include the maintenance of dedicated archives (copies of data which are frozen at some point in time). Stale data allow cloud providers to offer a variety of read operations with different semantics, e.g., read the most recent data, read data not older than / not younger than some timestamp t, or read data produced between t1and t2, or read data exactly as of t. These read operations can be supported by a read-only site using a stale replica. In this paper we present our approach to cloud data management, based on a recent protocol for data grids. We discuss in detail how the refresh of individual replicas is provided in a completely distributed way. Finally, we present the results of a performance evaluation in a data cloud setting.
Laura Cristiana Voicu, Heiko Schuldt, Yuri Breitbart, Hans-Jörg Schek
IEEE CLOUD4
2008 Toward replication in grids for digital libraries with freshness and correctness guarantees
abstract
Abstract Building digital libraries (DLs) on top of data grids while facilitating data access and minimizing access overheads is challenging. To achieve this, replication in a Grid has to provide dedicated features that are only partly supported by existing Grid environments. First, it must provide transparent and consistent access to distributed data. Second, it must dynamically control the creation and maintenance of replicas. Third, it should allow higher replication granularities, i.e. beyond individual files. Fourth, users should be able to specify their freshness demands, i.e. whether they need most recent data or are satisfied with slightly outdated data. Finally, all these tasks must be performed efficiently. This paper presents an approach that will finally allow one to build a fully integrated and self‐managing replication subsystem for data grids that will provide all the above features. Our approach is to start with an accepted replication protocol for database clusters, namely PDBREP, and to adapt it to the grid. Copyright © 2008 John Wiley & Sons, Ltd.
Fuat Akal, Heiko Schuldt, Hans-Jörg Schek
Concurr. Comput. Pract. Exp.3
2008 Introduction to the special issue on database and information retrieval integration
W. Bruce Croft, Hans-Jörg Schek
VLDB J.2
2006 Efficient and Coordinated Checkpointing for Reliable Distributed Data Stream Management
Gert Brettlecker, Heiko Schuldt, Hans-Jörg Schek
ADBIS3
2005 How can we support Grid Transactions? Towards Peer-to-Peer Transaction Processing
Can Türker, Klaus Haller, Christoph Schuler, Hans-Jörg Schek
CIDR4
2005 The Hyperdatabase Network - New Middleware for Searching and Maintaining the Information Space
Hans-Jörg Schek
SOFSEM1
2005 Fine-Grained Replication and Scheduling with Freshness and Correctness Guarantees
Fuat Akal, Can Türker, Hans-Jörg Schek, Yuri Breitbart, Torsten Grabs, Lourens E. Veen
VLDB3
2005 Peer-to-peer Execution of (transactional) Processes
abstract
Standards like SOAP, WSDL, and UDDI facilitate the proliferation of services. Based on these technologies, processes are a means to combine services to applications and to provide new value-added services. For large information systems, a centralized process engine is no longer appropriate due to limited scalability. Instead, in this paper, we propose a distributed and decentralized process engine that routes process instances directly from one peer to the next. Such a peer-to-peer process execution promises good scalability characteristics since it is able to dynamically balance the load of processes and services among all available service providers. Therefore, navigation costs only accumulate on peers that are directly involved in the execution. However, this requires sophisticated strategies for the replication of meta-data for peer-to-peer process execution. Especially, replication mechanisms should avoid frequent accesses to global information repositories. In our system, called OSIRIS (Open Service Infrastructure for Reliable and Integrated Process Support), we deploy a publish/subscribe-based replication scheme together with freshness predicates to significantly reduce replication costs. This way, OSIRIS can support process-based applications in a dynamically evolving system without limiting scalability and correctness. Experiments have shown very promising results with respect to scalability. In addition, OSIRIS provides a flexible infrastructure that can be extended seamlessly in a modular way. This paper demonstrates the extension towards distributed concurrency control.
Christoph Schuler, Heiko Schuldt, Can Türker, Roger Weber, Hans-Jörg Schek
Int. J. Cooperative Inf. Syst.5
2004 Scalable Peer-to-Peer Process Management - The OSIRIS Approach
abstract
The functionality of applications is increasingly being made available by services. General concepts and standards like SOAP, WSDL, and UDDI support the discovery and invocation of single Web services. State-of-the-art process management is conceptually based on a centralized process manager. The resources of this coordinator limit the number of concurrent process executions, especially since the coordinator has to persistently store each state change for recovery purposes. In this paper, we overcome this limitation by executing processes in a peer-to-peer way exploiting all nodes of the system. By distributing the execution and navigation costs, we can achieve a higher degree of scalability allowing for a much larger throughput of processes compared to centralized solutions. This paper describes our prototype system OSIRIS, which implements such a true peer-to-peer process execution. We further present very promising results verifying the advantages over centralized process management in terms of scalability.
Christoph Schuler, Roger Weber, Heiko Schuldt, Hans-Jörg Schek
ICWS4
2004 PowerDB-IR - Scalable Information Retrieval and Storage with a Cluster of Databases
Torsten Grabs, Klemens Böhm, Hans-Jörg Schek
Knowl. Inf. Syst.3
2003 Peer-to-Peer Process Execution with Osiris
Christoph Schuler, Roger Weber, Heiko Schuldt, Hans-Jörg Schek
ICSOC4
2003 Transactional Peer-to-Peer Information Processing: The AMOR Approach
Klaus Haller, Heiko Schuldt, Hans-Jörg Schek
Mobile Data Management3
2003 The Lowell Report
abstract
No abstract available.
Jim Gray 0001, Hans-Jörg Schek, Michael Stonebraker, Jeffrey D. Ullman
SIGMOD Conference2
2003 WebService Composition with O'GRAPE and OSIRIS
Roger Weber, Christoph Schuler, Patrick Neukomm, Heiko Schuldt, Hans-Jörg Schek
VLDB5
2002 OLAP Query Evaluation in a Database Cluster: A Performance Study on Intra-Query Parallelism
Fuat Akal, Klemens Böhm, Hans-Jörg Schek
ADBIS3
2002 Infrastructure for Information Spaces
Hans-Jörg Schek, Heiko Schuldt, Christoph Schuler, Roger Weber
ADBIS1
2002 XMLTM: efficient transaction management for XML documents
abstract
A common approach to storage and retrieval of XML documents is to store them in a database, together with materialized views on their content. The advantage over "native" XML storage managers seems to be that transactions and concurrency are for free, next to other benefits. But a closer look and preliminary experiments reveal that this results in poor performance of concurrent queries and updates. The reason is that database lock contention hinders parallelism unnecessarily. We therefore investigate concurrency control at the semantic, i.e., XML level and describe a respective transaction manager XMLTM. It features a new locking protocol DGLOCK. It generalizes the protocol for locking on directed acyclic graphs by adding simple predicate locking on the content of elements, e.g., on their text. Instead of using the original XML documents, we propose to take advantage of an abstraction of the XML document collection known as DataGuides. XMLTM allows to run XML processing at the underlying database at low ANSI isolation degrees and to release database locks early without sacrificing correctness in this setting. We have built a complete prototype system that is implemented on top of the XML Extender for IBM DB2. Our evaluation shows that our approach consistently yields performance improvements by an order of magnitude. We stress that our approach can also be implemented within a native XML storage manager, and we expect even better performance.
Torsten Grabs, Klemens Böhm, Hans-Jörg Schek
CIKM3
2002 Hyperdatabases: Infrastructure for the Information Space
Hans-Jörg Schek
EDBT1
2002 FAS - A Freshness-Sensitive Coordination Middleware for a Cluster of OLAP Components
Uwe Röhm, Klemens Böhm, Hans-Jörg Schek, Heiko Schuldt
VLDB3
2002 Atomicity and isolation for transactional processes
abstract
Processes are increasingly being used to make complex application logic explicit. Programming using processes has significant advantages but it poses a difficult problem from the system point of view in that the interactions between processes cannot be controlled using conventional techniques. In terms of recovery, the steps of a process are different from operations within a transaction. Each one has its own termination semantics and there are dependencies among the different steps. Regarding concurrency control, the flow of control of a process is more complex than in a flat transaction. A process may, for example, partially roll back its execution or may follow one of several alternatives. In this article, we deal with the problem of atomicity and isolation in the context of processes. We propose a unified model for concurrency control and recovery for processes and show how this model can be implemented in practice, thereby providing a complete framework for developing middleware applications using processes.
Heiko Schuldt, Gustavo Alonso, Catriel Beeri, Hans-Jörg Schek
ACM Trans. Database Syst.4
2001 PowerDB-IR - Information Retrieval on Top of a Database Cluster
abstract
Our current concern is a scalable infrastructure for information retrieval (IR) with up-to-date retrieval results in the presence of frequent, continuous updates. Timely processing of updates is important with novel application domains, e.g., e-commerce. We want to use off-the-self hardware and software as much as possible. These issues are challenging, given the additional requirement that the resulting system must scale well. We have built PowerDB-IR, a system that has the characteristics sought. This paper describes its design, implementation, and evaluation. PowerDB-IR is a coordination layer for a database cluster. The rationale behind a database cluster is to 'scale-out', i.e., to add further cluster nodes, whenever necessary for better performance. We build on IR-to-database mappings and service decomposition to support high-level parallelism. We follow a three-tier architecture with the database cluster as the bottom layer for storage management. The middle tier provides IR-specific processing and update services. PowerDB-IR has the following features: It allows to insert and retrieve documents concurrently, and it ensures freshness with almost no overhead. Alternative physical data organization schemes provide adequate performance for different workloads. Query processing techniques for the different data organizations efficiently integrate the ranked retrieval results from the cluster nodes. We have run extensive experiments with our prototype using commercial database systems and middleware software products. The main result is that PowerDB-IR shows surprisingly ideal scalability and low response times.
Torsten Grabs, Klemens Böhm, Hans-Jörg Schek
CIKM3
2001 High-level Parallelism in a Database Cluster: A Feasibility Study Using Document Services
abstract
Our concern is the design of a scalable infrastructure for complex application services. We want to find out if a cluster of commodity database systems is well-suited as such an infrastructure. To this end, we have carried out a feasibility study based on document services, e.g. document insertion and retrieval. We decompose a service request into short parallel database transactions. Our system, implemented as an extension of a transaction processing monitor, routes the short transactions to the appropriate database systems in the cluster. Routing depends on the data distribution that we have chosen. To avoid bottlenecks, we distribute document functionality, such as term extraction, over the cluster. Extensive experiments show the following. (1) A relatively small number of components - for example eight components $already suffices to cope with high workloads of more than 100 concurrently active clients. (2) Speedup and throughput increase linearly for insertion operations when increasing the cluster size. These observations also hold when bundling service invocations into transactions at the semantic layer. A specialized coordinator component then implements semantic serializability and atomicity. Our experiments show that such a coordinator has minimal impact on CPU resource consumption and on response times.
Torsten Grabs, Klemens Böhm, Hans-Jörg Schek
ICDE3
2001 Cache-Aware Query Routing in a Cluster of Databases
abstract
We investigate query routing techniques in a cluster of databases for a query-dominant environment. The objective is to decrease query response time. Each component of the cluster runs an off-the-shelf DBMS and holds a copy of the whole database. The cluster has a coordinator that routes each query to an appropriate component. Considering queries of realistic complexity, e.g., TPC-R, this article addresses the following questions: Can routing benefit from caching effects due to previous queries? Since our components are black-boxes, how can we approximate their cache content? How to route a query, given such cache approximations? To answer these questions, we have developed a cache-aware query router that is based on signature approximations of queries. We report on experimental evaluations with the TPC-R benchmark using our PowerDB database cluster prototype. Our main result is that our approach of cache approximation routing is better than state-of-the-art strategies by a factor of two with regard to mean response time.
Uwe Röhm, Klemens Böhm, Hans-Jörg Schek
ICDE3
2001 Evolution of Database Technology: Hyperdatabases
abstract
Our vision is that hyperdatabases become available that extend and evolve from database technology. Hyperdatabases move up to a higher level, closer to the applications. A hyperdatabase manages distributed objects and software components as well as workflows, in analogy to a database system that manages data and transactions. In short, hyperdatabases will provide "higher order data independence", e.g., immunity of applications against changes in the implementation of components and workload transparency. Such an evolution of database technology should keep its pivotal role as infrastructure for application development for data-intensive, central and distributed application. The hyperdatabase concept abstracts from the host of current infrastructures and middleware technology.
Hans-Jörg Schek
IDEAS1
2001 Fast Evaluation Techniques for Complex Similarity Queries
Klemens Böhm, Michael Mlivoncic, Hans-Jörg Schek, Roger Weber
VLDB3
2001 Special Issue on The 2nd Web Information Systems Engineering Conference (WISE'01)
M. Tamer Özsu, Hans-Jörg Schek, Katsumi Tanaka, Yanchun Zhang
World Wide Web2
2000 OLAP Query Routing and Physical Design in a Database Cluster
Uwe Röhm, Klemens Böhm, Hans-Jörg Schek
EDBT3
2000 Evaluating the Coordination Overhead of Replica Maintenance in a Cluster of Databases
Klemens Böhm, Torsten Grabs, Uwe Röhm, Hans-Jörg Schek
Euro-Par4
2000 Interactive-Time Similarity Search for Large Image Collections Using Parallel VA-Files
Roger Weber, Klemens Böhm, Hans-Jörg Schek
ICDE3
2000 Panel: Future Directions of Database Research - The VLDB Broadening Strategy, Part 1
Hans-Jörg Schek
VLDB1
2000 Hyperdatabases
abstract
When relational database systems were introduced twenty years ago (1980), they were an infrastructure and main platform for application development. With today's information systems, the database system is a storage manager, far away from the applications. Our vision is that hyperdatabases become available that move up and extend database concepts to a higher level, closer to the applications. A hyperdatabase manages distributed objects and software components as well as workflows, in analogy to a database system that manages data and transactions. In short, hyperdatabases, also called "higher order databases", will provide "higher order data independence", e.g., immunity of applications against changes in the implementation of components and workload transparency. They will be the infrastructure for distributed information systems engineering of the future, and they are an abstraction from the host of current infrastructures and middleware technology. The article elaborates on this vision and outlines concrete projects at ETHZ such as PowerDB, a database cluster project. It shows how an efficient document engine can be built on top of a database cluster. A further project studies transactional process management as a layer on top of database transactions. Image similarity and multimedia components is another project where a hyperdatabase coordinates specialized components such as feature extraction and indexing services in a distributed environment.
Hans-Jörg Schek, Klemens Böhm, Torsten Grabs, Uwe Röhm, Heiko Schuldt, Roger Weber
WISE1
2000 Automatic Genration of Reliable E-Commerce Payment Processes
abstract
The most important phase in e-commerce interactions is the payment, due to the transfer of sensitive information (e.g., credit card numbers). A couple of requirements exist both from the point of view of a customer and from the merchant's perspective. This set of requirements is even enlarged when complex interactions are considered in which a customer purchases goods originating from different merchants within one single e-commerce transaction. We show how all these different requirements of payment interactions can be seamlessly integrated in transactional payment processes. These processes are generated automatically based on the customer's specification of the e-commerce transaction (involved participants, means of payment, etc.). We present the basic structure of such payment processes, how the requirements are mapped into these processes and how they can be generated automatically. Furthermore, we present the architecture of a payment coordinator that has been implemented within the INVENT project. This payment coordinator controls the execution of transactional payment processes, thereby keeping track of the interactions with the various participants.
Heiko Schuldt, Andrei Popovici, Hans-Jörg Schek
WISE3
1999 Architecture of a Networked Image Search and Retrieval System
abstract
Large scale networked image retrieval systems face a number of problems that are not fully satisfied by current systems. On one hand, integrated solutions that store all image data centrally are often limited in terms of scalability and autonomy of data providers. On the other hand, WWW-based search engines proved to be fairly scalable, and data providers retain their autonomy. However, such engines often confront users with links to servers that are not available or to images that no longer exist, i.e., they are unable to keep their meta-database consistent with the repositories' contents. Furthermore, existing solutions often neglect the cost of image delivery. The considerable variations in the effective bandwidth in today's Internet lead to highly unpredictable response times, which are often intolerable from the user's point of view.
Roger Weber, Jürg Bolliger, Thomas R. Gross, Hans-Jörg Schek
CIKM4
1999 A Generalized Transaction Theory for Database and Non-database Tasks
Armin Fessler, Hans-Jörg Schek
Euro-Par2
1999 Transactions in Stack, Fork, and Join Composite Systems
Gustavo Alonso, Armin Fessler, Guy Pardon, Hans-Jörg Schek
ICDT4
1999 Transactional Coordination Agents for Composite Systems
abstract
Composite systems are collections of autonomous, heterogeneous and distributed software applications. In these systems, data dependencies are continuously violated by local operations, and therefore coordination processes are necessary to guarantee overall correctness and consistency. Such coordination processes must be endowed with some form of execution guarantees, which require the participating subsystems to have certain database functionality (such as atomicity of local operations, order preservation, and either compensation of operations or the deferment of their commit). However, this functionality is not present in many applications and must be implemented by a transactional coordination agent coupled with the application. In this paper, we discuss the requirements to be met by the applications and their associated transactional coordination agents. We identify a minimal set of functionalities which the applications must provide in order to participate in transactional coordination processes, and we also discuss how the missing database functionality can be added to arbitrary applications using transactional coordination agents. Then, we identify the structure of a generic transactional coordination agent and provide an implementation example of a transactional coordination agent tailored to SAP R/3.
Heiko Schuldt, Hans-Jörg Schek, Gustavo Alonso
IDEAS2
1999 Correctness in General Configurations of Transactional Components
abstract
From a transactional point of view, composite systems are component based applications in which each component has its own transaction management logic.These systems are highly relevant in practice since they are likely to be the standard architecture for many future distributed applications.Unfortunately,there is no appropriate conceptual framework in which to reason about such systems.Following up on existing work that addressed special cases of composite systems, in this paper we tackle the problem of general composite systems, i.e., those with arbitrary configurations.We propose a correctness criterion, develop a new proof technique that allows us to address arbitrary configurations, and discuss several important issues related to concurrency control in distributed systems.
Gustavo Alonso, Armin Fessler, Guy Pardon, Hans-Jörg Schek
PODS4
1999 Concurrency Control and Recovery in Transactional Process Management
abstract
The unified theory of concurrency control and recovery integrates atomicity and isolation within a common framework, thereby avoiding many of the shortcomings resulting from treating them as orthogonal problems.This theory can be applied to the traditional read/write model as well as to semantically rich operations.In this paper, we extend the unified theory by applying it to generalized process structures, i.e., arbitrary partially ordered sequences of transaction invocations.lJsing the extended unified theory, our goal is to provide a more flexible handling of concurrent processes while allowing: as much parallelism as possible.Unlike in the original unified theory, we take into account that not all activities of a process might be compensatable and the fact that these process structures require transactional properties more general than in traditional ACID transactions.We provide a correctness criterion for transactional processes and identity the key points in which the more flexible structure of transactional processes implies differences from traditional transactions.
Heiko Schuldt, Gustavo Alonso, Hans-Jörg Schek
PODS3
1998 A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces
Roger Weber, Hans-Jörg Schek, Stephen Blott
VLDB2
1998 Unifying Concurrency Control and Recovery of Transactions with Semantically Rich Operations
Radek Vingralek, Haiyan Hasse-Ye, Yuri Breitbart, Hans-Jörg Schek
Theor. Comput. Sci.4
1997 Correctness and Parallelism of Composite Systems
abstract
In recent years, databases have started to be used as intelligent repositories for a variety of semantically-richer systems. A consequence of such architectures is that transaction scheduling takes place throughout composite systems consisting of layered subsystems. Such transaction architectures have been studied extensively. Existing theory, however, limits the degree of parallelism, and makes a number of simplifying assumptions which cannot be taken for granted in practice. This paper proposes a new model and correctness criterion, stack conflict consistency, for composite transactional systems. The main contribution of the new model is to establish the correctness conditions under which higher degrees of parallelism can be achieved between operations of the same transaction, as well as between conflicting operations of different transactions, in a uniform way. This possibility, although hinted at previously, has not yet been exploited in practical composite systems. Hence, we hope...
Gustavo Alonso, Stephen Blott, Armin Fessler, Hans-Jörg Schek
PODS4
1997 Distributed Processing over Stand-alone Systems and Applications
Gustavo Alonso, Claus Hagen, Hans-Jörg Schek, Markus Tresch
VLDB3
1996 An Open Storage System for Abstract Objects
abstract
Database systems must become more open to retain their relevance as a technology of choice and necessity. Openness implies not only databases exporting their data, but also exporting their services. This is as true in classical application areas as in non-classical (GIS, multimedia, design, etc).This paper addresses the problem of exporting storage-management services of indexing, replication and basic query processing. We describe an abstract-object storage model which provides the basic mechanism, 'likeness', through which these services are applied uniformly to internally-stored, internally-defined data, and to externally-stored, externally-defined data. Managing external data requires the coupling of external operations to the database system. We discuss the interfaces and protocols required of these to achieve correct resource management and admit efficient realisation. Throughout, we demonstrate our solutions in the area of semi-structured file management; in our case, geospatial metadata files.
Stephen Blott, Lukas Relly, Hans-Jörg Schek
SIGMOD Conference3
1996 Intra-Transaction Parallelism in the Mapping of an Object Model to a Relational Multi-Processor System
Michael Rys, Moira C. Norrie, Hans-Jörg Schek
VLDB3
1995 Unified Transaction Model for Semantically Rich Operations
Radek Vingralek, Haiyan Ye, Yuri Breitbart, Hans-Jörg Schek
ICDT4
1994 A Unified Approach to Concurrency Control and Transaction Recovery (Extended Abstract)
Gustavo Alonso, Radek Vingralek, Divyakant Agrawal, Yuri Breitbart, Amr El Abbadi, Hans-Jörg Schek, Gerhard Weikum
EDBT6
1994 Semantics-Based Multilevel Transaction Management in Federated Systems
abstract
A federated database management system (FDBMS) is a special type of distributed database system that enables existing local databases, in a heterogeneous environment, to maintain a high degree of autonomy. One of the key problems in this setting is the coexistence of local transactions and global transactions, where the latter access and manipulate data of multiple local databases. In modeling FDBMS transaction executions the authors propose a more realistic model than the traditional read/write model; in their model a local database exports high-level operations which are the only operations distributed global transactions can execute to access data in the shared local databases. Such restrictions are not unusual in practice as, for example, no airline or bank would ever permit foreign users to execute ad hoc queries against their databases for fear of compromising autonomy. The proposed architecture can be elegantly modeled using the multilevel nested transaction model for which a sound theoretical foundation exists to prove concurrent executions correct. A multilevel scheduler that is able to exploit the semantics of exported operations can significantly increase concurrency by ignoring pseudo conflicts. A practical scheduling mechanism for FDBMSs is described that offers the potential for greater performance and more flexibility than previous approaches based on the read/write model.>
Andrew Deacon, Hans-Jörg Schek, Gerhard Weikum
ICDE2
1994 Unifying concurrency control and recovery of transactions
Gustavo Alonso, Radek Vingralek, Divyakant Agrawal, Yuri Breitbart, Amr El Abbadi, Hans-Jörg Schek, Gerhard Weikum
Inf. Syst.6
1993 Towards a Unified Theory of Concurrency Control and Recovery
abstract
The classical theory of transaction management is based on two different and independent criteria for the correct execution of transactions.The first criterion, serializability, ensures correct execution of parallel transactions under the assumption that no failures occur.The second criterion, strictness, ensures correct recovery from failures.In this paper we develop a unified model that allows reasoning about the correctness of concurrency control and recovery within the same framework.We introduce the correctness criteria of (prefix-) reducibility and (prefix-) expanded serializability and investigate their relationships to the classical criteria.An important advantage of our model is that it captures schedules with semantically rich ADT actions in addition to classical read/write schedules.
Hans-Jörg Schek, Gerhard Weikum, Haiyan Ye
PODS1
1992 Utilization of External Foreign Computation Services
Hans-Jörg Schek
ICDE1
1991 Interoperability In Multidatabases: Semantic and System Issues (Panel)
Yuri Breitbart, Hector Garcia-Molina, Witold Litwin, Nick Roussopoulos, Hans-Jörg Schek, Gio Wiederhold
VLDB5
1990 A Relational Object Model
Marc H. Scholl, Hans-Jörg Schek
ICDT2
1990 The DASDBS Project: Objectives, Experiences, and Future Prospects
abstract
A retrospective of the Darmstadt database system project, also known as DASDBS, is presented. The project is aimed at providing data management support for advanced applications, such as geo-scientific information systems and office automation. Similar to the dichotomy of RSS and RDS in System R, a layered architectural approach was pursued: a storage management kernel serves as the lowest common denominator of the requirements of the various applications classes, and a family of application-oriented front-ends provides semantically richer functions on top of the kernel. The lessons that were learned from building the DASDBS system are discussed. Particular emphasis is placed on the following issues: the role of nested relations, the experiences with using object buffers for coupling the system with the programming-language environment and the learning process in implementing multilevel transactions.>
Hans-Jörg Schek, Heinz-Bernhard Paul, Marc H. Scholl, Gerhard Weikum
IEEE Trans. Knowl. Data Eng.1
1989 A Signature Access Method for the Starburst Database System
Walter W. Chang, Hans-Jörg Schek
VLDB2
1989 A Frame-Based Knowledge Representation Model and its Mapping to Nested Relations
Ulrich Reimer, Hans-Jörg Schek
Data Knowl. Eng.2
1988 Multi-Level Transaction Management, Theoretical Art or Practical Need ?
Catriel Beeri, Hans-Jörg Schek, Gerhard Weikum
EDBT2
1987 Architecture and Implementation of the Darmstadt Database Kernel System
abstract
The multi-layered architecture of the DArmStadt Data Base System (DASDBS) for advanced applications is introduced DASDBS is conceived as a family of application-specific database systems on top of a common database kernel system. The main design problem considered here is, What features are common enough to be integrated into the kernel and what features are rather application-specific? Kernel features must be simple enough to be efficiently implemented and to serve a broad class of clients, yet powerful enough to form a convenient basis for application-oriented layers. Our kernel provides mechanisms to efficiently store hierarchically structured complex objects, and offers operations which are set-oriented and can be processed in a single scan through the objects. To achieve high concurrency in a layered system, a multi-level transaction methodology is applied. First experiences with our current implementation and some lessons we have learned from it are reported.
Heinz-Bernhard Paul, Hans-Jörg Schek, Marc H. Scholl, Gerhard Weikum, Uwe Deppisch
SIGMOD Conference2
1987 Supporting Flat Relations by a Nested Relational Kernel
Marc H. Scholl, Heinz-Bernhard Paul, Hans-Jörg Schek
VLDB3
1986 The relational model with relation-valued attributes
Hans-Jörg Schek, Marc H. Scholl
Inf. Syst.1
1984 Nested Transactions in a Combined IRS-DBMS Architecture
Hans-Jörg Schek
SIGIR1
1984 Architectural Issues of Transaction Management in Multi-Layered Systems
Gerhard Weikum, Hans-Jörg Schek
VLDB2
1982 Remarks on the Algebra of Non First Normal Form Relations
abstract
Usually, the first normal form condition of the relational model of data is imposed. Presently, a broader class of data base applications like office information systems is considered where this restriction is not convenient. Therefore, an extension of the relational model is proposed consisting of Non First Normal Form (NF2) relations. The relational algebra is enriched mainly by so called nest and unnest operations which transform between NF2 relations and the usual ones. We state some properties of these operations and some rules which occur in combination with the operations of the usual relational algebra. Since we propose to use the NF2 model also for the internal data model these rules are important not only for theoretical reasons but also for a practical implementation.
Gerhard Jaeschke, Hans-Jörg Schek
PODS2
1982 Data Structures for an Integrated Data Base Management and Information Retrieval System
Hans-Jörg Schek, Peter Pistor
VLDB1
1980 Methods for the Administration of Textual Data in Database Systems
Hans-Jörg Schek
SIGIR1