Robert Schuler

dblp:65/6901 · also Robert E. Schuler · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0002-8956-5707ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 8 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 It's the Data Stupid
abstract
Artificial Intelligence and Machine Learning have emerged as a promising approach to scientific investigations, but there is a persistent shortage of high-quality, properly annotated datasets suitable for training models. Here, we outline some of the widely reported characteristics for making AI-ready data and compare that with FAIR data. We discuss the limitations of traditional data repositories and the challenges associated with establishing a data repository that can grow with scientific communities and accommodate rapid evolution in research priorities. Finally, we introduce the SCALE principles for repository design that offer a proven framework for creating sustainable, scalable data repositories that can adapt to new data models and methodologies, ensuring that software infrastructure serves research needs rather than constraining them.
Carl Kesselman, Robert Schuler
eScience2
2024 Creating Thriving Data-Centric Communities from Basic Research to Commercial Applications
abstract
The ability to accumulate and analyze large quantities of data is rapidly becoming a competitive advantage not only in science but in the broader economy as well. Advances such as AlphaFold, the AI-based protein prediction tool and ChatGPT the large language model-based chat bot, have ignited enormous excitement in science and industry for leveraging data and computational techniques to solve important problems. However, what is typically lost in all the excitement is the fact that such startling achievements were only possible after a critical mass of high quality data existed to train models using machine learning algorithms. Both examples relied on open data sources that were generated painstakingly by user communities over the course of decades. We argue that in order to unlock future high impact data science achievements like these will require a culture of and skill set for data management, sharing and reuse. In this paper, we describe our work within the dental, oral, and craniofacial community to create such a data sharing community that has grown out of basic research to increasingly touch on clinical sciences and even commercial applications.
Robert Schuler, Carl Kesselman
e-Science1
2023 Let's Put the Science in eScience
abstract
The underlying premise behind eScience is that computational methods and data-driven approaches can contribute to scientific discovery on a par with, or even superior to, traditional experimental methods; that the combination of computers, software, and extant data collections are the modern equivalent to the scientific instruments that have led to our understanding of fundamental laws in physics, chemistry, biology, and other domains. However, a robust methodology for making the results of eScience activities “scientific” is lacking, with significant consequences. In this brief paper we propose a shift in perspective as to what it means to create an eScience-based result and how the scientific validity of eScience experiments might be improved.
Carl Kesselman, Robert Schuler, Ian T. Foster
e-Science2
2023 Database Evolution, by Scientists, for Scientists: A Case Study
abstract
Database management systems have been used to great advantage for industry usage scenarios. As science becomes increasingly dependent on carefully organized and curated data to inform and drive new discoveries, the need for database management systems has grown significantly. The long standing “20 questions” method was advocated in early studies of applying relational databases for science in order to elicit requirements for designing and developing information systems for scientific data. It has been observed, however, that database designs become outdated within months of usage leading to degradation in the quality of the database schema. In addition, there is limited evidence that scientists themselves have the tools and processes necessary to develop and maintain scientific databases without reliance on database administrators. Beyond learning to query databases, scientists need tools to create and evolve databases and guidance on how to apply those tools to develop information systems. In this paper, we present a simplified methodology for database evolution for scientists and a case study of database evolution by a scientist in the context of a research database for cell modeling. We include a detailed analysis of the activities and processes employed by the scientist during the schema evolution. Our results show that a scientist can successfully evolve a complex information system driven by new research requirements.
Robert Schuler, Jitin Singla, Brinda Vallat, Kate L. White, Helen M. Berman, Carl Kesselman
e-Science1
2022 Managing Database-Application Co-Evolution in a Scientific Data Ecosystem
abstract
Scientific databases used for organizing, archiving, collaborating and sharing research data depend on a well-defined schema to accurately reflect the scientific domain and on database-driven applications for supporting key user interactions with the database. Applications that interact with a database typically depend on some form of schema mappings, such as object-relational mappings, to inform the application of how to query and manipulate the database. The presence of schema mappings, however, further exacerbates the already difficult task of evolving the database schema. Database migration utilities provide some help by coordinating schema evolution scripts with application code changes, but only automate the simplest schema mapping changes. In this paper, we present an approach to coupled database-application evolution by extending a database evolution language with model management operations. We introduce a novel set of model management operations and define their semantics and then describe how they may be integrated into schema modification operators. We then present an evaluation of the concepts from real-world usage of model mappings in scientific database deployments.
Robert Schuler, Carl Kesselman
e-Science1
2021 CHiSEL: a user-oriented framework for simplifing database evolution
Robert Schuler, Carl Kesselman
Distributed Parallel Databases1
2020 Towards Co-Evolution of Data-Centric Ecosystems
abstract
Database evolution is a notoriously difficult task, and it is exacerbated by the necessity to evolve database-dependent applications. As science becomes increasingly dependent on sophisticated data management, the need to evolve an array of database-driven systems will only intensify. In this paper, we present an architecture for data-centric ecosystems that allows the components to seamlessly co-evolve by centralizing the models and mappings at the data service and pushing model-adaptive interactions to the database clients. Boundary objects fill the gap where applications are unable to adapt and need a stable interface to interact with the components of the ecosystem. Finally, evolution of the ecosystem is enabled via integrated schema modification and model management operations. We present use cases from actual experiences that demonstrate the utility of our approach.
Robert Schuler, Karl Czajkowski, Mike D'Arcy, Hongsuda Tangmunarunkit, Carl Kesselman
SSDBM1
2019 Toward FAIR Knowledge Turns in Bioinformatics
abstract
Sharing of bioinformatics data within research communities holds the promise of facilitating more rapid discovery, yet the volume of data is growing at a pace exponentially greater than what traditional biocuration can support. We present here an approach that we have used to empower data producing researchers to curate high quality shared data that is ready for reuse and re-analysis.
Robert Schuler, Alejandro Bugacov, Matthew Blow, Carl Kesselman
BIBM1
2019 A High-level User-oriented Framework for Database Evolution
abstract
Databases are well suited to the task of describing and organizing research datasets, however, the difficulties of using database management systems effectively have resulted in their limited usage among domain scientists. Scientists operate in an environment that is changing steadily with new experimental protocols, instruments, and discoveries that impact what datasets they generate and how they describe and organize them. In order to manage datasets for a scientific application, scientists need to routinely revise their database schemas to reflect these changes. Unfortunately, evolving a database is one of the well-known and most difficult aspects of database usage. The conventional data definition and manipulation languages offer relatively low-level programming abstractions to perform complex database evolution tasks, and therefore require specialized technical skills not possessed by most domain scientists. A simplified means of expressing database evolution operations can reduce the effort of keeping the scientific database in sync with changing requirements. This paper presents a high-level, user-oriented, schema evolution framework with an algebra of specialized schema modification operators. The approach allows introduction of novel operators as motivated by new requirements and is amenable to well established optimization techniques for efficient planning and execution. We present the framework and its implementation, and we demonstrate its utility in an exemplar use case and performance evaluation.
Robert Schuler, Carl Kesselman
SSDBM1
2018 ERMrest: a web service for collaborative data management
abstract
The foundation of data oriented scientific collaboration is the ability for participants to find, access and reuse data created during the course of an investigation, what has been referred to as the FAIR principles. In this paper, we describe ERMrest, a collaborative data management service that promotes data oriented collaboration by enabling FAIR data management throughout the data life cycle. ERMrest is a RESTful web service that promotes discovery and reuse by organizing diverse data assets into a dynamic entity relationship model. We present details on the design and implementation of ERMrest, data on its performance and its use by a range of collaborations to accelerate and enhance their scientific output.
Karl Czajkowski, Carl Kesselman, Robert Schuler, Hongsuda Tangmunarunkit
SSDBM3
2018 Towards an efficient and effective framework for the evolution of scientific databases
abstract
Database systems are well suited to scientific data management and analysis workloads, however, a database must evolve to keep pace with changing requirements and adjust to changes in the domain conceptualization as applications mature. Evolving a database (i.e., updating its schema and instance data) is one of the greatest challenges in database maintenance and the difficulties are compounded by the lack of sufficient tools to support scientists. This paper presents a schema evolution framework based on an algebraic approach that introduces extended and higher-level composite relational operators tailored to the task of schema evolution. These higher-level operators simplify the task of evolving a database for non-expert users, while enabling efficient evaluation of schema evolution expressions.
Robert Schuler, Carl Kesselman
SSDBM1
2017 Experiences with DERIVA: An Asset Management Platform for Accelerating eScience
abstract
The pace of discovery in eScience is increasingly dependent on a scientist's ability to acquire, curate, integrate, analyze, and share large and diverse collections of data. It is all too common for investigators to spend inordinate amounts of time developing ad hoc procedures to manage their data. In previous work, we presented Deriva, a Scientific Asset Management System, designed to accelerate data driven discovery. In this paper, we report on the use of Deriva in a number of substantial and diverse eScience applications. We describe the lessons we have learned, both from the perspective of the Deriva technology, as well as the ability and willingness of scientists to incorporate Scientific Asset Management into their daily workflows.
Alejandro Bugacov, Karl Czajkowski, Carl Kesselman, Robert Schuler, Hongsuda Tangmunarunkit
eScience5
2017 ERMRest: A Collaborative Data Catalog with Fine Grain Access Control
abstract
Creating and maintaining an accurate description of data assets and the relationships between assets is a critical aspect of making data findable, accessible, interoperable, and reusable (FAIR). Typically, such metadata are created and maintained in a data catalog by a curator as part of data publication. However, allowing metadata to be created and maintained by data producers as the data is generated rather then waiting for publication can have significant advantages in terms of productivity and repeatability. The responsibilities for metadata management need not fall on any one individual, but rather may be delegated to appropriate members of a collaboration, enabling participants to edit or maintain specific attributes, to describe relationships between data elements, or to correct errors. To support such collaborative data editing, we have created ERMrest, a relational data service for the Web that enables the creation, evolution and navigation of complex models used to describe and structure diverse file or relational data objects. A key capability of ERMrest is its ability to control operations down to the level of individual data elements, i.e. fine-grained access control, so that many different modes of data-oriented collaboration can be supported. In this paper we introduce ERMrest and describe its fine-grained access control capabilities that support collaborative editing. ERMrest is in daily use in many data driven collaborations and we describe a sample policy that is based on a common biocuration pattern.
Karl Czajkowski, Carl Kesselman, Robert Schuler
eScience3
2016 Accelerating data-driven discovery with scientific asset management
abstract
The overhead and burden of managing data in complex discovery processes involving experimental protocols with numerous data-producing and computational steps has become the gating factor that determines the pace of discovery. The lack of comprehensive systems to capture, manage, organize and retrieve data throughout the discovery life cycle leads to significant overheads on scientists' time and effort, reduced productivity, lack of reproducibility, and an absence of data sharing. In “creative fields” like digital photography and music, digital asset management (DAM) systems for capturing, managing, curating and consuming digital assets like photos and audio recordings, have fundamentally transformed how these data are used. While asset management has not taken hold in eScience applications, we believe that transformation similar to that observed in the creative space could be achieved in scientific domains if appropriate ecosystems of asset management tools existed to capture, manage, and curate data throughout the scientific discovery process. In this paper, we introduce DERIVA, a framework and infrastructure for asset management in eScience and present initial results from its usage in active research use cases.
Robert Schuler, Carl Kesselman, Karl Czajkowski
eScience1
2014 Digital asset management for heterogeneous biomedical data in an era of data-intensive science
abstract
Biomedical research depends upon increasingly high throughput instruments and sophisticated data analytics. In spite of the significant overhead of handling research data, there is little support for researchers to manage and organize data for purposes of exploration, analysis, and ultimately publication. Shared file systems with metadata coded into directory hierarchies and spreadsheets are the common practice. In this paper, we present a digital asset management approach and system for streamlining data operations and reducing data management overheads for biomedical researchers. It consists of data management tasks including storage, archival, annotation, search, cataloging, publication, and collaboration. We present a preliminary performance evaluation of a key component of the system, and we demonstrate the utility of this approach in a pilot deployment and user study.
Robert Schuler, Carl Kesselman, Karl Czajkowski
BIBM1
2014 Efficient Data Staging Using Performance-Based Adaptation and Policy-Based Resource Allocation
abstract
Before scientific analyses run on shared infrastructure, such as the Open Science Grid or XSEDE, scientists must often transfer or stage key data sets those resources. Often these datasets consist of many files that may be transferred by multiple clients in parallel. We study two techniques that improve the use of available resources for these large, long-running, multi-file transfers. First, we adapt transfer parameters for multi-file transfers based on recent transfer performance. Second, we use VO and site policies to influence the allocation of system resources for transfers, such as available transfer streams. We describe our system design and summarize its implementation and performance.
Ann L. Chervenak, Alex Sim, Junmin Gu, Robert Schuler, Nandan Hirpathak
PDP4
2011 Enabling collaborative research using the Biomedical Informatics Research Network (BIRN)
abstract
OBJECTIVE: As biomedical technology becomes increasingly sophisticated, researchers can probe ever more subtle effects with the added requirement that the investigation of small effects often requires the acquisition of large amounts of data. In biomedicine, these data are often acquired at, and later shared between, multiple sites. There are both technological and sociological hurdles to be overcome for data to be passed between researchers and later made accessible to the larger scientific community. The goal of the Biomedical Informatics Research Network (BIRN) is to address the challenges inherent in biomedical data sharing. MATERIALS AND METHODS: BIRN tools are grouped into 'capabilities' and are available in the areas of data management, data security, information integration, and knowledge engineering. BIRN has a user-driven focus and employs a layered architectural approach that promotes reuse of infrastructure. BIRN tools are designed to be modular and therefore can work with pre-existing tools. BIRN users can choose the capabilities most useful for their application, while not having to ensure that their project conforms to a monolithic architecture. RESULTS: BIRN has implemented a new software-based data-sharing infrastructure that has been put to use in many different domains within biomedicine. BIRN is actively involved in outreach to the broader biomedical community to form working partnerships. CONCLUSION: BIRN's mission is to provide capabilities and services related to data sharing to the biomedical research community. It does this by forming partnerships and solving specific, user-driven problems whose solutions are then available for use by other groups.
Karl G. Helmer, José Luis Ambite, Joseph Ames, Rachana Ananthakrishnan, Gully A. P. C. Burns, Ann L. Chervenak, Ian T. Foster, Liming Lee, David B. Keator, Fabio Macciardi, Ravi K. Madduri, John-Paul Navarro, Steven G. Potkin, Bruce R. Rosen, Seth Ruffins, Robert Schuler, Jessica A. Turner, Arthur W. Toga, Christina Williams, Carl Kesselman
J. Am. Medical Informatics Assoc.16
2009 The Globus Replica Location Service: Design and Experience
abstract
Distributed computing systems employ replication to improve overall system robustness, scalability, and performance. A replica location service (RLS) offers a mechanism to maintain and provide information about physical locations of replicas. This paper defines a design framework for RLSs that supports a variety of deployment options. We describe the RLS implementation that is distributed with the Globus toolkit and is in production use in several grid deployments. Features of our modular implementation include the use of soft-state protocols to populate a distributed index and Bloom filter compression to reduce overheads for distribution of index information. Our performance evaluation demonstrates that the RLS implementation scales well for individual servers with millions of entries and up to 100 clients. We describe the characteristics of existing RLS deployments and discuss how RLS has been integrated with higher-level data management services.
Ann L. Chervenak, Robert Schuler, Matei Ripeanu, Muhammad Ali Amer, Shishir Bharathi, Ian T. Foster, Adriana Iamnitchi, Carl Kesselman
IEEE Trans. Parallel Distributed Syst.2