Bertram Ludäscher

dblp:l/BertramLudascher · DBLP profile ↗
← Back
74ranked-venue papers
14as first author
8since 2021 · last 2025
0000-0001-9140-936XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 46 · 10 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 2 since 2021Systems, architecture and hardware · 8 · 2 first-authorArtificial intelligence and machine learning · 6 · 4 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Theory of computation · 4 · 1 since 2021
YearPublicationVenuePosition
2025 Winning by Numbers: Connecting Strong Admissibility to Optimal Play in Argumentation
Shawn Bowers, Martin Caminada, Bertram Ludäscher
ECSQARU3
2025 LogicLM: Robust Application of Large Language Models with Logic Programming for Data Analytics
Evgeny S. Skvortsov, Shayan Mirjafari, Ojaswa Garg, Yilin Xia, Shawn Bowers, Bertram Ludäscher
EDBT6
2025 AF-XRAY: Visual Explanation and Resolution of Ambiguity in Legal Argumentation Frameworks
abstract
Argumentation frameworks (AFs) provide formal approaches for legal reasoning, but identifying sources of ambiguity and explaining argument acceptance remains challenging for non-experts. We present AF-XRAY, an open-source toolkit for exploring, analyzing, and visualizing abstract AFs in legal reasoning. AF-XRAY introduces: (i) layered visualizations based on game-theoretic argument length revealing well-founded derivation structures; (ii) classification of attack edges by semantic roles (primary, secondary, blunders); (iii) overlay visualizations of alternative 2-valued solutions on ambiguous 3-valued grounded semantics; and (iv) identification of critical attack sets whose suspension resolves undecided arguments. Through systematic generation of critical attack sets, AF-XRAY transforms ambiguous scenarios into grounded solutions, enabling users to pinpoint specific causes of ambiguity and explore alternative resolutions. We use real-world legal cases (e.g., Wild Animals as modeled by Bench-Capon) to show that our tool supports teleological legal reasoning by revealing how different assumptions lead to different justified conclusions.
Yilin Xia, Heng Zheng 0001, Shawn Bowers, Bertram Ludäscher
ICAIL4
2025 Modeling U.S. Supreme Court Briefs with Computational Argumentation
abstract
We show how to represent U.S. Supreme Court briefs as a structured argument model that relies entirely on the briefs’ existing structure.
Heng Zheng 0001, Dexter Williams, Bertram Ludäscher
JURIX3
2025 Natural Language to Logica: Towards Interactive and Explainable Data Analytics
Ojaswa Garg, Shayan Mirjafari, Yilin Xia, Shawn Bowers, Bertram Ludäscher, Evgeny S. Skvortsov
LOPSTR5
2024 Layered Visualization of Argumentation Frameworks
abstract
We propose a new layered visualization in PyArg for grounded labelings of abstract argumentation frameworks. Argument nodes are colored according to their label (IN, OUT, or UNDEC) and have a new length annotation, which is derived from provenance subgraphs. New edge annotations explain an attack-edge’s role in determining the value (label) of nodes in an argumentation framework.
Yilin Xia, Daphne Odekerken, Shawn Bowers, Bertram Ludäscher
COMMA4
2024 Logica: Declarative Data Science for Mere Mortals
Evgeny S. Skvortsov, Yilin Xia, Bertram Ludäscher
EDBT3
2024 SciPrompt: Knowledge-augmented Prompting for Fine-grained Categorization of Scientific Topics
abstract
Prompt-based fine-tuning has become an essential method for eliciting information encoded in pre-trained language models for a variety of tasks, including text classification.For multi-class classification tasks, prompt-based fine-tuning under low-resource scenarios has resulted in performance levels comparable to those of fully fine-tuning methods.Previous studies have used crafted prompt templates and verbalizers, mapping from the label terms space to the class space, to solve the classification problem as a masked language modeling task.However, cross-domain and finegrained prompt-based fine-tuning with an automatically enriched verbalizer remains unexplored, mainly due to the difficulty and costs of manually selecting domain label terms for the verbalizer, which requires humans with domain expertise.To address this challenge, we introduce SCIPROMPT, a framework designed to automatically retrieve scientific topic-related terms for low-resource text classification tasks.To this end, we select semantically correlated and domain-specific label terms within the context of scientific literature for verbalizer augmentation.Furthermore, we propose a new verbalization strategy that uses correlation scores as additional weights to enhance the prediction performance of the language model during model tuning.Our method outperforms stateof-the-art, prompt-based fine-tuning methods on scientific text classification tasks under few and zero-shot settings, especially in classifying fine-grained and emerging scientific topics 1 .
Zhiwen You, Kanyao Han, Bertram Ludäscher, Jana Diesner
EMNLP4
2020 Approximate Summaries for Why and Why-not Provenance
abstract
Why and why-not provenance have been studied extensively in recent years. However, why-not provenance and --- to a lesser degree --- why provenance can be very large, resulting in severe scalability and usability challenges. We introduce a novel approximate summarization technique for provenance to address these challenges. Our approach uses patterns to encode why and why-not provenance concisely. We develop techniques for efficiently computing provenance summaries that balance informativeness, conciseness, and completeness. To achieve scalability, we integrate sampling techniques into provenance capture and summarization. Our approach is the first to both scale to large datasets and generate comprehensive and meaningful summaries.
Seokki Lee, Bertram Ludäscher, Boris Glavic
Proc. VLDB Endow.2
2019 Application of BagIt-Serialized Research Object Bundles for Packaging and Re-Execution of Computational Analyses
abstract
In this paper we describe our experience adopting the Research Object Bundle (RO-Bundle) format with BagIt serialization (BagIt-RO) for the design and implementation of "tales" in the Whole Tale platform. A tale is an executable research object intended for the dissemination of computational scientific findings that captures information needed to facilitate understanding, transparency, and re-execution for review and computational reproducibility at the time of publication. We describe the Whole Tale platform and requirements that led to our adoption of BagIt-RO, specifics of our implementation, and discuss migrating to the emerging Research Object Crate (RO-Crate) standard.
Kyle Chard, Thomas Thelen, Matthew J. Turk, Craig Willis, Niall Gaffney, Matthew B. Jones, Kacper Kowalik, Bertram Ludäscher, Timothy M. McPhillips, Jarek Nabrzyski, Victoria Stodden, Ian J. Taylor
eScience8
2019 Reproducibility by Other Means: Transparent Research Objects
abstract
Research Objects have the potential to significantly enhance the reproducibility of scientific research. One important way Research Objects can do this is by encapsulating the means for re-executing the computational components of studies, thus supporting the new form of reproducibility enabled by digital computing-exact repeatability. However, Research Objects also can make scientific research more reproducible by supporting transparency, a component of reproducibility orthogonal to re-executability. We describe here our vision for making Research Objects more transparent by providing means for disambiguating claims about reproducibility generally, and computational repeatability specifically. We show how support for science-oriented queries can enable researchers to assess the reproducibility of Research Objects and the individual methods and results they encapsulate.
Timothy M. McPhillips, Craig Willis, Michael R. Gryk, Santiago Núñez-Corrales, Bertram Ludäscher
eScience5
2019 Computing environments for reproducibility: Capturing the "Whole Tale"
abstract
The act of sharing scientific knowledge is rapidly evolving away from traditional articles and presentations to the delivery of executable objects that integrate the data and computational details (e.g., scripts and workflows) upon which the findings rely. This envisioned coupling of data and process is essential to advancing science but faces technical and institutional barriers. The Whole Tale project aims to address these barriers by connecting computational, data-intensive research efforts with the larger research process—transforming the knowledge discovery and dissemination process into one where data products are united with research articles to create “living publications” or tales. The Whole Tale focuses on the full spectrum of science, empowering users in the long tail of science, and power users with demands for access to big data and compute resources. We report here on the design, architecture, and implementation of the Whole Tale environment.
Adam Brinckman, Kyle Chard, Niall Gaffney, Mihael Hategan, Matthew B. Jones, Kacper Kowalik, Sivakumar Kulasekaran, Bertram Ludäscher, Bryce D. Mecum, Jarek Nabrzyski, Victoria Stodden, Ian J. Taylor, Matthew J. Turk, Kandace Turner
Future Gener. Comput. Syst.8
2019 Verbalizing phylogenomic conflict: Representation of node congruence across competing reconstructions of the neoavian explosion
abstract
Phylogenomic research is accelerating the publication of landmark studies that aim to resolve deep divergences of major organismal groups. Meanwhile, systems for identifying and integrating the products of phylogenomic inference-such as newly supported clade concepts-have not kept pace. However, the ability to verbalize node concept congruence and conflict across multiple, in effect simultaneously endorsed phylogenomic hypotheses, is a prerequisite for building synthetic data environments for biological systematics and other domains impacted by these conflicting inferences. Here we develop a novel solution to the conflict verbalization challenge, based on a logic representation and reasoning approach that utilizes the language of Region Connection Calculus (RCC-5) to produce consistent alignments of node concepts endorsed by incongruent phylogenomic studies. The approach employs clade concept labels to individuate concepts used by each source, even if these carry identical names. Indirect RCC-5 modeling of intensional (property-based) node concept definitions, facilitated by the local relaxation of coverage constraints, allows parent concepts to attain congruence in spite of their differentially sampled children. To demonstrate the feasibility of this approach, we align two recent phylogenomic reconstructions of higher-level avian groups that entail strong conflict in the "neoavian explosion" region. According to our representations, this conflict is constituted by 26 instances of input "whole concept" overlap. These instances are further resolvable in the output labeling schemes and visualizations as "split concepts", which provide the labels and relations needed to build truly synthetic phylogenomic data environments. Because the RCC-5 alignments fundamentally reflect the trained, logic-enabled judgments of systematic experts, future designs for such environments need to promote a culture where experts routinely assess the intensionalities of node concepts published by our peers-even and especially when we are not in agreement with each other.
Nico M. Franz, Lukas J. Musher, Joseph W. Brown, Shizhuo Yu, Bertram Ludäscher
PLoS Comput. Biol.5
2019 PUG: a framework and practical implementation for why and why-not provenance
Seokki Lee, Bertram Ludäscher, Boris Glavic
VLDB J.2
2018 Provenance Summaries for Answers and Non-Answers
abstract
Explaining why an answer is (not) in the result of a query has proven to be of immense importance for many applications. However, why-not provenance, and to a lesser degree also why-provenance, can be very large, even for small input datasets. The resulting scalability and usability issues have limited the applicability of provenance. We present PUG , a system for why and why-not provenance that applies a range of novel techniques to overcome these challenges. Specifically, PUG limits provenance capture to what is relevant to explain a (missing) result of interest and uses an efficient sampling-based summarization method to produce compact explanations for (missing) answers. Using two real-world datasets, we demonstrate how a user can draw meaningful insights from explanations produced by PUG.
Seokki Lee, Bertram Ludäscher, Boris Glavic
Proc. VLDB Endow.2
2017 A SQL-Middleware Unifying Why and Why-Not Provenance for First-Order Queries
abstract
Explaining why an answer is in the result of a query or why it is missing from the result is important for many applications including auditing, debugging data and queries, and answering hypothetical questions about data. Both types of questions, i.e., why and why-not provenance, have been studied extensively. In this work, we present the first practical approach for answering such questions for queries with negation (firstorder queries). Our approach is based on a rewriting of Datalog rules (called firing rules) that captures successful rule derivations within the context of a Datalog query. We extend this rewriting to support negation and to capture failed derivations that explain missing answers. Given a (why or why-not) provenance question, we compute an explanation, i.e., the part of the provenance that is relevant to answer the question. We introduce optimizations that prune parts of a provenance graph early on if we can determine that they will not be part of the explanation for a given question. We present an implementation that runs on top of a relational database using SQL to compute explanations. Our experiments demonstrate that our approach scales to large instances and significantly outperforms an earlier approach which instantiates the full provenance to compute explanations.
Seokki Lee, Sven Köhler 0003, Bertram Ludäscher, Boris Glavic
ICDE3
2016 Introducing Explorer of Taxon Concepts with a case study on spider measurement matrix building
abstract
BACKGROUND: Taxonomic descriptions are traditionally composed in natural language and published in a format that cannot be directly used by computers. The Exploring Taxon Concepts (ETC) project has been developing a set of web-based software tools that convert morphological descriptions published in telegraphic style to character data that can be reused and repurposed. This paper introduces the first semi-automated pipeline, to our knowledge, that converts morphological descriptions into taxon-character matrices to support systematics and evolutionary biology research. We then demonstrate and evaluate the use of the ETC Input Creation - Text Capture - Matrix Generation pipeline to generate body part measurement matrices from a set of 188 spider morphological descriptions and report the findings. RESULTS: From the given set of spider taxonomic publications, two versions of input (original and normalized) were generated and used by the ETC Text Capture and ETC Matrix Generation tools. The tools produced two corresponding spider body part measurement matrices, and the matrix from the normalized input was found to be much more similar to a gold standard matrix hand-curated by the scientist co-authors. Special conventions utilized in the original descriptions (e.g., the omission of measurement units) were attributed to the lower performance of using the original input. The results show that simple normalization of the description text greatly increased the quality of the machine-generated matrix and reduced edit effort. The machine-generated matrix also helped identify issues in the gold standard matrix. CONCLUSIONS: ETC Text Capture and ETC Matrix Generation are low-barrier and effective tools for extracting measurement values from spider taxonomic descriptions and are more effective when the descriptions are self-contained. Special conventions that make the description text less self-contained challenge automated extraction of data from biodiversity descriptions and hinder the automated reuse of the published knowledge. The tools will be updated to support new requirements revealed in this case study.
Hong Cui, Dongfang Xu, Steven S. Chong, Martin Ramirez, Thomas Rodenhausen, James A. Macklin, Bertram Ludäscher, Robert A. Morris 0002, Eduardo M. Soto, Nicolás Mongiardino Koch
BMC Bioinform.7
2012 Win-move is coordination-free (sometimes)
abstract
In a recent paper by Hellerstein [15], a tight relationship was conjectured between the number of strata of a Datalog¬ program and the number of "coordination stages" required for its distributed computation. Indeed, Ameloot et al. [9] showed that a query can be computed by a coordination-free relational transducer network iff it is monotone, thus answering in the affirmative a variant of Hellerstein's CALM conjecture, based on a particular definition of coordination-free computation. In this paper, we present three additional models for declarative networking. In these variants, relational transducers have limited access to the way data is distributed. This variation allows transducer networks to compute more queries in a coordination-free manner: e.g., a transducer can check whether a ground atom A over the input schema is in the "scope" of the local node, and then send either A or ¬A to other nodes.
Daniel Zinn, Todd J. Green, Bertram Ludäscher
ICDT3
2012 Database Support for Exploring Scientific Workflow Provenance Graphs
Manish Kumar Anand, Shawn Bowers, Bertram Ludäscher
SSDBM3
2012 Workflows for microarray data processing in the Kepler environment
abstract
BACKGROUND: Microarray data analysis has been the subject of extensive and ongoing pipeline development due to its complexity, the availability of several options at each analysis step, and the development of new analysis demands, including integration with new data sources. Bioinformatics pipelines are usually custom built for different applications, making them typically difficult to modify, extend and repurpose. Scientific workflow systems are intended to address these issues by providing general-purpose frameworks in which to develop and execute such pipelines. The Kepler workflow environment is a well-established system under continual development that is employed in several areas of scientific research. Kepler provides a flexible graphical interface, featuring clear display of parameter values, for design and modification of workflows. It has capabilities for developing novel computational components in the R, Python, and Java programming languages, all of which are widely used for bioinformatics algorithm development, along with capabilities for invoking external applications and using web services. RESULTS: We developed a series of fully functional bioinformatics pipelines addressing common tasks in microarray processing in the Kepler workflow environment. These pipelines consist of a set of tools for GFF file processing of NimbleGen chromatin immunoprecipitation on microarray (ChIP-chip) datasets and more comprehensive workflows for Affymetrix gene expression microarray bioinformatics and basic primer design for PCR experiments, which are often used to validate microarray results. Although functional in themselves, these workflows can be easily customized, extended, or repurposed to match the needs of specific projects and are designed to be a toolkit and starting point for specific applications. These workflows illustrate a workflow programming paradigm focusing on local resources (programs and data) and therefore are close to traditional shell scripting or R/BioConductor scripting approaches to pipeline design. Finally, we suggest that microarray data processing task workflows may provide a basis for future example-based comparison of different workflow systems. CONCLUSIONS: We provide a set of tools and complete workflows for microarray data analysis in the Kepler environment, which has the advantages of offering graphical, clear display of conceptual steps and parameters and the ability to easily integrate other resources such as remote data and web services.
Thomas Stropp, Timothy M. McPhillips, Bertram Ludäscher, Mark Bieda
BMC Bioinform.3
2011 Towards Reliable, Performant Workflows for Streaming-Applications on Cloud Platforms
abstract
Scientific workflows are commonplace in eScience applications. Yet, the lack of integrated support for data models, including streaming data, structured collections and files, is limiting the ability of workflows to support emerging applications in energy informatics that are stream oriented. This is compounded by the absence of Cloud data services that support reliable and performant streams. In this paper, we propose and present a scientific workflow framework that supports streams as first-class data, and is optimized for performant and reliable execution across desktop and Cloud platforms. The workflow framework features and its empirical evaluation on a private Eucalyptus cloud are presented.
Daniel Zinn, Quinn J. Hart, Timothy M. McPhillips, Bertram Ludäscher, Yogesh L. Simmhan, Michail Giakkoupis, Viktor Prasanna 0001
CCGRID4
2011 Scientific workflow design 2.0: Demonstrating streaming data collections in Kepler
abstract
Scientific workflow systems are used to integrate existing software components (actors) into larger analysis pipelines to perform in silico experiments. Current approaches for handling data in nested-collection structures, as required in many scientific domains, lead to many record-management actors (shims) that make the workflow structure overly complex, and as a consequence hard to construct, evolve and maintain. By constructing and executing workflows from bioinformatics and geosciences in the Kepler system, we will demonstrate how COMAD (Collection-Oriented Modeling and Design), an extension of conventional workflow design, addresses these shortcomings. In particular, COMAD provides a hierarchical data stream model (as in XML) and a novel declarative configuration language for actors that functions as a middleware layer between the workflow's data model (streaming nested collections) and the actor's data model (base data and lists thereof). Our approach allows actor developers to focus on the internal actor processing logic oblivious to the workflow structure. Actors can then be re-used in various workflows simply by adapting actor configurations. Due to streaming nested collections and declarative configurations, COMAD workflows can usually be realized as linear data processing pipelines, which often reflect the scientific data analysis intention better than conventional designs. This linear structure not only simplifies actor insertions and deletions (workflow evolution), but also decreases the overall complexity of the workflow, reducing future effort in maintenance.
Lei Dou, Daniel Zinn, Timothy M. McPhillips, Sven Köhler 0003, Sean Riddle, Shawn Bowers, Bertram Ludäscher
ICDE7
2011 ProPub: Towards a Declarative Approach for Publishing Customized, Policy-Aware Provenance
Saumen C. Dey, Daniel Zinn, Bertram Ludäscher
SSDBM3
2011 Improving Workflow Fault Tolerance through Provenance-Based Recovery
Sven Köhler 0003, Sean Riddle, Daniel Zinn, Timothy M. McPhillips, Bertram Ludäscher
SSDBM5
2010 Techniques for efficiently querying scientific workflow provenance graphs
abstract
A key advantage of scientific workflow systems over traditional scripting approaches is their ability to automatically record data and process dependencies introduced during workflow runs. This information is often represented through provenance graphs, which can be used by scientists to better understand, reproduce, and verify scientific results. However, while most systems record and store data and process dependencies, few provide easy-to-use and efficient approaches for accessing and querying provenance information. Instead, users formulate provenance graph queries directly against physical data representations (e.g., relational, XML, or RDF), leading to queries that are difficult to express and expensive to evaluate. We address these problems through a high-level query language tailored for expressing provenance graph queries. The language is based on a general model of provenance supporting scientific workflows that process XML data and employ update semantics. Query constructs are provided for querying both structure and lineage information. Unlike other languages that return sets of nodes as answers, our query language is closed, i.e., answers to lineage queries are sets of lineage dependencies (edges) allowing answers to be further queried. We provide a formal semantics for the language and present novel techniques for efficiently evaluating lineage queries. Experimental results on real and synthetic provenance traces demonstrate that our lineage based optimizations outperform an in-memory and standard database implementation by orders of magnitude. We also show that our strategies are feasible and can significantly reduce both provenance storage size and query execution time when compared with standard approaches.
Manish Kumar Anand, Shawn Bowers, Bertram Ludäscher
EDBT3
2010 Provenance browser: Displaying and querying scientific workflow provenance graphs
abstract
This demonstration presents an interactive provenance browser for visualizing and querying data dependency (lineage) graphs produced by scientific workflow runs. The browser allows users to explore different views of provenance as well as to express complex and recursive graph queries through a high-level query language (QLP). Answers to QLP queries are lineage preserving in that queries return sets of lineage dependencies (denoting provenance graphs), which can be further queried and visually displayed (as graphs) in the browser. By combining provenance visualization, navigation, and query, the provenance browser can enable scientists to more easily access and explore scientific workflow provenance information.
Manish Kumar Anand, Shawn Bowers, Bertram Ludäscher
ICDE3
2010 XML-based computation for scientific workflows
abstract
Scientific workflows are increasingly used to rapidly integrate existing algorithms to create larger and more complex programs. However, designing workflows using purely dataflow-oriented computation models introduces a number of challenges, including the need to use low-level components to mediate and transform data (so-called shims) and large numbers of additional ¿wires¿ for routing data to components within a workflow. To address these problems, we employ Virtual Data Assembly Lines (VDAL), a modeling paradigm that can eliminate most shims and reduce wiring complexity. We show how a VDAL design can be implemented using existing XML technologies and how static analysis can provide significant help to scientists during workflow design and evolution, e.g., by displaying actor dependencies or by detecting so-called unproductive actors.
Daniel Zinn, Shawn Bowers, Bertram Ludäscher
ICDE3
2010 WATERS: a Workflow for the Alignment, Taxonomy, and Ecology of Ribosomal Sequences
abstract
BACKGROUND: For more than two decades microbiologists have used a highly conserved microbial gene as a phylogenetic marker for bacteria and archaea. The small-subunit ribosomal RNA gene, also known as 16 S rRNA, is encoded by ribosomal DNA, 16 S rDNA, and has provided a powerful comparative tool to microbial ecologists. Over time, the microbial ecology field has matured from small-scale studies in a select number of environments to massive collections of sequence data that are paired with dozens of corresponding collection variables. As the complexity of data and tool sets have grown, the need for flexible automation and maintenance of the core processes of 16 S rDNA sequence analysis has increased correspondingly. RESULTS: We present WATERS, an integrated approach for 16 S rDNA analysis that bundles a suite of publicly available 16 S rDNA analysis software tools into a single software package. The "toolkit" includes sequence alignment, chimera removal, OTU determination, taxonomy assignment, phylogentic tree construction as well as a host of ecological analysis and visualization tools. WATERS employs a flexible, collection-oriented 'workflow' approach using the open-source Kepler system as a platform. CONCLUSIONS: By packaging available software tools into a single automated workflow, WATERS simplifies 16 S rDNA analyses, especially for those without specialized bioinformatics, programming expertise. In addition, WATERS, like some of the newer comprehensive rRNA analysis tools, allows researchers to minimize the time dedicated to carrying out tedious informatics steps and to focus their attention instead on the biological interpretation of the results. One advantage of WATERS over other comprehensive tools is that the use of the Kepler workflow system facilitates result interpretation and reproducibility via a data provenance sub-system. Furthermore, new "actors" can be added to the workflow as desired and we see WATERS as an initial seed for a sizeable and growing repository of interoperable, easy-to-combine tools for asking increasingly complex microbial ecology questions.
Amber L. Hartman, Sean Riddle, Timothy M. McPhillips, Bertram Ludäscher, Jonathan A. Eisen
BMC Bioinform.4
2010 Parallelizing XML data-streaming workflows via MapReduce
Daniel Zinn, Shawn Bowers, Sven Köhler 0003, Bertram Ludäscher
J. Comput. Syst. Sci.4
2009 Scientific Workflows: Business as Usual?
Bertram Ludäscher, Mathias Weske, Timothy M. McPhillips, Shawn Bowers
BPM1
2009 Efficient provenance storage over nested data collections
abstract
Scientific workflow systems are increasingly used to automate complex data analyses, largely due to their benefits over traditional approaches for workflow design, optimization, and provenance recording. Many workflow systems employ a simple dependency model to represent the provenance of data produced by workflow runs. Although commonly adopted, this model does not capture explicit data dependencies introduced by "provenance-aware" processes, and it can lead to inefficient storage when workflow data is complex or structured. We present a provenance model, extending the conventional approach, that supports (i) explicit data dependencies and (ii) nested data collections. Our model adopts techniques from reference-based XML versioning, adding annotations for process and data dependencies. We present strategies and reduction techniques to store immediate and transitive provenance information within our model, and examine trade-offs among update time, storage size, and query response time. We evaluate our approach on real-world and synthetic workflow execution traces, demonstrating significant reductions in storage size, while also reducing the time required to store and query provenance information.
Manish Kumar Anand, Shawn Bowers, Timothy M. McPhillips, Bertram Ludäscher
EDBT4
2009 X-CSR: Dataflow Optimization for Distributed XML Process Pipelines
abstract
Abstract — XML process networks are a simple, yet power-ful programming paradigm for loosely coupled, coarse-grained dataflow applications such as data-centric scientific workflows. We describe a framework called ∆-XML that is well-suited for applications in which pipelines of data processors modify parts (“deltas”) of XML data collections while keeping the overall collection structure intact. We show how to optimize the execution of ∆-XML process networks by minimizing the data shipping cost in distributed settings. This X-CSR1 optimization employs static type inference based on XML Schema to determine the XML stream fragments that are relevant to a processor, allowing irrelevant fragments to be bypassed (“shipped”) to downstream pipeline steps. Finally, we present evaluation results for a real-world scientific workflow, which shows the practical feasibility of X-CSR. A long version of this paper is available as [1]. I.
Daniel Zinn, Shawn Bowers, Timothy M. McPhillips, Bertram Ludäscher
ICDE4
2009 Exploring Scientific Workflow Provenance Using Hybrid Queries over Nested Data and Lineage Graphs
Manish Kumar Anand, Shawn Bowers, Timothy M. McPhillips, Bertram Ludäscher
SSDBM4
2009 What Makes Scientific Workflows Scientific?
Bertram Ludäscher
SSDBM1
2009 Scientific workflow design for mere mortals
Timothy M. McPhillips, Shawn Bowers, Daniel Zinn, Bertram Ludäscher
Future Gener. Comput. Syst.4
2008 Provenance in collection-oriented scientific workflows
abstract
Abstract We describe a provenance model tailored to scientific workflows based on the collection‐oriented modeling and design paradigm. Our implementation within the Kepler scientific workflow system captures the dependencies of data and collection creation events on preexisting data and collections, and embeds these provenance records within the data stream. A provenance query engine operates on self‐contained workflow traces representing serializations of the output data stream for particular workflow runs. We demonstrate this approach in our response to the first provenance challenge. Copyright © 2007 John Wiley & Sons, Ltd.
Shawn Bowers, Timothy M. McPhillips, Bertram Ludäscher
Concurr. Comput. Pract. Exp.3
2008 From computation models to models of provenance: the RWS approach
abstract
Abstract Scientific workflows often benefit from or even require advanced modeling constructs, e.g. nesting of subworkflows, cycles for executing loops, data‐dependent routing, and pipelined execution. In such settings, an often overlooked aspect of provenance takes center stage: a suitable model of provenance (MoP) for scientific workflows should be based upon the underlying model of computation (MoC) used for executing the workflows. We can derive an adequate MoP from a MoC (such as Kahn's process networks) by taking into account the assumptions that a MoC entails, and by recording the observables which it affords. In this way, a MoP captures or at least better approximates ‘real’ data dependencies for workflows with advanced modeling constructs. As a specific instance, we elaborate on the Read–Write–ReSet model, a simple and flexible MoP suitable for a number of different MoCs. Copyright © 2007 John Wiley & Sons, Ltd.
Bertram Ludäscher, Norbert Podhorszki, Ilkay Altintas, Shawn Bowers, Timothy M. McPhillips
Concurr. Comput. Pract. Exp.1
2008 Special Issue: The First Provenance Challenge
abstract
Abstract The first Provenance Challenge was set up in order to provide a forum for the community to understand the capabilities of different provenance systems and the expressiveness of their provenance representations. To this end, a functional magnetic resonance imaging workflow was defined, which participants had to either simulate or run in order to produce some provenance representation, from which a set of identified queries had to be implemented and executed. Sixteen teams responded to the challenge, and submitted their inputs. In this paper, we present the challenge workflow and queries, and summarize the participants' contributions. Copyright © 2007 John Wiley & Sons, Ltd.
Luc Moreau 0001, Bertram Ludäscher, Ilkay Altintas, Roger S. Barga, Shawn Bowers, Steven P. Callahan, George Chin, Ben Clifford, Shirley Cohen, Sarah Cohen Boulakia, Susan B. Davidson, Ewa Deelman, Luciano A. Digiampietri, Ian T. Foster, Juliana Freire, James Frew, Joe Futrelle, Tara Gibson, Yolanda Gil, Carole A. Goble, Jennifer Golbeck, Paul Groth, David A. Holland, Jihie Kim, David Koop, Ales Krenek, Timothy M. McPhillips, Gaurang Mehta, Simon Miles, Dominic Metzger, Steve Munroe, James D. Myers, Beth Plale, Norbert Podhorszki, Varun Ratnakar, Emanuele Santos, Carlos Scheidegger, Karen Schuchardt, Margo I. Seltzer, Yogesh L. Simmhan, Cláudio T. Silva, Peter Slaughter, Eric G. Stephan, Robert Stevens 0001, Daniele Turi, Huy T. Vo, Michael Wilde, Jun Zhao 0003, Yong Zhao 0009
Concurr. Comput. Pract. Exp.2
2007 A knowledge environment for the biodiversity and ecological sciences
William K. Michener, James Beach, Matthew B. Jones, Bertram Ludäscher, Deana D. Pennington, Ricardo Scachetti Pereira, Arcot Rajasekar, Mark Schildhauer
J. Intell. Inf. Syst.4
2007 Rewriting queries using views with access patterns under integrity constraints
Alin Deutsch, Bertram Ludäscher, Alan Nash
Theor. Comput. Sci.2
2006 Scientific Workflows: More e-Science Mileage from Cyberinfrastructure
abstract
We view scientific workflows as the domain scientist's way to harness cyberinfrastructure for e-Science. Domain scientists are often interested in "end-to-end" frameworks which include data acquisition, transformation, analysis, visualization, and other steps. While there is no lack of technologies and standards to choose from, a simple, unified framework combining data modeling and processoriented modeling and design of scientific workflows has yet to emerge. Towards this end, we introduce a number of concepts such as models of computation and provenance, actor-oriented modeling, adapters, hybrid types, and higher-order components, and then outline a particular composition of some of these concepts, yielding a promising new synthesis for describing scientific workflows, i.e., Collection-Oriented Modeling and Design (COMAD).
Bertram Ludäscher, Shawn Bowers, Timothy M. McPhillips, Norbert Podhorszki
e-Science1
2006 S04 - Introduction to scientific workflow management and the Kepler system
abstract
A scientific workflow combines data and processes into a configurable, structured set of steps that implement semi-automated computational solutions of a scientific data management or analysis problem. Scientific workflow systems provide graphical user interfaces to combine different technologies along with efficient methods for using them with the goal to increase the efficiency of the scientists. This tutorial provides an introduction to scientific workflow construction and management (Part I) and includes a detailed hands-on session (Part II) using the Kepler system. It is intended for an audience with a computational science background. It will cover principles and foundations of scientific workflows, Kepler environment installation, workflow construction using Kepler library components, and workflow execution management that uses Kepler facilities to provide process and data monitoring and provenance information, as well as high speed data movement solutions. This tutorial also incorporates hands-on exercises and application examples from different scientific disciplines.
Ilkay Altintas, Bertram Ludäscher, Scott Klasky, Mladen A. Vouk
SC2
2006 An Extensible Infrastructure for Processing Distributed Geospatial Data Streams
abstract
Although the processing of data streams has been the focus of many research efforts in several areas, the case of remotely sensed streams in scientific contexts has received little attention. We present an extensible architecture to compose streaming image processing pipelines spanning multiple nodes on a network using a scientific workflow approach. This architecture includes (i) a mechanism for stream query dispatching so new streams can be dynamically generated from within individual processing nodes as a result of local or remote requests, and (ii) a mechanism for making the resulting streams externally available. As complete processing image pipelines can be cascaded across multiple interconnected nodes in a dynamic, scientist-driven way, the approach facilitates the reuse of data and the scalability of computations. We demonstrate the advantages of our infrastructure with a toolset of stream operators acting on remotely sensed data streams for realtime change detection
Carlos Rueda, Michael Gertz 0001, Bertram Ludäscher, Bernd Hamann
SSDBM3
2006 Scientific workflow management and the Kepler system
abstract
Abstract Many scientific disciplines are now data and information driven, and new scientific knowledge is often gained by scientists putting together data analysis and knowledge discovery ‘pipelines’. A related trend is that more and more scientific communities realize the benefits of sharing their data and computational services, and are thus contributing to a distributed data and computational community infrastructure (a.k.a. ‘the Grid’). However, this infrastructure is only a means to an end and ideally scientists should not be too concerned with its existence. The goal is for scientists to focus on development and use of what we call scientific workflows . These are networks of analytical steps that may involve, e.g., database access and querying steps, data analysis and mining steps, and many other steps including computationally intensive jobs on high‐performance cluster computers. In this paper we describe characteristics of and requirements for scientific workflows as identified in a number of our application projects. We then elaborate on Kepler, a particular scientific workflow system, currently under development across a number of scientific data management projects. We describe some key features of Kepler and its underlying Ptolemy II system, planned extensions, and areas of future research. Kepler is a community‐driven, open source project, and we always welcome related projects and new contributors to join. Copyright © 2005 John Wiley & Sons, Ltd.
Bertram Ludäscher, Ilkay Altintas, Chad Berkley, Dan Higgins, Efrat Jaeger, Matthew B. Jones, Edward A. Lee, Yang Zhao 0020
Concurr. Comput. Pract. Exp.1
2005 Actor-Oriented Design of Scientific Workflows
Shawn Bowers, Bertram Ludäscher
ER2
2005 Rewriting Queries Using Views with Access Patterns Under Integrity Constraints
Alin Deutsch, Bertram Ludäscher, Alan Nash
ICDT2
2005 Incorporating Semantics in Scientific Workflow Authoring
Chad Berkley, Shawn Bowers, Matthew B. Jones, Bertram Ludäscher, Mark Schildhauer
SSDBM4
2005 A Scientific Workflow Approach to Distributed Geospatial Data Processing using Web Services
Efrat Jaeger, Ilkay Altintas, Bertram Ludäscher, Deana D. Pennington, William K. Michener
SSDBM4
2005 Creating and Providing Data Management Services for the Biological and Ecological Sciences: Science Environment for Ecological Knowledge
Samantha Romanello, James Beach, Shawn Bowers, Matthew B. Jones, Bertram Ludäscher, William K. Michener, Deana D. Pennington, Arcot Rajasekar, Mark Schildhauer
SSDBM5
2004 Processing Unions of Conjunctive Queries with Negation under Limited Access Patterns
Alan Nash, Bertram Ludäscher
EDBT2
2004 Web Service Composition Through Declarative Queries: The Case of Conjunctive Queries with Union and Negation
abstract
A Web service operation can be seen as a function op: X/sub 1/,..., X/sub n/ /spl rarr/ Y/sub 1/,..., Y/sub m/ having an input message (request) with n arguments (parts), and an output message (response) with m parts. We study the problem of deciding whether a query Q is feasible, i.e., whether there exists a logically equivalent query Q' that can be executed observing the limited access patterns given by the Web service (source) relations. Executability depends on the specific syntactic form of a query, while feasibility is a more "robust" semantic notion, involving all equivalent queries (i.e., reorderings, minimized queries, etc). We show that deciding query feasibility (called "stability") is NP-complete for conjunctive queries (CQ) and for conjunctive queries with union (UCQ).
Bertram Ludäscher, Alan Nash
ICDE1
2004 A Web Service Composition and Deployment Framework for Scientific Workflows
abstract
The article presents the Web services framework in the Kepler scientific workflow system and illustrates them with a real-world example.
Ilkay Altintas, Efrat Jaeger, Bertram Ludäscher, Ashraf Memon
ICWS4
2004 Processing First-Order Queries under Limited Access Patterns
abstract
We study the problem of answering queries over sources with limited access patterns. Given a first-order query Q, the problem is to decide whether there is an equivalent query which can be executed observing the access patterns restrictions. If so, we say that Q is feasible. We define feasible for first-order queries---previous definitions handled only some existential cases---and characterize the complexity of many first-order query classes. For each of them, we show that deciding feasibility is as hard as deciding containment. Since feasibility is undecidable in many cases and hard to decide in some others, we also define an approximation to it which can be computed in NP for any first-order query and in P for unions of conjunctive queries with negation. Finally, we outline a practical overall strategy for processing first-order queries under limited access patterns.
Alan Nash, Bertram Ludäscher
PODS2
2004 Kepler: An Extensible System for Design and Execution of Scientific Workflows
abstract
Most scientists conduct analyses and run models in several different software and hardware environments, mentally coordinating the export and import of data from one environment to another. The Kepler scientific workflow system provides domain scientists with an easy-to-use yet powerful system for capturing scientific workflows (SWFs). SWFs are a formalization of the ad-hoc process that a scientist may go through to get from raw data to publishable results. Kepler attempts to streamline the workflow creation and execution process so that scientists can design, execute, monitor, re-run, and communicate analytical procedures repeatedly with minimal effort. Kepler is unique in that it seamlessly combines high-level workflow design with execution and runtime interaction, access to local and remote data, and local and remote service invocation. SWFs are superficially similar to business process workflows but have several challenges not present in the business workflow scenario. For example, they often operate on large, complex and heterogeneous data, can be computationally intensive and produce complex derived data products that may be archived for use in reparameterized runs or other workflows. Moreover, unlike business workflows, SWFs are often dataflow-oriented as witnessed by a number of recent academic systems (e.g., DiscoveryNet, Taverna and Triana) and commercial systems (Scitegic/Pipeline-Pilot, Inforsense). In a sense, SWFs are often closer to signal-processing and data streaming applications than they are to control-oriented business workflow applications.
Ilkay Altintas, Chad Berkley, Efrat Jaeger, Matthew B. Jones, Bertram Ludäscher, Steve Mock
SSDBM5
2004 On Integrating Scientific Resources through Semantic Registration
Shawn Bowers, Bertram Ludäscher
SSDBM3
2003 BIRN-M: A Semantic Mediator for Solving Real-World Neuroscience Problems
abstract
No abstract available.
Amarnath Gupta, Bertram Ludäscher, Maryann E. Martone
SIGMOD Conference2
2003 A Modeling and Execution Environment for Distributed Scientific Workflows
abstract
We illustrate how a domain scientist can perform a complex scientific task by interleaving data access, querying, and manipulation, as well as analytical steps and computations in complex, problem specific ways. We show how our system is used by a geneticist for solving the problem of discovering the so-called "co-regulated" genes by interlinking data and computation from several Web sites, local computations, as well as local and remote databases. The main distinctive features of our system (compared, e.g., to the ZOO environment (Ioannidis et al., 1996)) include: (i) executable workflows run as Web services; (ii) abstract workflows employ concept names and semantic types that are higher-level (and thus more "scientist friendly") than executable workflows; and (iii) our system supports automatic translation of the latter into the former.
Ilkay Altintas, Sangeeta Bhagwanani, David Buttler, Sandeep Chandra, Zhengang Cheng, Matthew Coleman, Terence Critchlow, Amarnath Gupta, Ling Liu 0001, Bertram Ludäscher, Calton Pu, Reagan W. Moore, Arie Shoshani, Mladen A. Vouk
SSDBM11
2003 Compiling Abstract Scientific Workflows into Web Service Workflows
abstract
The authors present an approach for compiling "scientist-friendly" abstract workflow specifications into real-world executable workflows of Web service invocations, using a set of abstract-as-view definitions from a repository of abstract tasks. There have been a number of proposals and systems for scientific workflow management. However, our approach features unique aspects, in particular: the separation of abstract and concrete executable workflows; and the use of database mediation techniques to automatically translate AWFs into EWFs.
Bertram Ludäscher, Ilkay Altintas, Amarnath Gupta
SSDBM1
2003 Towards a formalization of disease-specific ontologies for neuroinformatics
Amarnath Gupta, Bertram Ludäscher, Jeffrey S. Grethe, Maryann E. Martone
Neural Networks2
2002 Navigating Virtual Information Sources with Know-ME
Xufei Qian, Bertram Ludäscher, Maryann E. Martone, Amarnath Gupta
EDBT2
2002 Registering Scientific Information Sources for Semantic Mediation
Amarnath Gupta, Bertram Ludäscher, Maryann E. Martone
ER2
2002 A Transducer-Based XML Query Processor
Bertram Ludäscher, Pratik Mukhopadhyay, Yannis Papakonstantinou
VLDB1
2002 Understanding the global semantics of referential actions using logic rules
abstract
Referential actions are specialized triggers for automatically maintaining referential integrity in databases. While the local effects of referential actions can be grasped easily, it is far from obvious what the global semantics of a set of interacting referential actions should be. In particular, when using procedural execution models, ambiguities due to the execution ordering can occur. No global, declarative semantics of referential actions has yet been defined.We show that the well-known logic programming semantics provide a natural global semantics of referential actions that is based on their local characterization: To capture the global meaning of a set RA of referential actions, we first define their abstract (but non-constructive) intended semantics . Next, we formalize RA as a logic program P RA . The declarative, logic programming semantics of P RA then provide the constructive, global semantics of the referential actions. So, we do not define a semantics for referential actions, but we show that there exists a unique natural semantics if one is ready to accept (i) the intuitive local semantics of local referential actions, (ii) the formalization of those and of the local "effect-propagating" rules, and (iii) the well-founded or stable model semantics from logic programming as "reasonable" global semantics for local rules.We first focus on the subset of referential actions for deletions only. We prove the equivalence of the logic programming semantics and the abstract semantics via a game-theoretic characterization, which provides additional insight into the meaning of interacting referential actions. In this case a unique maximal admissible solution exists , computable by a ptime algorithm.Second, we investigate the general case---including modifications. We show that in this case there can be multiple maximal admissible subsets and that all maximal admissible subsets can be characterized as 3-valued stable models of P RA . We show that for a given set of user requests, in the presence of referential actions of the form ON UPDATE CASCADE, the admissibility check and the computation of the subsequent database state, and (for non-admissible updates) the derivation of debugging hints all are in ptime. Thus, full referential actions can be implemented efficiently.
Wolfgang May, Bertram Ludäscher
ACM Trans. Database Syst.2
2001 Model-Based Mediation with Domain Maps
abstract
Proposes an extension to current view-based mediator systems called "model-based mediation", in which views are defined and executed at the level of conceptual models (CMs) rather than at the structural level. Structural integration and lifting of data to the conceptual level is "pushed down" from the mediator to wrappers which, in our system, export the classes, associations, constraints and query capabilities of a source. Another novel feature of our architecture is the use of domain maps - semantic nets of concepts and relationships that are used to mediate across sources from multiple worlds (i.e. whose data are related in indirect and often complex ways). As part of registering a source's CM with the mediator, the wrapper creates a "semantic index" of its data into the domain map. We show that these indexes not only semantically correlate the multiple-worlds data, and thereby support the definition of the integrated CM, but they are also useful during query processing, for example, to select relevant sources. A first prototype of the system has been implemented for a complex neuroscience mediation problem.
Bertram Ludäscher, Amarnath Gupta, Maryann E. Martone
ICDE1
2001 Towards a federated neuroscientific knowledge management system using brain atlases
Gully A. P. C. Burns, Klaas E. Stephan, Bertram Ludäscher, Amarnath Gupta, Rolf Kötter
Neurocomputing3
2000 Navigation-Driven Evaluation of Virtual Mediated Views
Bertram Ludäscher, Yannis Papakonstantinou, Pavel E. Velikhov
EDBT1
2000 Knowledge-Based Integration of Neuroscience Data Sources
abstract
The need for information integration is paramount in many biological disciplines, because of the large heterogeneity in both the types of data involved and in the diversity of approaches (physiological, anatomical, biochemical, etc.) taken by biologists to study the same or correlated phenomena. However, the very heterogeneity makes the task of information integration very difficult since two approaches studying different aspects of the same phenomena may not even share common attributes in their schema description. The paper develops a wrapper-mediator architecture which extends the conventional data- and view-oriented information mediation approach by incorporating additional knowledge modules that bridge the gap between the heterogeneous data sources. The semantic integration of the disparate local data sources employs F-logic as a data and knowledge representation and reasoning formalism. We show that the rich object oriented modeling features of F-logic together with its declarative rule language and the uniform treatment of data and metadata (schema information) make it an ideal candidate for complex integration tasks. We substantiate this claim by elaborating on our integration architecture and illustrating the approach using real world examples from the neuroscience domain. The complete integration framework is currently under development; a first prototype establishing the viability of the approach is operational.
Amarnath Gupta, Bertram Ludäscher, Maryann E. Martone
SSDBM2
2000 Model-Based Information Integration in a Neuroscience Mediator System
Bertram Ludäscher, Amarnath Gupta, Maryann E. Martone
VLDB1
2000 Games and total Datalog¬ queries
Jörg Flum, Max Kubierschky, Bertram Ludäscher
Theor. Comput. Sci.3
1999 XML-Based Information Mediation with MIX
abstract
The MIX mediator system, MIXm, is developed as part of the MIX Project at the San Diego Supercomputer Center, and the University of California, San Diego.1 MIXm uses XML as the common model for data exchange. Mediator views are expressed in XMAS (XML Matching And Structuring Language), a declarative XML query language. To facilitate user-friendly query formulation and for optimization purposes, MIXm employs XML DTDs as a structural description (in effect, a “schema”) of the exchanged data. The novel features of the system include:
Chaitanya K. Baru, Amarnath Gupta, Bertram Ludäscher, Richard Marciano, Yannis Papakonstantinou, Pavel E. Velikhov, Vincent Chu
SIGMOD Conference3
1998 Referential Actions: From Logical Semantics to Implementation
Bertram Ludäscher, Wolfgang May
EDBT1
1998 Managing Semistructured Data with FLORID: A Deductive Object-Oriented Perspective
Bertram Ludäscher, Rainer Himmeröder, Georg Lausen, Wolfgang May, Christian Schlepphorst
Inf. Syst.1
1997 Total and Partial Well-Founded Datalog Coincide
Jörg Flum, Max Kubierschky, Bertram Ludäscher
ICDT3
1997 Referential Actions as Logic Rules
abstract
Referential actions are specialized triggers used to automatically maintain referential integrity.While their local behavior can be grasped easily, it is far from clear what the combined effect of a set of referential actions, i.e., their global semantics should be.For example, different execution orders may lead to ambiguities in determining the final set of updates to be applied.To resolve these problems, we propose an abstract logical framework for rule-based maintenance of referential integrity: First, we identify desirable abstract properties like admissibility of updates which lead to a non-constructive global semantics of referential actions.We obtain a constructive definition by formalizing a set of referential actions RA as logical rules, and show that the declarative semantics of the resulting logic program PRA captures the intended abstract semantics: The well-founded model of PRA yields a unique set of updates, which is a safe, sceptical approximation of the set of all maximal admissible updates; the thud truth-value undefined is assigned to all controversiaI updates.Finally, we show how to obtain a characterization of all maximal admissible subsets of a given set of updates using certain maximal stable models.-.-.--... .
Bertram Ludäscher, Wolfgang May, Georg Lausen
PODS1