Daniel Zinn

dblp:66/780 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 4 first-authorSystems, architecture and hardware · 2 · 1 first-authorTheory of computation · 2 · 1 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Distributed computing theory · 50% Algorithmic game theory and mechanism design · 28% Logic in computer science · 22%
Databases, data mining, and information retrieval
6 papers
Data models and query languages · 35% Distributed and cloud data management · 24% Data integration and cleaning · 20%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 100%
Network and information security
1 paper
Network security · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed computing theory › distributed computability
coordination-free computation
0.422016
Weaker Forms of Monotonicity for Declarative Networking: A More Fine-Grained Answer to the CALM-Conjecture · ACM Trans. Database Syst. 2016
Weaker forms of monotonicity for declarative networking: a more fine-grained answer to the calm-conjecture · PODS 2014
Computational science and engineering
scientific workflow
0.332011
Scientific workflow design 2.0: Demonstrating streaming data collections in Kepler · ICDE 2011
XML-based computation for scientific workflows · ICDE 2010
X-CSR: Dataflow Optimization for Distributed XML Process Pipelines · ICDE 2009
Data models and query languages
datalog
0.212016
Weaker Forms of Monotonicity for Declarative Networking: A More Fine-Grained Answer to the CALM-Conjecture · ACM Trans. Database Syst. 2016
Algorithmic game theory and mechanism design › social choice
monotonicity
0.212016
Weaker Forms of Monotonicity for Declarative Networking: A More Fine-Grained Answer to the CALM-Conjecture · ACM Trans. Database Syst. 2016
Distributed and cloud data management
declarative networking
0.212014
Weaker forms of monotonicity for declarative networking: a more fine-grained answer to the calm-conjecture · PODS 2014
Logic in computer science › logic programming
datalog
0.212014
Weaker forms of monotonicity for declarative networking: a more fine-grained answer to the calm-conjecture · PODS 2014
Query processing and optimization
data flow optimization
0.112009
X-CSR: Dataflow Optimization for Distributed XML Process Pipelines · ICDE 2009
Network security
content filtering
0.112007
ConceptDoppler: a weather tracker for internet censorship · CCS 2007
Network security › censorship
internet censorship
0.112007
ConceptDoppler: a weather tracker for internet censorship · CCS 2007
Data integration and cleaning
data transformation
0.012010
XML-based computation for scientific workflows · ICDE 2010
Data models and query languages › XML data management
XML data processing
0.012009
X-CSR: Dataflow Optimization for Distributed XML Process Pipelines · ICDE 2009

Methods — techniques the papers use, named apart from their topics

relational transducer networks · 0.9datalog · 0.5monotonicity analysis · 0.4declarative configuration · 0.2collection-oriented modeling and design · 0.2static analysis · 0.2XML-based modeling · 0.2static type inference · 0.2XML Schema · 0.2
YearPublicationVenuePosition
2017 Datalog Queries Distributing over Components
abstract
We investigate the class D of queries that distribute over components. These are the queries that can be evaluated by taking the union of the query results over the connected components of the database instance. We show that it is undecidable whether a (positive) Datalog program distributes over components. Additionally, we show that connected Datalog¬ (the fragment of Datalog¬ where all rules are connected) provides an effective syntax for Datalog¬ programs that distribute over components under the stratified as well as under the well-founded semantics. As a corollary, we obtain a simple proof for one of the main results in previous work [Zinn et al. 2012], namely that the classic win-move query is in F 2 (a particular class of coordination-free queries).
Tom J. Ameloot, Bas Ketsman, Frank Neven, Daniel Zinn
ACM Trans. Comput. Log.4
2016 Weaker Forms of Monotonicity for Declarative Networking: A More Fine-Grained Answer to the CALM-Conjecture
abstract
The CALM-conjecture, first stated by Hellerstein [2010] and proved in its revised form by Ameloot et al. [2013] within the framework of relational transducer networks, asserts that a query has a coordination-free execution strategy if and only if the query is monotone. Zinn et al. [2012] extended the framework of relational transducer networks to allow for specific data distribution strategies and showed that the nonmonotone win-move query is coordination-free for domain-guided data distributions. In this article, we extend the story by equating increasingly larger classes of coordination-free computations with increasingly weaker forms of monotonicity and present explicit Datalog variants that capture each of these classes. One such fragment is based on stratified Datalog where rules are required to be connected with the exception of the last stratum. In addition, we characterize coordination-freeness as those computations that do not require knowledge about all other nodes in the network, and therefore, can not globally coordinate. The results in this article can be interpreted as a more fine-grained answer to the CALM-conjecture.
Tom J. Ameloot, Bas Ketsman, Frank Neven, Daniel Zinn
ACM Trans. Database Syst.4
2015 Datalog Queries Distributing over Components
abstract
We investigate the class D of queries that distribute over components. These are the queries that can be evaluated by taking the union of the query results over the connected components of the database instance. We show that it is undecidable whether a (positive) Datalog program distributes over components. Additionally, we show that connected Datalog with Negation (the fragment of Datalog with Negation where all rules are connected) provides an effective syntax for Datalog with Negation programs that distribute over components under the stratified as well as under the well-founded semantics. As a corollary, we obtain a simple proof for one of the main results in previous work [Zinn, Green, and Ludäscher, ICDT2012], namely, that the classic win-move query is in F_2 (a particular class of coordination-free queries).
Tom J. Ameloot, Bas Ketsman, Frank Neven, Daniel Zinn
ICDT4
2014 Weaker forms of monotonicity for declarative networking: a more fine-grained answer to the calm-conjecture
abstract
The CALM-conjecture, first stated by Hellerstein [23] and proved in its revised form by Ameloot et al. [13] within the framework of relational transducer networks, asserts that a query has a coordination-free execution strategy if and only if the query is monotone. Zinn et al. [32] extended the framework of relational transducer networks to allow for specific data distribution strategies and showed that the nonmonotone win-move query is coordination-free for domain-guided data distributions. In this paper, we complete the story by equating increasingly larger classes of coordination-free computations with increasingly weaker forms of monotonicity and make Datalog variants explicit that capture each of these classes. One such fragment is based on stratified Datalog where rules are required to be connected with the exception of the last stratum. In addition, we characterize coordination-freeness as those computations that do not require knowledge about all other nodes in the network, and therefore, can not globally coordinate. The results in this paper can be interpreted as a more fine-grained answer to the CALM-conjecture.
Tom J. Ameloot, Bas Ketsman, Frank Neven, Daniel Zinn
PODS4
2012 Win-move is coordination-free (sometimes)
abstract
In a recent paper by Hellerstein [15], a tight relationship was conjectured between the number of strata of a Datalog¬ program and the number of "coordination stages" required for its distributed computation. Indeed, Ameloot et al. [9] showed that a query can be computed by a coordination-free relational transducer network iff it is monotone, thus answering in the affirmative a variant of Hellerstein's CALM conjecture, based on a particular definition of coordination-free computation. In this paper, we present three additional models for declarative networking. In these variants, relational transducers have limited access to the way data is distributed. This variation allows transducer networks to compute more queries in a coordination-free manner: e.g., a transducer can check whether a ground atom A over the input schema is in the "scope" of the local node, and then send either A or ¬A to other nodes.
Daniel Zinn, Todd J. Green, Bertram Ludäscher
ICDT1
2011 Towards Reliable, Performant Workflows for Streaming-Applications on Cloud Platforms
abstract
Scientific workflows are commonplace in eScience applications. Yet, the lack of integrated support for data models, including streaming data, structured collections and files, is limiting the ability of workflows to support emerging applications in energy informatics that are stream oriented. This is compounded by the absence of Cloud data services that support reliable and performant streams. In this paper, we propose and present a scientific workflow framework that supports streams as first-class data, and is optimized for performant and reliable execution across desktop and Cloud platforms. The workflow framework features and its empirical evaluation on a private Eucalyptus cloud are presented.
Daniel Zinn, Quinn J. Hart, Timothy M. McPhillips, Bertram Ludäscher, Yogesh L. Simmhan, Michail Giakkoupis, Viktor Prasanna 0001
CCGRID1
2011 Scientific workflow design 2.0: Demonstrating streaming data collections in Kepler
abstract
Scientific workflow systems are used to integrate existing software components (actors) into larger analysis pipelines to perform in silico experiments. Current approaches for handling data in nested-collection structures, as required in many scientific domains, lead to many record-management actors (shims) that make the workflow structure overly complex, and as a consequence hard to construct, evolve and maintain. By constructing and executing workflows from bioinformatics and geosciences in the Kepler system, we will demonstrate how COMAD (Collection-Oriented Modeling and Design), an extension of conventional workflow design, addresses these shortcomings. In particular, COMAD provides a hierarchical data stream model (as in XML) and a novel declarative configuration language for actors that functions as a middleware layer between the workflow's data model (streaming nested collections) and the actor's data model (base data and lists thereof). Our approach allows actor developers to focus on the internal actor processing logic oblivious to the workflow structure. Actors can then be re-used in various workflows simply by adapting actor configurations. Due to streaming nested collections and declarative configurations, COMAD workflows can usually be realized as linear data processing pipelines, which often reflect the scientific data analysis intention better than conventional designs. This linear structure not only simplifies actor insertions and deletions (workflow evolution), but also decreases the overall complexity of the workflow, reducing future effort in maintenance.
Lei Dou, Daniel Zinn, Timothy M. McPhillips, Sven Köhler 0003, Sean Riddle, Shawn Bowers, Bertram Ludäscher
ICDE2
2011 ProPub: Towards a Declarative Approach for Publishing Customized, Policy-Aware Provenance
Saumen C. Dey, Daniel Zinn, Bertram Ludäscher
SSDBM2
2011 Improving Workflow Fault Tolerance through Provenance-Based Recovery
Sven Köhler 0003, Sean Riddle, Daniel Zinn, Timothy M. McPhillips, Bertram Ludäscher
SSDBM3
2010 XML-based computation for scientific workflows
abstract
Scientific workflows are increasingly used to rapidly integrate existing algorithms to create larger and more complex programs. However, designing workflows using purely dataflow-oriented computation models introduces a number of challenges, including the need to use low-level components to mediate and transform data (so-called shims) and large numbers of additional ¿wires¿ for routing data to components within a workflow. To address these problems, we employ Virtual Data Assembly Lines (VDAL), a modeling paradigm that can eliminate most shims and reduce wiring complexity. We show how a VDAL design can be implemented using existing XML technologies and how static analysis can provide significant help to scientists during workflow design and evolution, e.g., by displaying actor dependencies or by detecting so-called unproductive actors.
Daniel Zinn, Shawn Bowers, Bertram Ludäscher
ICDE1
2010 Parallelizing XML data-streaming workflows via MapReduce
Daniel Zinn, Shawn Bowers, Sven Köhler 0003, Bertram Ludäscher
J. Comput. Syst. Sci.1
2009 X-CSR: Dataflow Optimization for Distributed XML Process Pipelines
abstract
Abstract — XML process networks are a simple, yet power-ful programming paradigm for loosely coupled, coarse-grained dataflow applications such as data-centric scientific workflows. We describe a framework called ∆-XML that is well-suited for applications in which pipelines of data processors modify parts (“deltas”) of XML data collections while keeping the overall collection structure intact. We show how to optimize the execution of ∆-XML process networks by minimizing the data shipping cost in distributed settings. This X-CSR1 optimization employs static type inference based on XML Schema to determine the XML stream fragments that are relevant to a processor, allowing irrelevant fragments to be bypassed (“shipped”) to downstream pipeline steps. Finally, we present evaluation results for a real-world scientific workflow, which shows the practical feasibility of X-CSR. A long version of this paper is available as [1]. I.
Daniel Zinn, Shawn Bowers, Timothy M. McPhillips, Bertram Ludäscher
ICDE1
2009 Scientific workflow design for mere mortals
Timothy M. McPhillips, Shawn Bowers, Daniel Zinn, Bertram Ludäscher
Future Gener. Comput. Syst.3
2007 ConceptDoppler: a weather tracker for internet censorship
abstract
Article Share on ConceptDoppler: a weather tracker for internet censorshipCCS '07: Proceedings of the 14th ACM conference on Computer and communications securityOctober 2007 Pages 352–365https://doi.org/10.1145/1315245.1315290Online:28 October 2007Publication History 42citation1,097DownloadsMetricsTotal Citations42Total Downloads1,097Last 12 Months115Last 6 weeks7 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jedidiah R. Crandall, Daniel Zinn, Michael Byrd, Earl T. Barr, Rich East
CCS2
2007 Modeling and Querying Vague Spatial Objects Using Shapelets
Daniel Zinn, Jim Bosch, Michael Gertz 0001
VLDB1
2005 Processing Top-N Queries in P2P-based Web Integration Systems with Probabilistic Guarantees
Katja Hose, Marcel Karnstedt, Kai-Uwe Sattler, Daniel Zinn
WebDB4