Carl-Christian Kanne

dblp:k/CarlChristianKanne · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
0since 2021 · last 2012
0009-0009-0234-3289ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19 · 5 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
11 papers
Query processing and optimization · 54% Data models and query languages · 29% Distributed and cloud data management · 8%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Theoretical computer science
1 paper
Approximation and online algorithms · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 100%

Topics — the 13 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
partial evaluation
0.112012
Declarative error management for robust data-intensive applications · SIGMOD Conference 2012
Query processing and optimization
query execution
0.112012
Declarative error management for robust data-intensive applications · SIGMOD Conference 2012
Query processing and optimization › query optimization
algebraic query optimization
0.122006
Algebraic Optimization of Nested XPath Expressions · ICDE 2006
Full-fledged Algebraic XPath Processing in Natix · ICDE 2005
Data models and query languages
XML query languages
0.122006
Algebraic Optimization of Nested XPath Expressions · ICDE 2006
Full-fledged Algebraic XPath Processing in Natix · ICDE 2005
Query processing and optimization
cardinality estimation
0.112010
Histograms reloaded: the merits of bucket diversity · SIGMOD Conference 2010
Query processing and optimization › cardinality estimation
histogram
0.112010
Histograms reloaded: the merits of bucket diversity · SIGMOD Conference 2010
Compilers and program optimization › program transformation
source-to-source transformation
0.112010
TransScale: Scalability transformations for declarative applications · ICDE 2010
Approximation and online algorithms
approximation algorithms
0.112006
A Linear Time Algorithm for Optimal Tree Sibling Partitioning and Approximation Algorithms in Natix · VLDB 2006
Graph data management
path query
0.112005
Cost-Sensitive Reordering of Navigational Primitives · SIGMOD Conference 2005
Data models and query languages › XML query languages
XPath
0.112005
Full-fledged Algebraic XPath Processing in Natix · ICDE 2005
Data models and query languages › XML data management
XML database
0.122005
Anatomy of a native XML base management system · VLDB J. 2002
Full-fledged Algebraic XPath Processing in Natix · ICDE 2005
Data models and query languages › XML data management
XML data model
0.012002
Anatomy of a native XML base management system · VLDB J. 2002
Storage systems › data management
XML storage
0.012000
Efficient Storage of XML Data · ICDE 2000

Methods — techniques the papers use, named apart from their topics

program partitioning · 0.2distribution transformation · 0.2rule compilation · 0.2query planning · 0.2declarative error handling · 0.1linear-time algorithm · 0.1algebraic equivalences · 0.1sequential scan · 0.1iterator-based execution · 0.1asynchronous i/o · 0.1parameterizable split algorithm · 0.0
YearPublicationVenuePosition
2012 Declarative error management for robust data-intensive applications
abstract
We present an approach to declaratively manage run-time errors in data-intensive applications. When large volumes of raw data meet complex third-party libraries, deterministic run-time errors become likely, and existing query processors typically stop without returning a result when a run-time error occurs. The ability to degrade gracefully in the presence of run-time errors, and partially execute jobs, is typically limited to specific operators such as bulkloading.
Carl-Christian Kanne, Vuk Ercegovac
SIGMOD Conference1
2011 Querying Versioned Software Repositories
Dietrich Christopeit, Michael H. Böhlen, Carl-Christian Kanne, Arturas Mazeika
ADBIS3
2011 Declarative Serializable Snapshot Isolation
Christian Tilgner, Boris Glavic, Michael H. Böhlen, Carl-Christian Kanne
ADBIS4
2011 Demaq/Transscale: Automated distribution and scalability for declarative applications
Alexander Böhm 0002, Carl-Christian Kanne
Inf. Syst.2
2011 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis
Kevin S. Beyer, Vuk Ercegovac, Rainer Gemulla, Andrey Balmin, Mohamed Y. Eltabakh, Carl-Christian Kanne, Fatma Özcan 0001, Eugene J. Shekita
Proc. VLDB Endow.6
2010 TransScale: Scalability transformations for declarative applications
abstract
The goal of the Demaq/TransScale system is to automate the distribution of complex application processes to large numbers of hosts. We implement distribution as a source-level transformation that turns the distribution-unaware application specification for a single host into a set of programs that can be executed on the various machines of a cluster.
Alexander Böhm 0002, Erich Marth, Carl-Christian Kanne
ICDE3
2010 Histograms reloaded: the merits of bucket diversity
abstract
Virtually all histograms store for each bucket the number of distinct values it contains and their average frequency. In this paper, we question this paradigm. We start out by investigating the estimation precision of three commercial database systems which also follow the above paradigm. It turns out that huge errors are quite common. We then introduce new bucket types and investigate their accuracy when building optimal histograms with them. The results are ambiguous. There is no clear winner among the bucket types. At this point, we (1) switch to heterogeneous histograms, where different buckets of the same histogram possibly are of different types, and (2) design more bucket types. The nice consequence of introducing heterogeneous histograms is that we can guarantee decent upper error bounds while at the same time heterogeneous histograms require far less space than homogeneous histograms.
Carl-Christian Kanne, Guido Moerkotte
SIGMOD Conference1
2009 Processes Are Data: A Programming Model for Distributed Applications
Alexander Böhm 0002, Carl-Christian Kanne
WISE2
2008 The Demaq system: declarative development of distributed applications
abstract
The goal of the Demaq project is to investigate a novel way of thinking about distributed applications that are based on the asynchronous exchange of XML messages. Unlike today's solutions that rely on imperative programming languages and multi-tiered application servers, Demaq uses a declarative language for implementing the application logic as a set of rules. A rule compiler transforms the application specifications into execution plans against the message history. The plans are evaluated using our optimized runtime engine. This allows us to leverage existing knowledge about declarative query processing for optimizing distributed applications.
Alexander Böhm 0002, Erich Marth, Carl-Christian Kanne
SIGMOD Conference3
2007 Demaq: A Foundation for Declarative XML Message Processing
Alexander Böhm 0002, Carl-Christian Kanne, Guido Moerkotte
CIDR2
2006 A Declarative Control Language for Dependable XML Message Queues
abstract
We present a novel approach for the implementation of efficient and dependable Web service engines (WSEs). A WSE instance represents a single node in a distributed network of participants that communicate using XML messages. We introduce a fully declarative language custom-tailored to XML message processing that allows to specify business processes in a concise manner. To support the efficient and reliable evaluation of our language, we show how to augment a native, transactional XML data store with efficient and reliable XML message queues.
Alexander Böhm 0002, Carl-Christian Kanne, Guido Moerkotte
ARES2
2006 Natix Visual Interfaces
Alexander Böhm 0002, Matthias Brantner, Carl-Christian Kanne, Norman May, Guido Moerkotte
EDBT3
2006 Algebraic Optimization of Nested XPath Expressions
abstract
The XPath language incorporates powerful primitives for formulating queries containing nested subexpressions which are existentially or universally quantified. However, even the best published approaches for evaluating XPath have unsatisfactory performance when applied to nested queries. We examine optimization techniques that unnest complex XPath queries. For this purpose, we classify XPath expressions particularly with regard to properties that are relevant for unnesting. We present algebraic equivalences that transform nested expressions into unnested expressions. In our experiments we compare the evaluation times with existing XPath evaluators and the naive evaluation.
Matthias Brantner, Carl-Christian Kanne, Guido Moerkotte, Sven Helmer
ICDE2
2006 A Linear Time Algorithm for Optimal Tree Sibling Partitioning and Approximation Algorithms in Natix
Carl-Christian Kanne, Guido Moerkotte
VLDB1
2005 Full-fledged Algebraic XPath Processing in Natix
abstract
We present the first complete translation of XPath into an algebra, paving the way for a comprehensive, state-of-the-art XPath (and later on, XQuery) compiler based on algebraic optimization techniques. Our translation includes all XPath features such as nested expressions, position-based predicates and node-set functions. The translated algebraic expressions can be executed using the proven, scalable, iterator-based approach, as we demonstrate in form of a corresponding physical algebra in our native XML DBMS Natix. A first glance at performance results shows that even without further optimization of the expressions, we provide a competitive evaluation technique for XPath queries.
Matthias Brantner, Sven Helmer, Carl-Christian Kanne, Guido Moerkotte
ICDE3
2005 Cost-Sensitive Reordering of Navigational Primitives
abstract
We present a method to evaluate path queries based on the novel concept of partial path instances. Our method (1) maximizes performance by means of sequential scans or asynchronous I/O, (2) does not require a special storage format, (3) relies on simple navigational primitives on trees, and (4) can be complemented by existing logical and physical optimizations such as duplicate elimination, duplicate prevention and path rewriting.We use a physical algebra which separates those navigation operations that require I/O from those that do not. All I/O operations necessary for the evaluation of a path are isolated in a single operator, which may employ efficient I/O scheduling strategies such as sequential scans or asynchronous I/O.Performance results for queries from the XMark benchmark show that reordering the navigation operations can increase performance up to a factor of four.
Carl-Christian Kanne, Matthias Brantner, Guido Moerkotte
SIGMOD Conference1
2004 Timestamp-Based Protocols for Synchronizing Access on XML Documents
Sven Helmer, Carl-Christian Kanne, Guido Moerkotte
DEXA2
2002 Optimized Translation of XPath into Algebraic Expressions Parameterized by Programs Containing Navigational Primitives
abstract
We propose a new approach for the efficient evaluation of XPath expressions. This is important, since XPath is not only used as a simple, stand-alone query language, but is also an essential ingredient of XQuery and XSLT. The main idea of our approach is to translate XPath into algebraic expressions parameterized with programs. These programs are mainly built from navigational primitives like accessing the first child or the next sibling. The goals of the approach are: 1) to enable pipelined evaluation, 2) to avoid producing duplicate (intermediate) result nodes, 3) to visit as few document nodes as possible, and 4) to avoid visiting nodes more than once. This improves the existing approaches, because our method is highly efficient.
Sven Helmer, Carl-Christian Kanne, Guido Moerkotte
WISE2
2002 Anatomy of a native XML base management system
Thorsten Fiebig, Sven Helmer, Carl-Christian Kanne, Guido Moerkotte, Julia Neumann, Robert Schiele, Till Westmann
VLDB J.3
2000 Efficient Storage of XML Data
abstract
We introduce NATIX, an efficient, native repository for storing, retrieving and managing tree-structured large objects, preferably XML documents. In contrast to traditionallarge object (LOB) managers, we do not split at arbitrary byte positions but take the semantics of the underlying tree structure of XML documents into account.\nOur parameterizable split algorithm dynamically maintains physical records of size smaller than a page which contain sets of connected tree nodes. This not only improves efficiency by clustering subtrees but also facilitates their compact representation. Existing approaches to store XML documents either use flat files or map every single tree node onto a separate physical record. The increased flexibility of our approach results in higher efficiency. Performance measurements validate this claim.
Carl-Christian Kanne, Guido Moerkotte
ICDE1
1999 Electronic Biochemical Pathways
Carl-Christian Kanne, Falk Schreiber, Dietrich Trümbach
GD1