EDBT 2026 Demo / reviewers in the wild / expert
Ceriel J. H. Jacobs
dblp:73/6019
· DBLP profile ↗
19ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0002-4692-7245ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10Software engineering, systems software and programming languages · 4Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Graph data management · 38% Indexing and storage engines · 33% Knowledge graphs · 22% | |
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 57% Distributed systems · 30% Memory systems · 12% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 50% Runtime systems and virtual machines · 50% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Graph data management › graph storage
knowledge graph storage |
0.4 | 1 | 2020 | Adaptive Low-level Storage of Very Large Knowledge Graphs · WWW 2020 |
Indexing and storage engines
column store |
0.2 | 1 | 2016 | Column-Oriented Datalog Materialization for Large Knowledge Graphs · AAAI 2016 |
Parallel and multicore computing
parallel programming models |
0.2 | 4 | 2010 | Satin: A high-level and efficient grid programming model · ACM Trans. Program. Lang. Syst. 2010 Efficient Java RMI for parallel programming · ACM Trans. Program. Lang. Syst. 2001 Source-level global optimizations for fine-grain distributed shared memory systems · PPoPP 2001 |
Indexing and storage engines › storage management
storage architecture |
0.1 | 1 | 2020 | Adaptive Low-level Storage of Very Large Knowledge Graphs · WWW 2020 |
Parallel and multicore computing › load balancing
dynamic load balancing |
0.1 | 1 | 2010 | Satin: A high-level and efficient grid programming model · ACM Trans. Program. Lang. Syst. 2010 |
Distributed systems
grid computing |
0.1 | 1 | 2010 | Satin: A high-level and efficient grid programming model · ACM Trans. Program. Lang. Syst. 2010 |
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling |
0.1 | 1 | 2010 | Satin: A high-level and efficient grid programming model · ACM Trans. Program. Lang. Syst. 2010 |
Machine learning and data management
inference optimization |
0.1 | 1 | 2016 | Column-Oriented Datalog Materialization for Large Knowledge Graphs · AAAI 2016 |
Memory systems › shared memory
distributed shared memory |
0.1 | 3 | 2001 | Source-level global optimizations for fine-grain distributed shared memory systems · PPoPP 2001 A Task- and Data-Parallel Programming Language Based on Shared Objects · ACM Trans. Program. Lang. Syst. 1998 Performance Evaluation of the Orca Shared-Object System · ACM Trans. Comput. Syst. 1998 |
Distributed systems
fault tolerance |
0.0 | 1 | 2010 | Satin: A high-level and efficient grid programming model · ACM Trans. Program. Lang. Syst. 2010 |
Distributed systems
remote procedure call |
0.0 | 1 | 2001 | Efficient Java RMI for parallel programming · ACM Trans. Program. Lang. Syst. 2001 |
Memory systems › cache coherence
cache coherence protocol |
0.0 | 1 | 1998 | Performance Evaluation of the Orca Shared-Object System · ACM Trans. Comput. Syst. 1998 |
Distributed systems › replication › data replication
object replication |
0.0 | 1 | 1998 | A Task- and Data-Parallel Programming Language Based on Shared Objects · ACM Trans. Program. Lang. Syst. 1998 |
Distributed systems
object sharing |
0.0 | 1 | 1998 | A Task- and Data-Parallel Programming Language Based on Shared Objects · ACM Trans. Program. Lang. Syst. 1998 |
Parallel and multicore computing › parallel programming models
task and data parallelism |
0.0 | 1 | 1998 | A Task- and Data-Parallel Programming Language Based on Shared Objects · ACM Trans. Program. Lang. Syst. 1998 |
Storage systems › distributed storage
consistency semantics |
0.0 | 1 | 1998 | A Task- and Data-Parallel Programming Language Based on Shared Objects · ACM Trans. Program. Lang. Syst. 1998 |
Distributed systems
group communication |
0.0 | 1 | 1998 | Performance Evaluation of the Orca Shared-Object System · ACM Trans. Comput. Syst. 1998 |
Distributed systems › distributed coordination
message ordering |
0.0 | 1 | 1998 | Performance Evaluation of the Orca Shared-Object System · ACM Trans. Comput. Syst. 1998 |
Parallel and multicore computing
parallel programming runtimes |
0.0 | 1 | 1998 | A Task- and Data-Parallel Programming Language Based on Shared Objects · ACM Trans. Program. Lang. Syst. 1998 |
Methods — techniques the papers use, named apart from their topics
topology-aware storage · 0.4interlinked data structures · 0.4proactive caching · 0.2datalog · 0.2column-oriented layout · 0.2speculative parallelism · 0.1cluster-aware scheduling · 0.1asynchronous exceptions · 0.1static analysis · 0.1object-graph aggregation · 0.1native compilation · 0.1computation migration · 0.1compile-time type specialization · 0.1access-check batching · 0.1performance measurement · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Rewrite or Not Rewrite? ML-Based Algorithm Selection for Datalog Query Answering on Knowledge GraphsabstractQuery-driven reasoning techniques with Datalog rules, like Magic Sets (MS), are ideal for implementing query answering on Knowledge Graphs (KGs). For some queries, executing a rewriting procedure like MS is the best choice, but for others a non-rewriting procedure like Query-subquery (QSQ) can be faster. Choosing beforehand which procedure should be used is not trivial and mistakes can be costly. To address this problem, we describe a first-of-its-kind method that builds a Machine Learning (ML) model to predict whether a query should be answered with MS or with QSQ. Experiments on several well-known KGs show that our method can return accurate predictions, and this leads to a significant reduction of the response time of query answering. Unmesh Joshi, Ceriel J. H. Jacobs, Jacopo Urbani |
ECAI | 2 |
| 2020 | Adaptive Low-level Storage of Very Large Knowledge GraphsabstractThe increasing availability and usage of Knowledge Graphs (KGs) on the Web calls for scalable and general-purpose solutions to store this type of data structures. We propose Trident, a novel storage architecture for very large KGs on centralized systems. Trident uses several interlinked data structures to provide fast access to nodes and edges, with the physical storage changing depending on the topology of the graph to reduce the memory footprint. In contrast to single architectures designed for single tasks, our approach offers an interface with few low-level and general-purpose primitives that can be used to implement tasks like SPARQL query answering, reasoning, or graph analytics. Our experiments show that Trident can handle graphs with 1011 edges using inexpensive hardware, delivering competitive performance on multiple workloads. Jacopo Urbani, Ceriel J. H. Jacobs |
WWW | 2 |
| 2019 | VLog: A Rule Engine for Knowledge Graphs
David Carral, Irina Dragoste, Larry González, Ceriel J. H. Jacobs, Markus Krötzsch, Jacopo Urbani |
ISWC (2) | 4 |
| 2016 | Column-Oriented Datalog Materialization for Large Knowledge GraphsabstractThe evaluation of Datalog rules over large Knowledge Graphs (KGs) is essential for many applications. In this paper, we present a new method of materializing Datalog inferences, which combines a column-based memory layout with novel optimization methods that avoid redundant inferences at runtime. The pro-active caching of certain subqueries further increases efficiency. Our empirical evaluation shows that this approach can often match or even surpass the performance of state-of-the-art systems, especially under restricted resources. Jacopo Urbani, Ceriel J. H. Jacobs, Markus Krötzsch |
AAAI | 2 |
| 2015 | Cashmere: Heterogeneous Many-Core ComputingabstractNew generations of many-core hardware become available frequently and are typically attractive extensions for data-centers because of power-consumption and performance benefits. As a result, supercomputers and clusters are becoming heterogeneous and start to contain a variety of many-core devices. Obtaining performance from a homogeneous cluster-computer is already challenging, but achieving it from a heterogeneous cluster is even more demanding. Related work primarily focuses on homogeneous many-core clusters. In this paper we present Cashmere, a programming system for heterogeneous many-core clusters. Cashmere is a tight integration of two existing systems: Satin is a programming system that provides a divide- and-conquer programming model with automatic load-balancing and latency-hiding, while Many-Core Levels is a programming system that provides a powerful methodology to optimize computational kernels for varying types of many-core hardware. We evaluate our system with several classes of applications and show that Cashmere achieves high performance and good scalability. The efficiency of heterogeneous executions is comparable to the homogeneous runs and is >90% in three out of four applications. Pieter Hijma, Ceriel J. H. Jacobs, Rob van Nieuwpoort, Henri E. Bal |
IPDPS | 2 |
| 2015 | Stepwise-refinement for performance: a methodology for many-core programmingabstractSummary Many‐core hardware is targeted specifically at obtaining high performance, but reaching high performance is often challenging because hardware‐specific details have to be taken into account. Although there are many programming systems that try to alleviate many‐core programming, some providing a high‐level language, others providing a low‐level language for control, none of these systems have a clear and systematic methodology as a foundation. In this article, we proposestepwise‐refinement for performance: a novel, clear, and structured methodology for obtaining high performance on many‐cores. We present a system that supports this methodology, offers multiple levels of abstraction to provide programmers a trade‐off between high‐level and low‐level programming, and provides programmers detailed performance feedback. We evaluate our methodology with several widely varying compute kernels on two different many‐core architectures: a Graphical Processing Unit (GPU) and the Xeon Phi. We show that our methodology gives insight in the performance, and that in almost all cases, we gain a substantial performance improvement using our methodology. Copyright © 2015 John Wiley & Sons, Ltd. Pieter Hijma, Rob van Nieuwpoort, Ceriel J. H. Jacobs, Henri E. Bal |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | AJIRA: A Lightweight Distributed Middleware for MapReduce and Stream ProcessingabstractCurrently, MapReduce is the most popular programming model for large-scale data processing and this motivated the research community to improve its efficiency either with new extensions, algorithmic optimizations, or hardware. In this paper we address two main limitations of MapReduce: one relates to the model's limited expressiveness, which prevents the implementation of complex programs that require multiple steps or iterations. The other relates to the efficiency of its most popular implementations (e.g., Hadoop), which provide good resource utilization only for massive volumes of input, operating sub optimally for smaller or rapidly changing input. To address these limitations, we present AJIRA, a new middleware designed for efficient and generic data processing. At a conceptual level, AJIRA replaces the traditional map/reduce primitives by generic operators that can be dynamically allocated, allowing the execution of more complex batch and stream processing jobs. At a more technical level, AJIRA adopts a distributed, multi-threaded architecture that strives at minimizing overhead for non-critical functionality. These characteristics allow AJIRA to be used as a single programming model for both batch and stream processing. To this end, we evaluated its performance against Hadoop, Spark, Esper, and Storm, which are state of the art systems for both batch and stream processing. Our evaluation shows that AJIRA is competitive in a wide range of scenarios both in terms of processing time and scalability, making it an ideal choice where flexibility, extensibility, and the processing of both large and dynamic data with a single programming model are either desirable or even mandatory requirements. Jacopo Urbani, Alessandro Margara, Ceriel J. H. Jacobs, Spyros Voulgaris, Henri E. Bal |
ICDCS | 3 |
| 2013 | DynamiTE: Parallel Materialization of Dynamic RDF Data
Jacopo Urbani, Alessandro Margara, Ceriel J. H. Jacobs, Frank van Harmelen, Henri E. Bal |
ISWC (1) | 3 |
| 2012 | Generating synchronization statements in divide-and-conquer programs
Pieter Hijma, Rob van Nieuwpoort, Ceriel J. H. Jacobs, Henri E. Bal |
Parallel Comput. | 3 |
| 2010 | What Is the Price of Simplicity? - A Cross-Platform Evaluation of the SAGA API
Mathijs den Burger, Ceriel J. H. Jacobs, Thilo Kielmann, André Merzky, Ole Weidner, Hartmut Kaiser |
Euro-Par (1) | 2 |
| 2010 | Satin: A high-level and efficient grid programming modelabstractComputational grids have an enormous potential to provide compute power. However, this power remains largely unexploited today for most applications, except trivially parallel programs. Developing parallel grid applications simply is too difficult. Grids introduce several problems not encountered before, mainly due to the highly heterogeneous and dynamic computing and networking environment. Furthermore, failures occur frequently, and resources may be claimed by higher-priority jobs at any time. In this article, we solve these problems for an important class of applications: divide-and-conquer. We introduce a system called Satin that simplifies the development of parallel grid applications by providing a rich high-level programming model that completely hides communication. All grid issues are transparently handled in the runtime system, not by the programmer. Satin's programming model is based on Java, features spawn-sync primitives and shared objects, and uses asynchronous exceptions and an abort mechanism to support speculative parallelism. To allow an efficient implementation, Satin consistently exploits the idea that grids are hierarchically structured. Dynamic load-balancing is done with a novel cluster-aware scheduling algorithm that hides the long wide-area latencies by overlapping them with useful local work. Satin's shared object model lets the application define the consistency model it needs. If an application needs only loose consistency, it does not have to pay high performance penalties for wide-area communication and synchronization. We demonstrate how grid problems such as resource changes and failures can be handled transparently and efficiently. Finally, we show that adaptivity is important in grids. Satin can increase performance considerably by adding and removing compute resources automatically, based on the application's requirements and the utilization of the machines and networks in the grid. Using an extensive evaluation on real grids with up to 960 cores, we demonstrate that it is possible to provide a simple high-level programming model for divide-and-conquer applications, while achieving excellent performance on grids. At the same time, we show that the divide-and-conquer model scales better on large systems than the master-worker approach, since it has no single central bottleneck. Rob van Nieuwpoort, Gosia Wrzesinska, Ceriel J. H. Jacobs, Henri E. Bal |
ACM Trans. Program. Lang. Syst. | 3 |
| 2005 | Ibis: a flexible and efficient Java-based Grid programming environmentabstractAbstract In computational Grids, performance‐hungry applications need to simultaneously tap the computational power of multiple, dynamically available sites. The crux of designing Grid programming environments stems exactly from the dynamic availability of compute cycles: Grid programming environments (a) need to beportableto run on as many sites as possible, (b) they need to beflexibleto cope with different network protocols and dynamically changing groups of compute nodes, while (c) they need to provideefficient(local) communication that enables high‐performance computing in the first place. Existing programming environments are either portable (Java), or flexible (Jini, Java Remote Method Invocation or (RMI)), or they are highly efficient (Message Passing Interface). No system combines all three properties that are necessary for Grid computing. In this paper, we present Ibis, a new programming environment that combines Java's ‘run everywhere’ portability both with flexible treatment of dynamically available networks and processor pools, and with highly efficient, object‐based communication. Ibis can transfer Java objects very efficiently by combining streaming object serialization with a zero‐copy protocol. Using RMI as a simple test case, we show that Ibis outperforms existing RMI implementations, achieving up to nine times higher throughputs with trees of objects. Copyright © 2005 John Wiley & Sons, Ltd. Rob van Nieuwpoort, Jason Maassen, Gosia Wrzesinska, Rutger F. H. Hofman, Ceriel J. H. Jacobs, Thilo Kielmann, Henri E. Bal |
Concurr. Pract. Exp. | 5 |
| 2005 | Object combining: a new aggressive optimization for object intensive programsabstractAbstract Object combining tries to put objects together that have roughly the same life times in order to reduce strain on the memory manager and to reduce the number of pointer indirections during a program's execution. Object combining works by appending the fields of one object to another, allowing allocation and freeing of multiple objects with a single heap (de)allocation. Unlike object inlining, which will only optimize objects where one has a (unique) pointer to another, our optimization also works if there is no such relation. Object inlining also directly replaces the pointer by the inlined object's fields. Object combining leaves the pointer in place to allow more combining. Elimination of the pointer accesses is implemented in a separate compiler optimization pass. Unlike previous object inlining systems, reference field overwrites are allowed and handled, resulting in much more aggressive optimization. Our object combining heuristics also allow unrelated objects to be combined, for example, those allocated inside a loop; recursive data structures (linked lists, trees) can be allocated several at a time and objects that are always used together can be combined. As Java explicitly permits code to be loaded at runtime and allows the new code to contribute to a running computation, we do not require a closed‐world assumption to enable these optimizations (but it will increase performance). The main focus of object combining in this paper is on reducing object (de)allocation overhead, by reducing both garbage collection work and the number of object allocations. Reduction of memory management overhead causes execution time to be reduced by up to 35%. Indirection removal further reduces execution time by up to 6%. Copyright © 2005 John Wiley & Sons, Ltd. Ronald Veldema, Ceriel J. H. Jacobs, Rutger F. H. Hofman, Henri E. Bal |
Concurr. Pract. Exp. | 2 |
| 2001 | Source-level global optimizations for fine-grain distributed shared memory systemsabstractThis paper describes and evaluates the use of aggressive static analysis in Jackal, a fine-grain Distributed Shared Memory (DSM) system for Java. Jackal uses an optimizing, source-level compiler rather than the binary rewriting techniques employed by most other fine-grain DSM systems. Source-level analysis makes existing access-check optimizations (e.g., access-check batching) more effective and enables two novel fine-grain DSM optimizations: object-graph aggregation and automatic computation migration. Ronald Veldema, Rutger F. H. Hofman, Raoul Bhoedjang, Ceriel J. H. Jacobs, Henri E. Bal |
PPoPP | 4 |
| 2001 | Efficient Java RMI for parallel programmingabstractJava offers interesting opportunities for parallel computing. In particular, Java Remote Method Invocation (RMI) provides a flexible kind of remote procedure call (RPC) that supports polymorphism. Sun's RMI implementation achieves this kind of flexibility at the cost of a major runtime overhead. The goal of this article is to show that RMI can be implemented efficiently, while still supporting polymorphism and allowing interoperability with Java Virtual Machines (JVMs). We study a new approach for implementing RMI, using a compiler-based Java system called Manta. Manta uses a native (static) compiler instead of a just-in-time compiler. To implement RMI efficiently, Manta exploits compile-time type information for generating specialized serializers. Also, it uses an efficient RMI protocol and fast low-level communication protocols.A difficult problem with this approach is how to support polymorphism and interoperability. One of the consequences of polymorphism is that an RMI implementation must be able to download remote classes into an application during runtime. Manta solves this problem by using a dynamic bytecode compiler, which is capable of compiling and linking bytecode into a running application. To allow interoperability with JVMs, Manta also implements the Sun RMI protocol (i.e., the standard RMI protocol), in addition to its own protocol.We evaluate the performance of Manta using benchmarks and applications that run on a 32-node Myrinet cluster. The time for a null-RMI (without parameters or a return value) of Manta is 35 times lower than for the Sun JDK 1.2, and only slightly higher than for a C-based RPC protocol. This high performance is accomplished by pushing almost all of the runtime overhead of RMI to compile time. We study the performance differences between the Manta and the Sun RMI protocols in detail. The poor performance of the Sun RMI protocol is in part due to an inefficient implementation of the protocol. To allow a fair comparison, we compiled the applications and the Sun RMI protocol with the native Manta compiler. The results show that Manta's null-RMI latency is still eight times lower than for the compiled Sun RMI protocol and that Manta's efficient RMI protocol results in 1.8 to 3.4 times higher speedups for four out of six applications. Jason Maassen, Rob van Nieuwpoort, Ronald Veldema, Henri E. Bal, Thilo Kielmann, Ceriel J. H. Jacobs, Rutger F. H. Hofman |
ACM Trans. Program. Lang. Syst. | 6 |
| 1998 | Performance Evaluation of the Orca Shared-Object SystemabstractOrca is a portable, object-based distributed shared memory (DSM) system. This article studies and evaluates the design choices made in the Orca system and compares Orca with other DSMs. The article gives a quantitative analysis of Orca's coherence protocol (based on write-updates with function shipping), the totally ordered group communication protocol, the strategy for object placement, and the all-software, user-space architecture. Performance measurements for 10 parallel applications illustrate the trade-offs made in the design of Orca and show that essentially the right design decisions have been made. A write-update protocol with function shipping is effective for Orca, especially since it is used in combination with techniques that avoid replicating objects that have a low read/write ratio. The overhead of totally ordered group communication on application performance is low. The Orca system is able to make near-optimal decisions for object placement and replication. In addition, the article compares the performance of Orca with that of a page-based DSM (TreadMarks) and another object-based DSM (CRL). It also analyzes the communication overhead of the DSMs for several applications. All performance measurements are done on a 32-node Pentium Pro cluster with Myrinet and Fast Ethernet networks. The results show that Orca programs send fewer messages and less data than the TreadMarks and CRL programs and obtain better speedups. Henri E. Bal, Raoul Bhoedjang, Rutger F. H. Hofman, Ceriel J. H. Jacobs, Koen Langendoen, Tim Rühl |
ACM Trans. Comput. Syst. | 4 |
| 1998 | A Task- and Data-Parallel Programming Language Based on Shared ObjectsabstractMany programming languages support either task parallelism, but few languages provide a uniform framework for writing applications that need both types of parallelism or data parallelism. We present a programming language and system that integrates task and data parallelism using shared objects. Shared objects may be stored on one processor or may be replicated. Objects may also be partitioned and distributed on several processors.Task parallelism is achieved by forking processes remotely and have them communicate and synchronize through objects. Data parallelism is achieved by executing operations on partitioned objects in parallel. Writing task-and data-parallel applications with shared objects has several advantages. Programmers use the objects as if they were stored in a memory common to all processors. On distributed-memory machines, if objects are remote, replicated, or partitioned, the system takes care of many low-level details such as data transfers and consistency semantics. In this article, we show how to write task-and data-parallel programs with our shared object model. We also desribe a portable implementation of the model. To assess the performance of the system, we wrote several applications that use task and data parallelism and excuted them on a collection of Pentium Pros connected by Myrinet. The performance of these applications is also discussed in this article. Saniya Ben Hassen, Henri E. Bal, Ceriel J. H. Jacobs |
ACM Trans. Program. Lang. Syst. | 3 |
| 1997 | Performance of a High-Level Parallel Language on a High-Speed Network
Henri E. Bal, Raoul Bhoedjang, Rutger F. H. Hofman, Ceriel J. H. Jacobs, Koen Langendoen, Tim Rühl, Kees Verstoep |
J. Parallel Distributed Comput. | 4 |
| 1988 | A Programmer-friendly LL(1) Parser GeneratorabstractAbstract LL(1) grammars have the conceptual and practical advantage that they allow the compiler writer to view the grammar as a program; this allows a more natural positioning of semantic actions and a simple attribute mechanism. Resulting parsers can be constructed that achieve fully automatic error‐recovery, which allows the compiler writer to ignore totally the issue of syntax errors. Measurement shows that such parsers can be reasonably efficient. Dick Grune, Ceriel J. H. Jacobs |
Softw. Pract. Exp. | 2 |