John Liagouris

dblp:95/9170 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-4692-3022ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-authorSystems, architecture and hardware · 3 · 1 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Cloud and datacenter computing · 43% Distributed systems · 28% Performance modeling and evaluation · 14%
Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 29% Graph data management · 24% Spatial and temporal data management · 23%
Network and information security
4 papers
Cryptographic protocols and secure computation · 77% Privacy and data protection · 23%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic protocols and secure computation
secure multiparty computation
2.232025
ORQ: Complex Analytics on Private Data with Strong Security Guarantees · SOSP 2025
TVA: A multi-party computation system for secure and expressive time series analytics · USENIX Security Symposium 2023
SECRECY: Secure collaborative analytics in untrusted clouds · NSDI 2023
Cloud and datacenter computing › microservices
microservice applications
1.012026
Slowpoke: End-to-end Throughput Optimization Modeling for Microservice Applications · NSDI 2026
Query processing and optimization
secure query processing
0.912025
ORQ: Complex Analytics on Private Data with Strong Security Guarantees · SOSP 2025
Performance modeling and evaluation
benchmarking
0.612022
A new benchmark harness for systematic and robust evaluation of streaming state stores · EuroSys 2022
Graph data management
RDF data management
0.622019
SRX: efficient management of spatial RDF data · VLDB J. 2019
An Effective Encoding Scheme for Spatial RDF Data · Proc. VLDB Endow. 2014
Cloud and datacenter computing
cluster resource management and scheduling
0.422019
Three steps is all you need: fast, accurate, automatic scaling decisions for distributed streaming dataflows · OSDI 2018
Megaphone: Latency-conscious state migration for distributed streaming dataflows · Proc. VLDB Endow. 2019
Cloud and datacenter computing
cluster computing framework
0.412019
Lineage stash: fault tolerance off the critical path · SOSP 2019
Parallel and multicore computing › parallel computation models
distributed dataflow
0.412019
Megaphone: Latency-conscious state migration for distributed streaming dataflows · Proc. VLDB Endow. 2019
Distributed systems
fault tolerance
0.412019
Lineage stash: fault tolerance off the critical path · SOSP 2019
Distributed systems
state migration
0.412019
Megaphone: Latency-conscious state migration for distributed streaming dataflows · Proc. VLDB Endow. 2019
Distributed systems
stream processing
0.412019
Megaphone: Latency-conscious state migration for distributed streaming dataflows · Proc. VLDB Endow. 2019
Cloud and datacenter computing
autoscaling
0.312018
Three steps is all you need: fast, accurate, automatic scaling decisions for distributed streaming dataflows · OSDI 2018
Electronic design automation › timing analysis
critical path analysis
0.312018
SnailTrail: Generalizing Critical Paths for Online Analysis of Distributed Dataflows · NSDI 2018
Distributed systems › stream processing
distributed stream processing
0.312018
Three steps is all you need: fast, accurate, automatic scaling decisions for distributed streaming dataflows · OSDI 2018
Performance modeling and evaluation
queueing models
0.312026
Slowpoke: End-to-end Throughput Optimization Modeling for Microservice Applications · NSDI 2026
Cloud and datacenter computing › datacenter operations
datacenter monitoring
0.312017
Online Reconstruction of Structural Information from Datacenter Logs · EuroSys 2017
Distributed systems › observability › distributed monitoring
distributed tracing
0.312017
Online Reconstruction of Structural Information from Datacenter Logs · EuroSys 2017
Cloud and datacenter computing
log analysis
0.312017
Online Reconstruction of Structural Information from Datacenter Logs · EuroSys 2017
Data integration and cleaning
data provenance
0.212016
Explaining Outputs in Modern Data Analytics · Proc. VLDB Endow. 2016
Graph data management
graph indexing
0.212016
graphVizdb: A scalable platform for interactive large graph visualization · ICDE 2016
Graph data management
graph visualization
0.212016
graphVizdb: A scalable platform for interactive large graph visualization · ICDE 2016
Query processing and optimization
parallel query processing
0.212016
Explaining Outputs in Modern Data Analytics · Proc. VLDB Endow. 2016
Spatial and temporal data management
spatial indexing
0.212016
graphVizdb: A scalable platform for interactive large graph visualization · ICDE 2016
Privacy and data protection
privacy-preserving data analysis
0.212023
TVA: A multi-party computation system for secure and expressive time series analytics · USENIX Security Symposium 2023
Cloud and datacenter computing › cloud security
untrusted cloud
0.212023
SECRECY: Secure collaborative analytics in untrusted clouds · NSDI 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontology reasoning
0.212014
Efficient Identification of Implicit Facts in Incomplete OWL2-EL Knowledge Bases · Proc. VLDB Endow. 2014
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology › ontology reasoning
OWL reasoning
0.212014
Efficient Identification of Implicit Facts in Incomplete OWL2-EL Knowledge Bases · Proc. VLDB Endow. 2014
Data integration and cleaning › semantic integration
ontology-based data integration
0.212014
Efficient Identification of Implicit Facts in Incomplete OWL2-EL Knowledge Bases · Proc. VLDB Endow. 2014
Query processing and optimization
range query
0.212014
An Effective Encoding Scheme for Spatial RDF Data · Proc. VLDB Endow. 2014
Spatial and temporal data management › spatial query processing
spatial join
0.212014
An Effective Encoding Scheme for Spatial RDF Data · Proc. VLDB Endow. 2014

Methods — techniques the papers use, named apart from their topics

vectorized query engine · 1.7oblivious operators · 1.7multiparty computation · 0.9multi-party computation · 0.9secure computation · 0.7secret sharing · 0.7structural information recovery · 0.6log reconstruction · 0.6query optimization · 0.4migration granularity · 0.4lineage · 0.4indexing · 0.4checkpointing · 0.4ahead-of-time preparation · 0.4r-tree · 0.2differential dataflow · 0.2backward tracing · 0.2spatial encoding scheme · 0.2
YearPublicationVenuePosition
2026 Slowpoke: End-to-end Throughput Optimization Modeling for Microservice Applications
Yizheng Xie, Di Jin 0004, Oguzhan Çölkesen, Vasiliki Kalavri, John Liagouris, Nikos Vasilakis
NSDI5
2025 ORQ: Complex Analytics on Private Data with Strong Security Guarantees
abstract
We present Orq, a system that enables collaborative analysis of large private datasets using cryptographically secure multi-party computation (MPC). Orq protects data against semi-honest or malicious parties and can efficiently evaluate relational queries with multi-way joins and aggregations that have been considered notoriously expensive under MPC. To do so, Orq eliminates the quadratic cost of secure joins by leveraging the fact that, in practice, the structure of many real queries allows us to join records and apply the aggregations "on the fly" while keeping the result size bounded. On the system side, Orq contributes generic oblivious operators, a data-parallel vectorized query engine, a communication layer that amortizes MPC network costs, and a dataflow API for expressing relational analytics—all built from the ground up.
Eli Baum, Sam Buxbaum, Nitin Mathai, Muhammad Faisal 0001, Vasiliki Kalavri, Mayank Varia, John Liagouris
SOSP7
2023 SECRECY: Secure collaborative analytics in untrusted clouds
John Liagouris, Vasiliki Kalavri, Muhammad Faisal 0001, Mayank Varia
NSDI1
2023 TVA: A multi-party computation system for secure and expressive time series analytics
Muhammad Faisal 0001, Jerry Zhang, John Liagouris, Vasiliki Kalavri, Mayank Varia
USENIX Security Symposium3
2022 A new benchmark harness for systematic and robust evaluation of streaming state stores
abstract
Modern stream processing systems often rely on embedded key-value stores, like RocksDB, to manage the state of long-running computations. Evaluating the performance of these stores when used for streaming workloads is cumbersome as it requires the configuration and deployment of a stream processing system that integrates the respective store, and the execution of representative queries to collect measurements.
Esmail Asyabi, Yuanli Wang, John Liagouris, Vasiliki Kalavri, Azer Bestavros
EuroSys3
2020 In support of workload-aware streaming state management
Vasiliki Kalavri, John Liagouris
HotStorage2
2019 Lineage stash: fault tolerance off the critical path
abstract
As cluster computing frameworks such as Spark, Dryad, Flink, and Ray are being deployed in mission critical applications and on larger and larger clusters, their ability to tolerate failures is growing in importance. These frameworks employ two broad approaches for fault tolerance: checkpointing and lineage. Checkpointing exhibits low overhead during normal operation but high overhead during recovery, while lineage-based solutions make the opposite tradeoff.
Stephanie Wang, John Liagouris, Robert Nishihara, Philipp Moritz, Ujval Misra, Alexey Tumanov, Ion Stoica
SOSP2
2019 Megaphone: Latency-conscious state migration for distributed streaming dataflows
abstract
We design and implement Megaphone, a data migration mechanism for stateful distributed dataflow engines with latency objectives. When compared to existing migration mechanisms, Megaphone has the following differentiating characteristics: (i) migrations can be subdivided to a configurable granularity to avoid latency spikes, and (ii) migrations can be prepared ahead of time to avoid runtime coordination. Megaphone is implemented as a library on an unmodified timely dataflow implementation, and provides an operator interface compatible with its existing APIs. We evaluate Megaphone on established benchmarks with varying amounts of state and observe that compared to naïve approaches Megaphone reduces service latencies during reconfiguration by orders of magnitude without significantly increasing steady-state overhead.
Moritz Hoffmann 0001, Andrea Lattuada 0001, Frank McSherry, Vasiliki Kalavri, John Liagouris, Timothy Roscoe
Proc. VLDB Endow.5
2019 SRX: efficient management of spatial RDF data
Konstantinos Theocharidis, John Liagouris, Nikos Mamoulis, Panagiotis Bouros, Manolis Terrovitis
VLDB J.2
2018 SnailTrail: Generalizing Critical Paths for Online Analysis of Distributed Dataflows
Moritz Hoffmann 0001, Andrea Lattuada 0001, John Liagouris, Vasiliki Kalavri, Desislava C. Dimitrova, Sebastian Wicki, Zaheer Chothia, Timothy Roscoe
NSDI3
2018 Three steps is all you need: fast, accurate, automatic scaling decisions for distributed streaming dataflows
Vasiliki Kalavri, John Liagouris, Moritz Hoffmann 0001, Desislava C. Dimitrova, Matthew Forshaw, Timothy Roscoe
OSDI2
2017 Online Reconstruction of Structural Information from Datacenter Logs
abstract
Well-run datacenter application architectures are heavily instrumented to provide detailed traces of messages and remote invocations. Reconstructing user sessions, call graphs, transaction trees, and other structural information from these messages, a process known as sessionization, is the foundation for a variety of diagnostic, profiling, and monitoring tasks essential to the operation of the datacenter.
Zaheer Chothia, John Liagouris, Desislava C. Dimitrova, Timothy Roscoe
EuroSys2
2016 graphVizdb: A scalable platform for interactive large graph visualization
abstract
We present a novel platform for the interactive visualization of very large graphs. The platform enables the user to interact with the visualized graph in a way that is very similar to the exploration of maps at multiple levels. Our approach involves an offline preprocessing phase that builds the layout of the graph by assigning coordinates to its nodes with respect to a Euclidean plane. The respective points are indexed with a spatial data structure, i.e., an R-tree, and stored in a database. Multiple abstraction layers of the graph based on various criteria are also created offline, and they are indexed similarly so that the user can explore the dataset at different levels of granularity, depending on her particular needs. Then, our system translates user operations into simple and very efficient spatial operations (i.e., window queries) in the backend. This technique allows for a fine-grained access to very large graphs with extremely low latency and memory requirements and without compromising the functionality of the tool. Our web-based prototype supports three main operations: (1) interactive navigation, (2) multi-level exploration, and (3) keyword search on the graph metadata.
Nikos Bikakis, John Liagouris, Maria Krommyda, George Papastefanatos, Timos K. Sellis
ICDE2
2016 Explaining Outputs in Modern Data Analytics
abstract
We report on the design and implementation of a general framework for interactively explaining the outputs of modern data-parallel computations, including iterative data analytics. To produce explanations, existing works adopt a naive backward tracing approach which runs into known issues; naive backward tracing may identify: (i) too much information that is difficult to process, and (ii) not enough information to reproduce the output, which hinders the logical debugging of the program. The contribution of this work is twofold. First, we provide methods to effectively reduce the size of explanations based on the first occurrence of a record in an iterative computation. Second, we provide a general method for identifying explanations that are sufficient to reproduce the target output in arbitrary computations -- a problem for which no viable solution existed until now. We implement our approach on differential dataflow , a modern high-throughput, low-latency dataflow platform. We add a small (but extensible) set of rules to explain each of its data-parallel operators, and we implement these rules as differential dataflow operators themselves. This choice allows our implementation to inherit the performance characteristics of differential dataflow, and results in a system that efficiently computes and updates explanatory inputs even as the inputs of the reference computation change. We evaluate our system with various analytic tasks on real datasets, and we show that it produces concise explanations in tens of milliseconds, while remaining faster -- up to two orders of magnitude -- than even the best implementations that do not support explanations.
Zaheer Chothia, John Liagouris, Frank McSherry, Timothy Roscoe
Proc. VLDB Endow.2
2014 Disassociation for electronic health record privacy
Grigorios Loukides, John Liagouris, Aris Gkoulalas-Divanis, Manolis Terrovitis
J. Biomed. Informatics2
2014 An Effective Encoding Scheme for Spatial RDF Data
abstract
The RDF data model has recently been extended to support representation and querying of spatial information (i.e., locations and geometries), which is associated with RDF entities. Still, there are limited efforts towards extending RDF stores to efficiently support spatial queries, such as range selections (e.g., find entities within a given range) and spatial joins (e.g., find pairs of entities whose locations are close to each other). In this paper, we propose an extension for RDF stores that supports efficient spatial data management. Our contributions include an effective encoding scheme for entities having spatial locations, the introduction of on-the-fly spatial filters and spatial join algorithms, and several optimizations that minimize the overhead of geometry and dictionary accesses. We implemented the proposed techniques as an extension to the opensource RDF-3X engine and we experimentally evaluated them using real RDF knowledge bases. The results show that our system offers robust performance for spatial queries, while introducing little overhead to the original query engine.
John Liagouris, Nikos Mamoulis, Panagiotis Bouros, Manolis Terrovitis
Proc. VLDB Endow.1
2014 Efficient Identification of Implicit Facts in Incomplete OWL2-EL Knowledge Bases
abstract
Integrating incomplete and possibly inconsistent data from various sources is a challenge that arises in several application areas, especially in the management of scientific data. A rising trend for data integration is to model the data as axioms in the Web Ontology Language (OWL) and use inference rules to identify new facts. Although there are several approaches that employ OWL for data integration, there is little work on scalable algorithms able to handle large datasets that do not fit in main memory. The main contribution of this paper is an algorithm that allows the effective use of OWL for integrating data in an environment with limited memory. The core idea is to exhaustively apply a set of complex inference rules on large disk-resident datasets. To the best of our knowledge, this is the first work that proposes an I/O-aware algorithm for tackling with such an expressive subset of OWL like the one we address here. Previous approaches considered either simpler models (e.g. RDFS) or main-memory algorithms. In the paper we detail the proposed algorithm, prove its correctness, and experimentally evaluate it on real and synthetic data.
John Liagouris, Manolis Terrovitis
Proc. VLDB Endow.1
2013 RDivF: Diversifying Keyword Search on RDF Graphs
Nikos Bikakis, Giorgos Giannopoulos, John Liagouris, Dimitrios Skoutas 0001, Theodore Dalamagas 0001, Timos K. Sellis
TPDL3
2012 Privacy Preservation by Disassociation
abstract
In this work, we focus on protection against identity disclosure in the publication of sparse multidimensional data. Existing multidimensional anonymization techniques (a) protect the privacy of users either by altering the set of quasi-identifiers of the original data (e.g., by generalization or suppression) or by adding noise (e.g., using differential privacy) and/or (b) assume a clear distinction between sensitive and non-sensitive information and sever the possible linkage. In many real world applications the above techniques are not applicable. For instance, consider web search query logs. Suppressing or generalizing anonymization methods would remove the most valuable information in the dataset: the original query terms. Additionally, web search query logs contain millions of query terms which cannot be categorized as sensitive or non-sensitive since a term may be sensitive for a user and non-sensitive for another. Motivated by this observation, we propose an anonymization technique termed disassociation that preserves the original terms but hides the fact that two or more different terms appear in the same record. We protect the users' privacy by disassociating record terms that participate in identifying combinations. This way the adversary cannot associate with high probability a record with a rare combination of terms. To the best of our knowledge, our proposal is the first to employ such a technique to provide protection against identity disclosure . We propose an anonymization algorithm based on our approach and evaluate its performance on real and synthetic datasets, comparing it against other state-of-the-art methods based on generalization and differential privacy.
Manolis Terrovitis, John Liagouris, Nikos Mamoulis, Spiros Skiadopoulos
Proc. VLDB Endow.2
2011 Mobile Task Computing: Beyond Location-Based Services and EBooks
John Liagouris, Spiros Athanasiou, Alexandros Efentakis, Stefan Pfennigschmidt, Dieter Pfoser, Eleni Tsigka, Agnès Voisard
W2GIS1