EDBT 2026 Demo / reviewers in the wild / expert
Steffen Maass
dblp:182/6239
· DBLP profile ↗
5ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Memory systems · 52% Storage systems · 20% Processor architecture and microarchitecture · 8% | |
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory management
virtual memory |
0.4 | 1 | 2020 | ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020 |
Operating systems › i/o › i/o subsystem
i/o stack |
0.3 | 1 | 2018 | Solros: a data-centric operating system architecture for heterogeneous computing · EuroSys 2018 |
Operating systems › resource management › memory management
virtual memory |
0.3 | 1 | 2018 | LATR: Lazy Translation Coherence · ASPLOS 2018 |
Graph data management
graph processing |
0.3 | 1 | 2017 | Mosaic: Processing a Trillion-Edge Graph on a Single Machine · EuroSys 2017 |
Graph data management › graph processing
single-machine graph processing |
0.3 | 1 | 2017 | Mosaic: Processing a Trillion-Edge Graph on a Single Machine · EuroSys 2017 |
Storage systems
file systems |
0.2 | 1 | 2016 | Understanding Manycore Scalability of File Systems · USENIX ATC 2016 |
Operating systems › resource management
memory management |
0.1 | 1 | 2020 | ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020 |
Operating systems › resource management › memory management
page swapping |
0.1 | 1 | 2020 | ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020 |
Memory systems
memory disaggregation |
0.1 | 1 | 2020 | ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020 |
Memory systems › memory management › virtual memory
address translation |
0.1 | 1 | 2018 | LATR: Lazy Translation Coherence · ASPLOS 2018 |
Processor architecture and microarchitecture › special-purpose processor
coprocessor |
0.1 | 1 | 2018 | Solros: a data-centric operating system architecture for heterogeneous computing · EuroSys 2018 |
Parallel and multicore computing › graph processing
heterogeneous graph processing |
0.1 | 1 | 2017 | Mosaic: Processing a Trillion-Edge Graph on a Single Machine · EuroSys 2017 |
Performance modeling and evaluation › parallel performance evaluation
multicore scalability |
0.1 | 1 | 2016 | Understanding Manycore Scalability of File Systems · USENIX ATC 2016 |
Methods — techniques the papers use, named apart from their topics
vertex-centric and edge-centric processing · 0.6hybrid execution model · 0.6hilbert-ordered tiles · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | ECOTLB: Eventually Consistent TLBsabstractWe propose ecoTLB —software-based eventual translation lookaside buffer (TLB) coherence—which eliminates the overhead of the synchronous TLB shootdown mechanism in operating systems that use address space identifiers (ASIDs). With an eventual TLB coherence, ecoTLB improves the performance of free and page swap operations by removing the inter-processor interrupt (IPI) overheads incurred to invalidate TLB entries. We show that the TLB shootdown has implications for page swapping in particular in emerging, disaggregated data centers and demonstrate that ecoTLB can improve both the performance and the specific swapping policy decisions using ecoTLB ’s asynchronous mechanism. We demonstrate that ecoTLB improves the performance of real-world applications, such as Memcached and Make, that perform page swapping using Infiniswap , a solution for next generation data centers that use disaggregated memory, by up to 17.2%. Moreover, ecoTLB improves the 99th percentile tail latency of Memcached by up to 70.8% due to its asynchronous scheme and improved policy decisions. Furthermore, we show that recent features to improve security in the Linux kernel, like kernel page table isolation (KPTI), can result in significant performance overheads on architectures without support for specific instructions to clear single entries in tagged TLBs, falling back to full TLB flushes. In this scenario, ecoTLB is able to recover the performance lost for supporting KPTI due to its asynchronous shootdown scheme and its support for tagged TLBs. Finally, we demonstrate that ecoTLB improves the performance of free operations by up to 59.1% on a 120-core machine and improves the performance of Apache on a 16-core machine by up to 13.7% compared to baseline Linux, and by up to 48.2% compared to ABIS, a recent state-of-the-art research prototype that reduces the number of IPIs. Steffen Maass, Mohan Kumar Kumar, Taesoo Kim, Tushar Krishna, Abhishek Bhattacharjee |
ACM Trans. Archit. Code Optim. | 1 |
| 2018 | LATR: Lazy Translation CoherenceabstractWe propose LATR-lazy TLB coherence-a software-based TLB shootdown mechanism that can alleviate the overhead of the synchronous TLB shootdown mechanism in existing operating systems. By handling the TLB coherence in a lazy fashion, LATR can avoid expensive IPIs which are required for delivering a shootdown signal to remote cores, and the performance overhead of associated interrupt handlers. Therefore, virtual memory operations, such as free and page migration operations, can benefit significantly from LATR's mechanism. For example, LATR improves the latency of munmap() by 70.8% on a 2-socket machine, a widely used configuration in modern data centers. Real-world, performance-critical applications such as web servers can also benefit from LATR: without any application-level changes, LATR improves Apache by 59.9% compared to Linux, and by 37.9% compared to ABIS, a highly optimized, state-of-the-art TLB coherence technique. Mohan Kumar, Steffen Maass, Sanidhya Kashyap, Ján Veselý, Zi Yan, Taesoo Kim, Abhishek Bhattacharjee, Tushar Krishna |
ASPLOS | 2 |
| 2018 | Solros: a data-centric operating system architecture for heterogeneous computingabstractWe propose Solros---a new operating system architecture for heterogeneous systems that comprises fast host processors, slow but massively parallel co-processors, and fast I/O devices. A general consensus to fully drive such a hardware system is to have a tight integration among processors and I/O devices. Thus, in the Solros architecture, a co-processor OS (data-plane OS) delegates its services, specifically I/O stacks, to the host OS (control-plane OS). Our observation for such a design is that global coordination with system-wide knowledge (e.g., PCIe topology, a load of each co-processor) and the best use of heterogeneous processors is critical to achieving high performance. Hence, we fully harness these specialized processors by delegating complex I/O stacks on fast host processors, which leads to an efficient global coordination at the level of the control-plane OS. Changwoo Min, Woon-Hak Kang, Mohan Kumar, Sanidhya Kashyap, Steffen Maass, Heeseung Jo, Taesoo Kim |
EuroSys | 5 |
| 2017 | Mosaic: Processing a Trillion-Edge Graph on a Single MachineabstractProcessing a one trillion-edge graph has recently been demonstrated by distributed graph engines running on clusters of tens to hundreds of nodes. In this paper, we employ a single heterogeneous machine with fast storage media (e.g., NVMe SSD) and massively parallel coprocessors (e.g., Xeon Phi) to reach similar dimensions. By fully exploiting the heterogeneous devices, we design a new graph processing engine, named Mosaic, for a single machine. We propose a new locality-optimizing, space-efficient graph representation---Hilbert-ordered tiles, and a hybrid execution model that enables vertex-centric operations in fast host processors and edge-centric operations in massively parallel coprocessors. Steffen Maass, Changwoo Min, Sanidhya Kashyap, Woon-Hak Kang, Mohan Kumar, Taesoo Kim |
EuroSys | 1 |
| 2016 | Understanding Manycore Scalability of File Systems
Changwoo Min, Sanidhya Kashyap, Steffen Maass, Taesoo Kim |
USENIX ATC | 3 |