Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Steffen Maass

dblp:182/6239 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 52% Storage systems · 20% Processor architecture and microarchitecture · 8%
Software engineering, system software, and programming languages
3 papers
Operating systems · 100%
Databases, data mining, and information retrieval
1 paper
Graph data management · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management
virtual memory
0.412020
ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020
Operating systems › i/o › i/o subsystem
i/o stack
0.312018
Solros: a data-centric operating system architecture for heterogeneous computing · EuroSys 2018
Operating systems › resource management › memory management
virtual memory
0.312018
LATR: Lazy Translation Coherence · ASPLOS 2018
Graph data management
graph processing
0.312017
Mosaic: Processing a Trillion-Edge Graph on a Single Machine · EuroSys 2017
Graph data management › graph processing
single-machine graph processing
0.312017
Mosaic: Processing a Trillion-Edge Graph on a Single Machine · EuroSys 2017
Storage systems
file systems
0.212016
Understanding Manycore Scalability of File Systems · USENIX ATC 2016
Operating systems › resource management
memory management
0.112020
ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020
Operating systems › resource management › memory management
page swapping
0.112020
ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020
Memory systems
memory disaggregation
0.112020
ECOTLB: Eventually Consistent TLBs · ACM Trans. Archit. Code Optim. 2020
Memory systems › memory management › virtual memory
address translation
0.112018
LATR: Lazy Translation Coherence · ASPLOS 2018
Processor architecture and microarchitecture › special-purpose processor
coprocessor
0.112018
Solros: a data-centric operating system architecture for heterogeneous computing · EuroSys 2018
Parallel and multicore computing › graph processing
heterogeneous graph processing
0.112017
Mosaic: Processing a Trillion-Edge Graph on a Single Machine · EuroSys 2017
Performance modeling and evaluation › parallel performance evaluation
multicore scalability
0.112016
Understanding Manycore Scalability of File Systems · USENIX ATC 2016

Methods — techniques the papers use, named apart from their topics

vertex-centric and edge-centric processing · 0.6hybrid execution model · 0.6hilbert-ordered tiles · 0.6
YearPublicationVenuePosition
2020 ECOTLB: Eventually Consistent TLBs
abstract
We propose ecoTLB —software-based eventual translation lookaside buffer (TLB) coherence—which eliminates the overhead of the synchronous TLB shootdown mechanism in operating systems that use address space identifiers (ASIDs). With an eventual TLB coherence, ecoTLB improves the performance of free and page swap operations by removing the inter-processor interrupt (IPI) overheads incurred to invalidate TLB entries. We show that the TLB shootdown has implications for page swapping in particular in emerging, disaggregated data centers and demonstrate that ecoTLB can improve both the performance and the specific swapping policy decisions using ecoTLB ’s asynchronous mechanism. We demonstrate that ecoTLB improves the performance of real-world applications, such as Memcached and Make, that perform page swapping using Infiniswap , a solution for next generation data centers that use disaggregated memory, by up to 17.2%. Moreover, ecoTLB improves the 99th percentile tail latency of Memcached by up to 70.8% due to its asynchronous scheme and improved policy decisions. Furthermore, we show that recent features to improve security in the Linux kernel, like kernel page table isolation (KPTI), can result in significant performance overheads on architectures without support for specific instructions to clear single entries in tagged TLBs, falling back to full TLB flushes. In this scenario, ecoTLB is able to recover the performance lost for supporting KPTI due to its asynchronous shootdown scheme and its support for tagged TLBs. Finally, we demonstrate that ecoTLB improves the performance of free operations by up to 59.1% on a 120-core machine and improves the performance of Apache on a 16-core machine by up to 13.7% compared to baseline Linux, and by up to 48.2% compared to ABIS, a recent state-of-the-art research prototype that reduces the number of IPIs.
Steffen Maass, Mohan Kumar Kumar, Taesoo Kim, Tushar Krishna, Abhishek Bhattacharjee
ACM Trans. Archit. Code Optim.1
2018 LATR: Lazy Translation Coherence
abstract
We propose LATR-lazy TLB coherence-a software-based TLB shootdown mechanism that can alleviate the overhead of the synchronous TLB shootdown mechanism in existing operating systems. By handling the TLB coherence in a lazy fashion, LATR can avoid expensive IPIs which are required for delivering a shootdown signal to remote cores, and the performance overhead of associated interrupt handlers. Therefore, virtual memory operations, such as free and page migration operations, can benefit significantly from LATR's mechanism. For example, LATR improves the latency of munmap() by 70.8% on a 2-socket machine, a widely used configuration in modern data centers. Real-world, performance-critical applications such as web servers can also benefit from LATR: without any application-level changes, LATR improves Apache by 59.9% compared to Linux, and by 37.9% compared to ABIS, a highly optimized, state-of-the-art TLB coherence technique.
Mohan Kumar, Steffen Maass, Sanidhya Kashyap, Ján Veselý, Zi Yan, Taesoo Kim, Abhishek Bhattacharjee, Tushar Krishna
ASPLOS2
2018 Solros: a data-centric operating system architecture for heterogeneous computing
abstract
We propose Solros---a new operating system architecture for heterogeneous systems that comprises fast host processors, slow but massively parallel co-processors, and fast I/O devices. A general consensus to fully drive such a hardware system is to have a tight integration among processors and I/O devices. Thus, in the Solros architecture, a co-processor OS (data-plane OS) delegates its services, specifically I/O stacks, to the host OS (control-plane OS). Our observation for such a design is that global coordination with system-wide knowledge (e.g., PCIe topology, a load of each co-processor) and the best use of heterogeneous processors is critical to achieving high performance. Hence, we fully harness these specialized processors by delegating complex I/O stacks on fast host processors, which leads to an efficient global coordination at the level of the control-plane OS.
Changwoo Min, Woon-Hak Kang, Mohan Kumar, Sanidhya Kashyap, Steffen Maass, Heeseung Jo, Taesoo Kim
EuroSys5
2017 Mosaic: Processing a Trillion-Edge Graph on a Single Machine
abstract
Processing a one trillion-edge graph has recently been demonstrated by distributed graph engines running on clusters of tens to hundreds of nodes. In this paper, we employ a single heterogeneous machine with fast storage media (e.g., NVMe SSD) and massively parallel coprocessors (e.g., Xeon Phi) to reach similar dimensions. By fully exploiting the heterogeneous devices, we design a new graph processing engine, named Mosaic, for a single machine. We propose a new locality-optimizing, space-efficient graph representation---Hilbert-ordered tiles, and a hybrid execution model that enables vertex-centric operations in fast host processors and edge-centric operations in massively parallel coprocessors.
Steffen Maass, Changwoo Min, Sanidhya Kashyap, Woon-Hak Kang, Mohan Kumar, Taesoo Kim
EuroSys1
2016 Understanding Manycore Scalability of File Systems
Changwoo Min, Sanidhya Kashyap, Steffen Maass, Taesoo Kim
USENIX ATC3