Georgios Tziantzioulis

dblp:33/10064 · also Georgios Tziantzoulis · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-7614-3815ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Evaluation of MindPalace for Chip Design Tradeoffs on Function-as-a-Service
abstract
Function-as-a-Service (FaaS) continues to be a popular and growing way to access abundant cloud resources. Both the research community and industry have begun investigating designing new or modifying existing processors for FaaS. While there exists many traditional computer architecture simulators, the complex, full-system and full-stack nature of FaaS workloads has opened the door to needing more purpose-built FaaS-specific simulators along with easy to use, simulation-friendly, characteristic workloads. In order to fully study FaaS workloads, the simulator needs to run a full OS running many concurrent containers with a full-stack serverless platform like OpenWhisk. This paper characterizes and describes MindPalace which is a simulator and framework designed to investigate microarchitecture innovations for FaaS workloads. MindPalace utilizes both QEMU and ChampSim. With QEMU, MindPalace is able to run complex, full-system and full-stack workloads including running the Apache OpenWhisk platform. By leveraging ChampSim, MindPalace can conduct detailed computer architecture and micro-architecture research and has access to advanced micro-architecture models for branch predictors and prefetchers. In addition, MindPalace uses a VM with Apache OpenWhisk running. In this work, we compare MindPalace to other simulators and use MindPalace to investigate FaaS-specific architecture ideas including a design which reloads branch predictor state between function invocations. We also characterize MindPalace and study the accuracy of simulation versus real hardware in terms of branch prediction rates, cache miss rates, and IPC.
Kaifeng Xu, Georgios Tziantzioulis, David Wentzlaff
ISPASS2
2024 MindPalace: A Framework for Studying Microarchitecture Design of Function-as-a-Service
abstract
The increasing popularity of Function-as-a-Service (FaaS) applications has created new challenges for current microarchitecture designs due to their complex middleware and significant Operating System (OS) overhead, prohibiting the detailed study of new hardware designs. This work introduces MindPalace, an open-source full-system framework that developed to study the architectural behavior of large middleware serverless systems. MindPalace contains a simulation infrastructure and a pre-configured virtual machine (VM) to assist FaaS architecture research. The simulation tool leverages QEMU and ChampSim, combining a featureful full-system emulator with a rich microarchitecture simulator. The VM contains a full Ubuntu image with preinstalled Open Whisk, together with checkpoints for ready-to-run serverless applications. Combining them enables full-system FaaS architectural study of the memory hierarchy, branch predictors, prefetchers, and other microarchitectural components.
Kaifeng Xu, Georgios Tziantzioulis, David Wentzlaff
ISPASS2
2023 Supply Chain Aware Computer Architecture
abstract
Progressively and increasingly, our society has become more and more dependent on semiconductors and semiconductor-enabled products and services. The importance of chips and their supply chains has been highlighted during the 2020-present chip shortage caused by manufacturing disruptions and increased demand due to the COVID-19 pandemic. However, semiconductor supply chains are inherently vulnerable to disruptions and chip crises can easily recur in the future.
August Ning, Georgios Tziantzioulis, David Wentzlaff
ISCA2
2022 FracDRAM: Fractional Values in Off-the-Shelf DRAM
abstract
As one of the cornerstones of computing, dynamic random-access memory (DRAM) is prevalent across digital systems. Over the years, researchers have proposed modifications to DRAM macros or explored alternative uses of existing DRAM chips to extend the functionality of this ubiquitous media. This work expands on the latter, providing new insights and demonstrating new functionalities in unmodified, commodity DRAM. FracDRAM is the first work to show how fractional values can be stored in off-the-shelf DRAM. We propose two primitive operations built with specially timed DRAM command sequences, to either store fractional values to the entire DRAM row or to masked bits in a row. Utilizing fractional values, this work enables more modules to perform the in-memory majority operation, increases the stability of the existing in-memory majority operation, and builds a state-of-the-art DRAM-based PUF with unmodified DRAM. In total, 582 DDR3 chips from seven major vendors are evaluated and characterized under different environments in this work. FracDRAM breaks through the conventional binary abstraction of DRAM logic, and brings new functions to the existing DRAM macro.
Fei Gao 0016, Georgios Tziantzioulis, David Wentzlaff
MICRO2
2022 OPDB: A Scalable and Modular Design Benchmark
abstract
Progress in electronic design automation (EDA) has enabled us to manage the exponential growth of complexity in integrated circuits (ICs) allowed by Moore’s Law, and translate it into performance improvements. Furthermore, as Moore’s law slows down and with the collapse of Dennard scaling over the past decade, EDA tools become increasingly important in addressing inefficiencies in ICs. As new proposals and techniques are put forward to address the current and future issues of IC design, a concrete set of contemporary benchmarks needs to be used for their evaluation. Traditionally, EDA researchers have mainly relied on industry provided design benchmarks in evaluating the performance of their tools. However, due to their effort to maintain their competitive advantage, industry releases include older designs with limited information about the details of each design. This means that in multiple cases, new proposals are evaluated with significantly dated designs, often more than a decade old, especially in terms of scale. In this work, we describe OPDB: a scalable, modular, heterogeneous, and extensible design benchmark for the EDA community. OPDB leverages and extends the OpenPiton open-source, tile-based research infrastructure to create a surplus of design benchmarks that target different components. Due to the tiled nature of this architecture, OPDB benchmarks can be made arbitrarily large in order to evaluate the efficiency of EDA tools across different design scales and configurations. OPDB contains several accelerators enabling full SoC designs to be used as benchmarks for EDA tools.
Georgios Tziantzioulis, Ting-Jung Chang, Jonathan Balkind, Jinzheng Tu 0001, Fei Gao 0016, David Wentzlaff
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 ST2 GPU: An Energy-Efficient GPU Design with Spatio-Temporal Shared-Thread Speculative Adders
abstract
Modern GPUs employ thousands of cores, yielding higher performance but also higher power consumption. To meet performance targets while staying within a reasonable power budget, designers have to make these execution cores increasingly more power efficient. One way to increase their power efficiency is to employ power-efficient adders. In this paper, we observe that consecutive arithmetic computations from the same code location are highly correlated and propose ST2GPU, a GPU architecture that uses history-based speculative adders that produce guaranteed correct results while saving 70% of the nominal adder power. We estimate that ST2GPU saves 21% of the GPU chip energy with practically no performance and area overheads.
Vijay Kandiah, Ali Murat Gok, Georgios Tziantzioulis, Nikos Hardavellas
DAC3
2019 Paths to Fast Barrier Synchronization on the Node
abstract
Synchronization primitives like barriers heavily impact the performance of parallel programs. As core counts increase and granularity decreases, the value of enabling fast barriers increases. Through the evaluation of the performance of a variety of software implementations of barriers, we found the cost of software barriers to be on the order of tens of thousands of cycles on various incarnations of x64 hardware. We argue that reducing the latency of a barrier via hardware support will dramatically improve the performance of existing applications and runtimes, and would enable new execution models, including those which currently do not perform well on multicore machines. To support our argument, we first present the design, implementation, and evaluation of a barrier on the Intel HARP, a prototype that integrates an x64 processor and FPGA in the same package. This effort gives insight into the potential speed and compactness of hardware barriers, and suggests useful improvements to the HARP platform. Next, we turn to the processor itself and describe an x64 ISA extension for barriers, and how it could be implemented in the microarchitecture with minimal collateral changes. This design allows for barriers to be securely managed jointly between the OS and the application. Finally, we speculate on how barrier synchronization might be implemented on future photonics-based hardware.
Conor Hetland, Georgios Tziantzioulis, Brian Suchy, Michael Leonard, John Albers, Nikos Hardavellas, Peter A. Dinda
HPDC2
2019 Prospects for Functional Address Translation
abstract
Address translation fundamentally embodies a translation function that maps from virtual to physical addresses. In current systems, the translation function is encoded by the kernel in an in-memory radix tree structure (the page table hierarchy) which is then interpreted by the hardware (the pagewalker, pagewalk-caches, and TLBs). We consider implementing the translation function itself as reconfigurable hardware-does this make any sense? To study this question, we collected numerous in-situ Linux page tables for a wide range of workloads, including those from HPC, to serve as example translation functions. We then prototyped several potential mechanisms to implement the translation function, including inverted page tables with function-specific perfect hashing, translation functions directly implemented using Espresso-minimized PLAs, translation functions genetically-evolved in a language suitable for FPGA-like synthesis, and translation functions based on recovered/manufactured region (segment/mmap) lookup using multiplexor trees. Each mechanism was then evaluated using the Linux page tables, primarily for space and lookup speed. We report our findings and try to address the question.
Conor Hetland, Georgios Tziantzioulis, Brian Suchy, Kyle C. Hale, Nikos Hardavellas, Peter A. Dinda
MASCOTS2
2019 ComputeDRAM: In-Memory Compute Using Off-the-Shelf DRAMs
abstract
In-memory computing has long been promised as a solution to the "Memory Wall" problem. Recent work has proposed using chargesharing on the bit-lines of a memory in order to compute in-place and with massive parallelism, all without having to move data across the memory bus. Unfortunately, prior work has required modification to RAM designs (e.g. adding multiple row decoders) in order to open multiple rows simultaneously. So far, the competitive and low-margin nature of the DRAM industry has made commercial DRAM manufacturers resist adding any additional logic into DRAM. This paper addresses the need for in-memory computation with little to no change to DRAM designs. It is the first work to demonstrate in-memory computation with off-the-shelf, unmodified, commercial, DRAM. This is accomplished by violating the nominal timing specification and activating multiple rows in rapid succession, which happens to leave multiple rows open simultaneously, thereby enabling bit-line charge sharing. We use a constraint-violating command sequence to implement and demonstrate row copy, logical OR, and logical AND in unmodified, commodity, DRAM. Subsequently, we employ these primitives to develop an architecture for arbitrary, massively-parallel, computation. Utilizing a customized DRAM controller in an FPGA and commodity DRAM modules, we characterize this opportunity in hardware for all major DRAM vendors. This work stands as a proof of concept that in-memory computation is possible with unmodified DRAM modules and that there exists a financially feasible way for DRAM manufacturers to support in-memory compute.
Fei Gao 0016, Georgios Tziantzioulis, David Wentzlaff
MICRO2
2016 Lazy Pipelines: Enhancing quality in approximate computing
Georgios Tziantzioulis, Ali Murat Gok, S. M. Faisal, Nikos Hardavellas, Seda Ogrenci Memik, Srinivasan Parthasarathy 0001
DATE1
2015 Edge importance identification for energy efficient graph processing
abstract
Modern graphs are large, often containing billions of nodes and edges that demand huge amount of processing for analysis purposes. The algorithms processing these graphs often run for long time and consume substantial amount of energy. However, not all edges in the graphs are equally important. Some edges play critical role in maintaining the community and other interesting structures in the graph, while the rest are less important for analysis. Identifying edges as important and unimportant allows one to apply elastic fidelity computing when processing edges of low importance, hence saving significant amount of energy while processing large graphs. In this paper we propose a novel technique for identifying important edges in a graph using a fast method that exploits locality sensitive hashing. We then propose a framework for energy-efficient computing that applies elastic fidelity computing when processing edges of low importance and applies full fidelity computing when processing important edges. This allows the framework to deliver good results while saving energy when processing a large number of low-importance edges. Our proposed technique reduces the power consumption by 3-30% while still producing results that are within acceptable range of the full-accuracy results.
S. M. Faisal, Georgios Tziantzioulis, Ali Murat Gok, Nikos Hardavellas, Seda Ogrenci Memik, Srinivasan Parthasarathy 0001
IEEE BigData2
2015 b-HiVE: a bit-level history-based error model with value correlation for voltage-scaled integer and floating point units
abstract
Existing timing error models for voltage-scaled functional units ignore the effect of history and correlation among outputs, and the variation in the error behavior at different bit locations. We propose b-HiVE, a model for voltage-scaling-induced timing errors that incorporates these attributes and demonstrates their impact on the overall model accuracy. On average across several operations, b-HiVE's estimation is within 1--3% of comprehensive analog simulations, which corresponds to 5--17x higher accuracy (6--10x on average) than error models currently used in approximate computing research. To the best of our knowledge, we present the first bit-level error models of arithmetic units, and the first error models for voltage scaling of bitwise logic operations and floating-point units.
Georgios Tziantzioulis, Ali Murat Gok, S. M. Faisal, Nikos Hardavellas, Seda Ogrenci Memik, Srinivasan Parthasarathy 0001
DAC1
2014 GemFI: A Fault Injection Tool for Studying the Behavior of Applications on Unreliable Substrates
abstract
Dependable computing on unreliable substrates is the next challenge the computing community needs to overcome due to both manufacturing limitations in low geometries and the necessity to aggressively minimize power consumption. System designers often need to analyze the way hardware faults manifest as errors at the architectural level and how these errors affect application correctness. This paper introduces GemFI, a fault injection tool based on the cycle accurate full system simulator Gem5. GemFI provides fault injection methods and is easily extensible to support future fault models. It also supports multiple processor models and ISAs and allows fault injection in both functional and cycle-accurate simulations. GemFI offers fast-forwarding of simulation campaigns via check pointing. Moreover, it facilitates the parallel execution of campaign experiments on a network of workstations. In order to validate and evaluate GemFI, we used it to apply fault injection on a series of real-world kernels and applications. The evaluation indicates that its overhead compared with Gem5 is minimal (up to 3.3%), whereas optimizations such as fast-forwarding via check pointing and execution on NoWs can significantly reduce simulation time of a fault injection campaign.
Konstantinos Parasyris, Georgios Tziantzioulis, Christos D. Antonopoulos, Nikolaos Bellas
DSN2
2011 Significance driven computation on next-generation unreliable platforms
abstract
In this paper, we propose a design paradigm for energy efficient and variation-aware operation of next-generation multicore heterogeneous platforms. The main idea behind the proposed approach lies on the observation that not all operations are equally important in shaping the output quality of various applications and of the overall system. Based on such an observation, we suggest that all levels of the software design stack, including the programming model, compiler, operating system (OS) and runtime system should identify the critical tasks and ensure correct operation of such tasks by assigning them to dynamically adjusted reliable cores/units. Specifically, based on error rates and operating conditions identified by a sense-and-adapt (SeA) unit, the OS selects and sets the right mode of operation of the overall system. The run-time system identifies the critical/less-critical tasks based on special directives and schedules them to the appropriate units that are dynamically adjusted for highly-accurate/approximate operation by tuning their voltage/frequency. Units that execute less significant operations can operate at voltages less than what is required for correct operation and consume less power, if required, since such tasks do not need to be always exact as opposed to the critical ones. Such scheme can lead to energy efficient and reliable operation, while reducing the design cost and overheads of conventional circuit/micro-architecture level techniques.
Georgios Karakonstantis, Nikolaos Bellas, Christos D. Antonopoulos, Georgios Tziantzioulis, Vaibhav Gupta, Kaushik Roy 0001
DAC4