EDBT 2026 Demo / reviewers in the wild / expert
Kyriakos Stavrou
dblp:52/5993
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 67% Electronic design automation · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Runtime systems and virtual machines › binary translation
dynamic binary translation |
0.1 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Methods — techniques the papers use, named apart from their topics
speculative instruction-fusion optimization · 0.4cycle-accurate simulation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | HW/SW co-designed processors: Challenges, design choices and a simulation infrastructure for evaluationabstractImproving single thread performance is a key challenge in modern microprocessors especially because the traditional approach of increasing clock frequency and deep pipelining cannot be pushed further due to power constraints. Therefore, researchers have been looking at unconventional architectures to boost single thread performance without running into the power wall. HW/SW co-designed processors like Nvidia Denver, are emerging as a promising alternative. However, HW/SW co-designed processors need to address some key challenges such as startup delay, providing high performance with simple hardware, translation/optimization overhead, etc. before they can become mainstream. A fundamental requirement for evaluating different design choices and trade-offs to meet these challenges is to have a simulation infrastructure. Unfortunately, there is no such infrastructure available today. Building the aforementioned infrastructure itself poses significant challenges as it encompasses the complexities of not only an architectural framework but also of a compilation one. This paper identifies the key challenges that HW/SW codesigned processors face and the basic requirements for a simulation infrastructure targeting these architectures. Furthermore, the paper presents DARCO, a simulation infrastructure to enable research in this domain. Rakesh Kumar 0003, José Cano 0001, Aleksandar Brankovic, Demos Pavlou, Kyriakos Stavrou, Enric Gibert, Alejandro Martínez, Antonio Gonzalez |
ISPASS | 5 |
| 2014 | Speculative hardware/software co-designed floating-point multiply-add fusionabstractA Fused Multiply-Add (FMA) instruction is currently available in many general-purpose processors. It increases performance by reducing latency of dependent operations and increases precision by computing the result as an indivisible operation with no intermediate rounding. However, since the arithmetic behavior of a single-rounding FMA operation is different than independent FP multiply followed by FP add instructions, some algorithms require significant revalidation and rewriting efforts to work as expected when they are compiled to operate with FMA--a cost that developers may not be willing to pay. Because of that, abundant legacy applications are not able to utilize FMA instructions. In this paper we propose a novel HW/SW collaborative technique that is able to efficiently execute workloads with increased utilization of FMA, by adding the option to get the same numerical result as separate FP multiply and FP add pairs. In particular, we extended the host ISA of a HW/SW co-designed processor with a new Combined Multiply-Add (CMA) instruction that performs an FMA operation with an intermediate rounding. This new instruction is used by a transparent dynamic translation software layer that uses a speculative instruction-fusion optimization to transform FP multiply and FP add sequences into CMA instructions. The FMA unit has been slightly modified to support both single-rounding and double-rounding fused instructions without increasing their latency and to provide a conservative fall-back path in case of mispeculation. Evaluation on a cycle-accurate timing simulator showed that CMA improved SPECfp performance by 6.3% and reduced executed instructions by 4.7%. Marc Lupon, Enric Gibert, Grigorios Magklis, Sridhar Samudrala, Raúl Martínez, Kyriakos Stavrou, David R. Ditzel |
ASPLOS | 6 |
| 2014 | Warm-Up Simulation Methodology for HW/SW Co-Designed Processors
Aleksandar Brankovic, Kyriakos Stavrou, Enric Gibert, Antonio González 0001 |
CGO | 2 |
| 2009 | Programming Abstractions and Toolchain for Dataflow Multithreading ArchitecturesabstractThe need to exploit multi-core systems for parallel processing has revived the concept of dataflow. In particular, the dataflow multithreading architectures have proven to be good candidates for these systems. In this work we propose an abstraction layer that enables compiling and running a program written for an abstract dataflow multithreading architecture on different implementations. More specifically, we present a set of compiler directives that provide the programmer with the means to express most types of dependencies between code segments. In addition, we present the corresponding toolchain that transforms this code into a form that can be compiled for different implementations of the model. As a case study for this work, we present the usage of the toolchain for the TFlux and DTA architectures. Kyriakos Stavrou, Demos Pavlou, Marios Nikolaides, Panayiotis Petrides, Paraskevas Evripidou, Pedro Trancoso, Zdravko Popovic, Roberto Giorgi |
ISPDC | 1 |
| 2008 | TFlux: A Portable Platform for Data-Driven Multithreading on Commodity Multicore SystemsabstractIn this paper we present thread flux (TFlux), a complete system that supports the data-driven multithreading (DDM) model of execution. TFlux virtualizes any details of the underlying system therefore offering the same programming model independently of the architecture. To achieve this goal, TFlux has a runtime support that is built on top of a commodity operating system. Scheduling of threads is performed by the thread synchronization unit (TSU), which can be implemented either as a hardware or a software module. In addition, TFlux includes a preprocessor that, along with a set of simple compiler directives, allows the user to easily develop DDM programs. The preprocessor then automatically produces the TFlux code, which can be compiled using any commodity C compiler, therefore automatically producing code to any ISA. TFlux has been validated on three platforms. A Simics-based multicore system with a TSU hardware module (TFluxHard), a commodity 8-core Intel Core2 QuadCore-based system with a software TSU module (TFluxSoft), and a Cell/BE system with a software TSU module (TFluxCell). The experimental results show that the performance achieved is close to linear speedup, on average 21x for the 27 nodes TFluxHard, and 4.4x on a 6 nodes TFluxSoft and TFluxCell. Most importantly, the observed speedup is stable across the different platforms thus allowing the benefits of DDM to be exploited on different commodity systems. Kyriakos Stavrou, Marios Nikolaides, Demos Pavlou, Samer Arandi, Paraskevas Evripidou, Pedro Trancoso |
ICPP | 1 |
| 2008 | HelperCoreDB: Exploiting multicore technology to improve database performanceabstractDue to limitations in the traditional microprocessor design, such as high complexity and power, all current commercial high-end processors contain multiple cores on the same chip (multicore). This trend is expected to continue resulting in increasing number of cores on the chip. While these cores may be used to achieve higher throughput, improving the execution of a single application requires careful coding and usually results in smaller benefit as the scale of the system is increased. In this work we propose an alternative way to exploit the multicore technology where certain cores execute code which indirectly improves the performance of the application. We call this approach the Helper Core approach. The main contribution of this work is the proposal and evaluation of HelperCoreDB, a Helper Core designed to improve the performance of database workloads by performing efficient data prefetching. We validate the proposed approach using a HelperCoreDBimplementation for the PostgreSQL DBMS. The approach is evaluated with native execution on a dual core processor system using the standard TPC-H benchmark. The experimental results show a significant performance improvement specially considering that the baseline system is a modern optimized processor that includes advanced features such as hardware prefetching. In particular, for the 22 queries of TPC-H, our approach achieves a large reduction of the secondary cache misses, 75% on average, and an improvement in the execution time of up to 21.5%. Kostas Papadopoulos, Kyriakos Stavrou, Pedro Trancoso |
IPDPS | 2 |
| 2007 | HelperCore_DB: Exploiting Multicore Technology for Databases
Kostas Papadopoulos, Kyriakos Stavrou, Pedro Trancoso |
PACT | 2 |
| 2006 | Thermal-Aware Scheduling: A Solution for Future Chip Multiprocessors Thermal ProblemsabstractThe increased complexity and operating frequency in current microprocessors is resulting in a decrease in the performance improvements. In order to keep up with the expected performance gains, major manufacturers have started to offer chip-multiprocessor architectures. Nevertheless, the integration of several cores on the same chip leads to increased heat dissipation and consequently additional costs, decrease of the reliability, and performance loss, among others. In this paper we propose thermal-aware scheduling (TAS) a technique that aims to minimize all these problems. When assigning processes to cores, TAS takes their temperature into account avoiding thermal violation events. As a side effect, the performance is improved. Simulation results show that for a 25-core CMP, a simple TAS heuristic reduces the performance loss that is introduced by excessive temperature, from 52% to 18%. At the same time, TAS decreases the chip's temperature by 2.6degC Kyriakos Stavrou, Pedro Trancoso |
DSD | 1 |