Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Vasileios Spiliopoulos 0001

dblp:53/8110-1 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-authorSoftware engineering, systems software and programming languages · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction scheduling
0.312018
Static Instruction Scheduling for High Performance on Limited Hardware · IEEE Trans. Computers 2018
Processor architecture and microarchitecture › instruction scheduling
static instruction scheduling
0.312018
Static Instruction Scheduling for High Performance on Limited Hardware · IEEE Trans. Computers 2018
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor
0.112018
Static Instruction Scheduling for High Performance on Limited Hardware · IEEE Trans. Computers 2018

Methods — techniques the papers use, named apart from their topics

load clustering · 0.3compiler optimization · 0.3
YearPublicationVenuePosition
2018 Static Instruction Scheduling for High Performance on Limited Hardware
abstract
Complex out-of-order (OoO) processors have been designed to overcome the restrictions of outstanding long-latency misses at the cost of increased energy consumption. Simple, limited OoO processors are a compromise in terms of energy consumption and performance, as they have fewer hardware resources to tolerate the penalties of long-latency loads. In worst case, these loads may stall the processor entirely. We present Clairvoyance, a compiler based technique that generates code able to hide memory latency and better utilize simple OoO processors. By clustering loads found across basic block boundaries, Clairvoyance overlaps the outstanding latencies to increases memory-level parallelism. We show that these simple OoO processors, equipped with the appropriate compiler support, can effectively hide long-latency loads and achieve performance improvements for memory-bound applications. To this end, Clairvoyance tackles (i) statically unknown dependencies, (ii) insufficient independent instructions, and (iii) register pressure. Clairvoyance achieves a geomean execution time improvement of 14 percent for memory-bound applications, on top of standard O3 optimizations, while maintaining compute-bound applications' high-performance.
Kim-Anh Tran, Trevor E. Carlson, Konstantinos Koukos, Magnus Själander, Vasileios Spiliopoulos 0001, Stefanos Kaxiras, Alexandra Jimborean
IEEE Trans. Computers5
2017 Clairvoyance: look-ahead compile-time scheduling
Kim-Anh Tran, Trevor E. Carlson, Konstantinos Koukos, Magnus Själander, Vasileios Spiliopoulos 0001, Stefanos Kaxiras, Alexandra Jimborean
CGO5
2016 Multiversioned decoupled access-execute: the key to energy-efficient compilation of general-purpose programs
abstract
Computer architecture design faces an era of great challenges in an attempt to simultaneously improve performance and energy efficiency. Previous hardware techniques for energy management become severely limited, and thus, compilers play an essential role in matching the software to the more restricted hardware capabilities. One promising approach is software decoupled access-execute (DAE), in which the compiler transforms the code into coarse-grain phases that are well-matched to the Dynamic Voltage and Frequency Scaling (DVFS) capabilities of the hardware. While this method is proved efficient for statically analyzable codes, general-purpose applications pose significant challenges due to pointer aliasing, complex control flow and unknown runtime events. We propose a universal compile-time method to decouple general-purpose applications, using simple but efficient heuristics. Our solutions overcome the challenges of complex code and show that automatic decoupled execution significantly reduces the energy expenditure of irregular or memory-bound applications and even yields slight performance boosts. Overall, our technique achieves over 20% on average energy-delay-product (EDP) improvements (energy over 15% and performance over 5%) across 14 benchmarks from SPEC CPU 2006 and Parboil benchmark suites, with peak EDP improvements surpassing 70%.
Konstantinos Koukos, Per Ekemark, Georgios Zacharopoulos 0001, Vasileios Spiliopoulos 0001, Stefanos Kaxiras, Alexandra Jimborean
CC4
2016 Keep it cool and in time: With runtime monitoring to thermal-aware execution speeds for deadline constrained systems
Kai Lampka, Björn Forsberg, Vasileios Spiliopoulos 0001
J. Parallel Distributed Comput.3
2014 Fix the code. Don't tweak the hardware: A new compiler approach to Voltage-Frequency scaling
Alexandra Jimborean, Konstantinos Koukos, Vasileios Spiliopoulos 0001, David Black-Schaffer, Stefanos Kaxiras
CGO3
2013 Towards more efficient execution: a decoupled access-execute approach
abstract
The end of Dennard scaling is expected to shrink the range of DVFS in future nodes, limiting the energy savings of this technique. This paper evaluates how much we can increase the effectiveness of DVFS by using a software decoupled access-execute approach. Decoupling the data access from execution allows us to apply optimal voltage-frequency selection for each phase and therefore improve energy efficiency over standard coupled execution.
Konstantinos Koukos, David Black-Schaffer, Vasileios Spiliopoulos 0001, Stefanos Kaxiras
ICS3
2013 Introducing DVFS-Management in a Full-System Simulator
abstract
Dynamic Voltage and Frequency Scaling (DVFS) is an essential part of controlling the power consumption of any computer system, ranging from mobile phones to servers. DVFS efficiency relies on hardware-software co-optimization, thus using existing hardware cannot reveal the full optimization potential beyond the current implementation's characteristics. To explore the vast design space for DVFS efficiency, that straddles software and hardware, a simulation infrastructure must provide features that are not readily available today, for example: software controllable clock and voltage domains, support for the OS and the frequency scaling module of it, and an online power estimation methodology. As the main contribution, this work enables DVFS studies in a full-system simulator. We extend the gem5 simulator to support full-system DVFS modeling. By doing so, we enable energy-efficiency experiments to be performed in gem5 and we showcase such studies. Finally, we show that both existing and novel frequency governors for Linux and Android can be effortlessly integrated in the framework, and we evaluate the efficiency of different DVFS schemes.
Vasileios Spiliopoulos 0001, Akash Bagdia, Andreas Hansson 0001, Peter Aldworth, Stefanos Kaxiras
MASCOTS1
2012 Power-Sleuth: A Tool for Investigating Your Program's Power Behavior
abstract
Modern processors support aggressive power saving techniques to reduce energy consumption. However, traditional profiling techniques have mainly focused on performance, which does not accurately reflect the power behavior of applications. For example, the longest running function is not always the most energy-hungry function. Thus software developers cannot always take full advantage of these power-saving features. We present Power-Sleuth, a power/performance estimation tool which is able to provide a full description of an application's behavior for any frequency from a single profiling run. The tool combines three techniques: a power and a performance estimation model with a program phase detection technique to deliver accurate, per-phase, per-frequency analysis. Our evaluation (against real power measurements) shows that we can accurately predict power and performance across different frequencies with average errors of 3.5% and 3.9% respectively.
Vasileios Spiliopoulos 0001, Andreas Sembrant, Stefanos Kaxiras
MASCOTS1
2011 Poster: DVFS management in real-processors
abstract
We describe a framework for run-time adaptive dynamic voltage-frequency scaling in Linux systems. Our underlying methodology is based on a simple first-order processor performance model in which frequency scaling is expressed as a change (in cycles) of the main memory latency. Utilizing available performance monitoring hardware, we show that our model is powerful enough to i) predict with reasonable accuracy the effect of frequency scaling, and ii) predict the energy consumed by the core under different V/f combinations. To validate our approach we perform highly accurate, fine grained power measurements directly on the processor off-chip voltage regulator.
Vasileios Spiliopoulos 0001, Georgios Keramidas, Stefanos Kaxiras, Konstantinos Efstathiou 0002
ICS1