EDBT 2026 Demo / reviewers in the wild / expert
Georgia Antoniou
dblp:226/4147
· DBLP profile ↗
8ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0000-9027-2031ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging control-flow similarity to reduce branch predictor cold effects in microservicesabstractModern datacenter applications commonly adopt a microservice software architecture, where an application is decomposed into smaller interconnected microservices communicating via the network.These microservices often operate under strict latency requirements, rendering them particularly vulnerable to microarchitectural cold effects that may arise from the interleaved execution of services on cores or power-gating cores between invocations.Previous analyses of microservices find branch mispredictions due to cold predictor resources to be a significant contributor to performance degradation, indicating that the dynamic control flow must be very similar in the set and order of executed instructions across different requests.Our analysis of control-flow similarity across requests, using static and dynamic control flow information to determine dynamic control flow reconvergence, confirms that, indeed, a large portion of requests follow similar paths.Motivated by the above findings, we propose Similarity-based Branch Prediction (SBP), a hybrid predictor architecture that enhances conventional predictors with a similarity component.SBP leverages the control-flow similarity across microservice requests to predict control flow (branch direction and target) by utilizing the control flow of past executions encoded in a reference execution trace.We realize a specific instantiation of SBP, called CHESS, which combines a conventional history-based fetch predictor, a static-hint predictor, and a similarity predictor.CHESS judiciously applies similarity prediction for branches identified as hard-to-predict through conventional prediction techniques, effectively mitigating branch predictor cold-start effects while keeping the length of the reference trace practical.Evaluation through a suite of microservices shows that CHESS reduces branch MPKI by 94% over a cold fetch predictor and 78% over a state-of-the-art predictor, while requiring a modest 18.1KB of additional storage space.This enables CHESS to deliver performance that is, on average, within 95% of a warm baseline system. Haris Volos 0001, Stylianos Vassiliou, Georgia Antoniou, Davide B. Bartolini, Yiannakis Sazeides |
ISCA | 3 |
| 2025 | SAGA: A Surrogate Assisted Genetic Algorithm for Fast CPU Power Virus GenerationabstractPower viruses are stress programs useful for investigating power delivery, thermal and cooling challenges of processors. They are used for design-time optimization but also for in-the-field diagnosis and configuration. While it is possible to construct a power virus manually, the increasing complexity of modern CPUs makes this task highly challenging. Crafting a power virus demands extensive microarchitectural expertise, making the process both difficult and time-intensive. Automated frameworks, such as those based on Genetic Algorithms, address these challenges by reducing the reliance on manual expertise. However, they often suffer from long search times due to the large number of power measurements required to assign fitness values to each candidate solution and explore the virus search space. In this paper, we propose SAGA: a Surrogate-Assisted-Genetic-Algorithm framework that reduces the number of required power measurements. Our method employs a low cost to train surrogate function that helps to predict the values of features of a strong power virus. This ultimately enables the ranking of candidate power viruses without actually measuring their power. The experimental results using real hardware reveal that SAGA can reduce a power virus search time by up to 2x as compared to a state-of-the-art method. The SAGA efficacy is demonstrated on four different CPUs, two ARM and two x86, establishing the robustness and generality of SAGA across architectures. Panteleimonas Chatzimiltis, Georgia Antoniou, Haris Volos 0001, Yiannakis Sazeides |
ISPASS | 2 |
| 2024 | Agile C-states: A Core C-state Architecture for Latency Critical Applications Optimizing both Transition and Cold-Start LatencyabstractLatency-critical applications running in modern datacenters exhibit irregular request arrival patterns and are implemented using multiple services with strict latency requirements (30–250μs). These characteristics render existing energy-saving idle CPU sleep states ineffective due to the performance overhead caused by the state’s transition latency. Besides the state transition latency, another important contributor to the performance overhead of sleep states is the cold-start latency, or in other words, the time required to warm up the microarchitectural state (e.g., cache contents, branch predictor metadata) that is flushed or discarded when transitioning to a lower-power state. Both the transition latency and cold-start latency can be particularly detrimental to the performance of latency critical applications with short execution times. While prior work focuses on mitigating the effects of transition and cold-start latency by optimizing request scheduling, in this work we propose a redesign of the core C-state architecture for latency-critical applications. In particular, we introduce C6Awarm, a new Agile core C-state that drastically reduces the performance overhead caused by idle sleep state transition latency and cold-start latency while maintaining significant energy savings. C6Awarm achieves its goals by (1) implementing medium-grained power gating, (2) preserving the microarchitectural state of the core, and (3) keeping the clock generator and PLL active and locked. Our analysis for a set of microservices based on an Intel Skylake server shows that C6Awarm manages to reduce the energy consumption by up to 70% with limited performance degradation (at most 2%). Georgia Antoniou, Davide B. Bartolini, Haris Volos 0001, Marios Kleanthous, Zhe Wang 0023, Kleovoulos Kalaitzidis, Tom Rollet, Onur Mutlu, Yiannakis Sazeides, Jawad Haj-Yahya |
ACM Trans. Archit. Code Optim. | 1 |
| 2022 | AgilePkgC: An Agile System Idle State Architecture for Energy Proportional Datacenter ServersabstractModern user-facing applications deployed in datacenters use a distributed system architecture that exacerbates the latency requirements of their constituent microservices (30-250$\mu$s). Existing CPU power-saving techniques degrade the performance of these applications due to the long transition latency (order of 100$\mu$s) to wake up from a deep CPU idle state (C-state). For this reason, server vendors recommend only enabling shallow core C-states (e.g., CC1) for idle CPU cores, thus preventing the system from entering deep package C-states (e.g., PC6) when all CPU cores are idle. This choice, however, impairs server energy proportionality since power-hungry resources (e.g., IOs, uncore, DRAM) remain active even when there is no active core to use them. As we show, it is common for all cores to be idle due to the low average utilization (e.g., 5-20%) of datacenter servers running user-facing applications. We propose to reap this opportunity with AgilePkgC (APC), a new package C-state architecture that improves the energy proportionality of server processors running latency-critical applications. APC implements PC 1A (package C l agile), a new deep package C-state that a system can enter once all cores are in a shallow C-state (i.e., CC1) and has a nanosecond-scale transition latency. PC 1A is based on four key techniques. First, a hardware-based agile power management unit (APMU) rapidly detects when all cores enter a shallow core C-state (CC1) and triggers the system-level power savings control flow. Second, an IO Standby Mode (IOSM) places IO interfaces (e.g., PCIe, DMI, UPI, DRAM) in shallow (nanosecond-scale transition latency) low-power modes. Third, a CLM Retention (CLMR) mode rapidly reduces the CLM (Cache-and-home-agent, Last-level-cache, and Mesh network-on-chip) domain’s voltage to its retention level, drastically reducing its power consumption. Fourth, APC keeps all system PLLs active in PC 1A to allow nanosecond-scale exit latency by avoiding PLL re-locking overhead. Combining these techniques enables significant power savings while requiring less than 200ns transition latency, $\gt250\times$ faster than existing deep package C-states (e.g., PC6), making PC 1A practical for datacenter servers. Our evaluation based on an Intel Skylake-based server shows that APC reduces the energy consumption of Memcached by up to 41% (25% on average) with <0.1% performance degradation. APC provides similar benefits for other representative workloads. Georgia Antoniou, Haris Volos 0001, Davide B. Bartolini, Tom Rollet, Yiannakis Sazeides, Jawad Haj-Yahya |
MICRO | 1 |
| 2022 | AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server ApplicationsabstractUser-facing applications running in modern datacenters exhibit irregular request patterns and are implemented using a multitude of services with tight latency requirements (30–250$\mu$s). These characteristics render existing energy-conserving techniques ineffective when processors are idle due to the long transition time (order of 100$\mu$s) from a deep CPU core idle power state (C-state). While prior works propose management techniques to mitigate this inefficiency, we tackle it at its root with AgileWatts (AW): a new deep CPU core C-state architecture optimized for datacenter server processors targeting latency-sensitive applications.AW drastically reduces the transition latency from deep CPU core idle power states while retaining most of their power savings based on three key ideas. First, AW eliminates the latency (several microseconds) of savinglrestoring the core context when powering-off/-on the core in a deep idle state by i) implementing medium-grained power-gates, carefully distributed across the CPU core, and ii) reraining context in the power-ungated domain. Second, AW eliminates rhe flush latency (several tens of microseconds) of the LllL2 caches when entering a deep idle state by keeping LllL2 content power-ungated. A small control logic also remains ungated to serve cache coherence traffic. AW implements cache sleep-mode and leakage reduction for the power-ungated domain by lowering a core’s voltage to the minimum operational level. Third, using a state-of-the-art power efficient all-digital phase-locked loop (ADPLL) clock generator, AW keeps the PLL active and locked during the idle state, cutting microseconds of wake-up latency at negligible power cost.Our evaluation with an accurate industrial-grade simulator calibrated against an Intel Skylake server shows that AW reduces the energy consumprion of Memcached by up to 71% (35% on average) with<1% end-to-end performance degradation. We observe similar trends for other evaluated services (MySQL and Kafka). AW’s new deep C-states C6A and C6AE reduce transition-time by up to 900$\times$ as compared to the deepest existing idle state C6, while consuming only 7% and 5% of the active state (C0) power, respectively. Jawad Haj-Yahya, Haris Volos 0001, Davide B. Bartolini, Georgia Antoniou, Jeremie S. Kim, Zhe Wang 0023, Kleovoulos Kalaitzidis, Tom Rollet, Ye Geng, Onur Mutlu, Yiannakis Sazeides |
MICRO | 4 |
| 2020 | Performance Characterization of Simultaneous Multi-Threading and Index Partitioning for an Online Document Search ApplicationabstractThis work reports on the results of a performance characterization of Simultaneous Multi-Threading (SMT) and Index-Partitioning (IP) when executing an online document search application. One of the key paper findings is that SMT is suitable for a latency sensitive application. In particular, our analysis shows that while SMT degrades single-thread execution latency, it still improves overall response-latency (queueing plus execution latency) because the multiple SMT contexts increase available throughput and help decrease sufficiently queuing latency. SMT is particularly effective in reducing the queueing latency induced by IP. Additionally, we find that in every situation we have evaluated combining SMT and IP yields the same or better average and tail overall response-latency for the application, dataset and server type used in this study. This is true for both single and dual socket server configurations we have evaluated. Our analysis for the dual socket configuration reveals that IP benefits increase with higher utilization. This is in contrast to the typical belief that IP benefits diminish with higher utilization. This is a result of judicious pinning of index-threads and index-data to the hardware to reduce remote memory access. Georgia Antoniou, Zacharias Hadjilambrou, Yiannakis Sazeides |
ISPASS | 1 |
| 2019 | Comprehensive Characterization of an Open Source Document Search EngineabstractThis work performs a thorough characterization and analysis of the open source Lucene search library. The article describes in detail the architecture, functionality, and micro-architectural behavior of the search engine, and investigates prominent online document search research issues. In particular, we study how intra-server index partitioning affects the response time and throughput, explore the potential use of low power servers for document search, and examine the sources of performance degradation ands the causes of tail latencies. Some of our main conclusions are the following: (a) intra-server index partitioning can reduce tail latencies but with diminishing benefits as incoming query traffic increases, (b) low power servers given enough partitioning can provide same average and tail response times as conventional high performance servers, (c) index search is a CPU-intensive cache-friendly application, and (d) C-states are the main culprits for performance degradation in document search. Zacharias Hadjilambrou, Marios Kleanthous, Georgia Antoniou, Antoni Portero, Yiannakis Sazeides |
ACM Trans. Archit. Code Optim. | 3 |
| 2018 | Towards Joint Land Cover and Crop Type Mapping with Numerous ClassesabstractThe detailed, accurate and frequent land cover and crop-type mapping emerge as essential for several scientific communities and geospatial applications. This paper presents a methodology for the semi-automatic production of land cover and crop type maps using a highly analytic nomenclature of more than 40 classes. An intensive manual annotation procedure was carried out for the production of reference data. A class nomenclature based on CORINE land cover Level-3 was employed along with several additional crop-type classes. Multitemporal surface reflectance Landsat-8 data for the year of 2016 were used for all classification experiments with a linear SVM classifier. Quantitative and qualitative evaluation highlighted the efficiency of the proposed approach achieving high accuracy rates. Further analysis on individual classes' performance highlighted the challenges in the proposed classification scheme as well as important outcomes regarding the spectral behavior of the considered categories. Christina Karakizi, Georgia Antoniou, Konstantinos Karantzalos |
IGARSS | 2 |