VLDB 2026 Research / reviewers in the wild / expert
Nikolas Ioannou
dblp:06/7063
· DBLP profile ↗
21ranked-venue papers
5as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Storage systems · 69% Parallel and multicore computing · 19% Processor architecture and microarchitecture · 6% | |
| Artificial intelligence
3 papers |
Optimization for machine learning · 41% Kernel, tree and ensemble methods · 31% Efficient and distributed learning · 12% | |
| Computer networks
1 paper |
Datacenter networks · 100% |
Topics — the 27 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
flash and SSD |
1.1 | 3 | 2020 | Toward a Better Understanding and Evaluation of Tree Structures on Flash SSDs · Proc. VLDB Endow. 2020 FlashNet: Flash/Network Stack Co-Design · ACM Trans. Storage 2018 Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency Targets · ACM Trans. Storage 2018 |
Storage systems
key-value storage |
0.5 | 2 | 2019 | Reaping the performance of fast NVM storage with uDepot · FAST 2019 FlashNet: Flash/Network Stack Co-Design · ACM Trans. Storage 2018 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.4 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
Machine learning › Kernel, tree and ensemble methods
gradient boosting |
0.4 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
Machine learning › Optimization for machine learning › second-order optimization
newton method |
0.4 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
Storage systems › indexing
b+-tree |
0.4 | 1 | 2020 | Toward a Better Understanding and Evaluation of Tree Structures on Flash SSDs · Proc. VLDB Endow. 2020 |
Storage systems › key-value storage
LSM-tree |
0.4 | 1 | 2020 | Toward a Better Understanding and Evaluation of Tree Structures on Flash SSDs · Proc. VLDB Endow. 2020 |
Machine learning › Optimization for machine learning › coordinate descent
stochastic coordinate descent |
0.4 | 1 | 2019 | SySCD: A System-Aware Parallel Coordinate Descent Algorithm · NeurIPS 2019 |
Storage systems
non-volatile memory storage |
0.4 | 1 | 2019 | Reaping the performance of fast NVM storage with uDepot · FAST 2019 |
Parallel and multicore computing › parallel algorithms
parallel algorithm design |
0.4 | 1 | 2019 | SySCD: A System-Aware Parallel Coordinate Descent Algorithm · NeurIPS 2019 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2018 | Snap ML: A Hierarchical Framework for Machine Learning · NeurIPS 2018 |
Machine learning › Representation and self-supervised learning › representation learning
hierarchical learning |
0.3 | 1 | 2018 | Snap ML: A Hierarchical Framework for Machine Learning · NeurIPS 2018 |
Datacenter networks
RDMA |
0.3 | 1 | 2018 | FlashNet: Flash/Network Stack Co-Design · ACM Trans. Storage 2018 |
Storage systems › flash and SSD
flash memory management |
0.3 | 1 | 2018 | Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency Targets · ACM Trans. Storage 2018 |
Storage systems › flash and SSD › flash memory management
garbage collection |
0.3 | 1 | 2018 | Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency Targets · ACM Trans. Storage 2018 |
Storage systems › flash and SSD › flash memory management
wear leveling |
0.3 | 1 | 2018 | Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency Targets · ACM Trans. Storage 2018 |
Parallel and multicore computing
parallel programming models |
0.3 | 2 | 2012 | Autotuning Skeleton-Driven Optimizations for Transactional Worklist Applications · IEEE Trans. Parallel Distributed Syst. 2012 Complementing user-level coarse-grain parallelism with implicit speculative parallelism · MICRO 2011 |
Parallel and multicore computing › speculative parallelization
thread-level speculation |
0.3 | 2 | 2012 | Mixed speculative multithreaded execution models · ACM Trans. Archit. Code Optim. 2012 Complementing user-level coarse-grain parallelism with implicit speculative parallelism · MICRO 2011 |
High-performance computing › performance optimization
auto-tuning |
0.1 | 1 | 2012 | Autotuning Skeleton-Driven Optimizations for Transactional Worklist Applications · IEEE Trans. Parallel Distributed Syst. 2012 |
Processor architecture and microarchitecture › multithreading
speculative multithreading |
0.1 | 1 | 2012 | Mixed speculative multithreaded execution models · ACM Trans. Archit. Code Optim. 2012 |
Parallel and multicore computing
thread-level parallelism |
0.1 | 1 | 2012 | Mixed speculative multithreaded execution models · ACM Trans. Archit. Code Optim. 2012 |
Machine learning › Learning theory
generalization |
0.1 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2020 | Toward a Better Understanding and Evaluation of Tree Structures on Flash SSDs · Proc. VLDB Endow. 2020 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2011 | Complementing user-level coarse-grain parallelism with implicit speculative parallelism · MICRO 2011 |
Storage systems
storage reliability |
0.1 | 1 | 2018 | Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency Targets · ACM Trans. Storage 2018 |
Processor architecture and microarchitecture › multithreading
multithreaded execution |
0.0 | 1 | 2012 | Mixed speculative multithreaded execution models · ACM Trans. Archit. Code Optim. 2012 |
Energy-efficient computing › energy-efficient architecture
energy-efficient multicore |
0.0 | 1 | 2011 | Complementing user-level coarse-grain parallelism with implicit speculative parallelism · MICRO 2011 |
Methods — techniques the papers use, named apart from their topics
coordinate descent · 0.8NUMA-aware optimization · 0.8hierarchical communication · 0.7data streaming · 0.7GPU acceleration · 0.7random fourier features · 0.4performance measurement · 0.4newton descent · 0.4gradient boosting · 0.4benchmarking · 0.4read-threshold voltage optimization · 0.3data placement · 0.3cross-stack co-design · 0.3block calibration · 0.3RDMA · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Understanding modern storage APIs: a systematic study of libaio, SPDK, and io_uringabstractRecent high-performance storage devices have exposed software inefficiencies in existing storage stacks, leading to a new breed of I/O stacks. The newest storage API of the Linux kernel is io_uring. We perform one of the first in-depth studies of io_uring, and compare its performance and dis-/advantages with the established libaio and SPDK APIs. Our key findings reveal that (i) polling design significantly impacts performance; (ii) with enough CPU cores io_uring can deliver performance close to that of SPDK; and (iii) performance scalability over multiple CPU cores and devices requires careful consideration and necessitates a hybrid approach. Last, we provide design guidelines for developers of storage intensive applications. Diego Didona, Jonas Pfefferle, Nikolas Ioannou, Bernard Metzler, Animesh Trivedi |
SYSTOR | 3 |
| 2020 | Improving NAND flash performance with read heat separationabstractThe continuous growth in 3D-NAND flash storage density has primarily been enabled by 3D stacking and by increasing the number of bits stored per memory cell. Unfortunately, these desirable flash device design choices are adversely affecting reliability and latency characteristics. In particular, increasing the number of bits stored per cell results in having to apply additional voltage thresholds during each read operation, therefore increasing the read latency characteristics. While most NAND flash challenges can be mitigated through appropriate background processing, the flash read latency characteristics cannot be hidden and remains the biggest challenge, especially for the newest flash generations that store four bits per cell. In this paper, we introduce read heat separation (RHS), a new heat-aware data-placement technique that exploits the skew present in real-world workloads to place frequently read user data on low-latency flash pages. Although conceptually simple, such a technique is difficult to integrate in a flash controller, as it introduces a significant amount of complexity, requires more metadata, and is further constrained by other flash-specific peculiarities. To overcome these challenges, we propose a novel flash controller architecture supporting read heat-aware data placement. We first discuss the trade-offs that such a new design entails and analyze the key aspects that influence the efficiency of RHS. Through both, extensive simulations and an implementation we realized in a commercial enterprise-grade solid-state drive controller, we show that our architecture can indeed significantly reduce the average read latency. For certain workloads, it can reverse the system-level read latency trends when using recent multi-bit flash generations and hence outperform SSDs using previous faster flash generations. Roman A. Pletka, Nikolaos Papandreou, Radu Stoica, Haralampos Pozidis, Nikolas Ioannou, Timothy Fisher, Aaron Fry, Kip Ingram, Andrew Walls |
MASCOTS | 5 |
| 2020 | SnapBoost: A Heterogeneous Boosting MachineabstractModern gradient boosting software frameworks, such as XGBoost and LightGBM, implement Newton descent in a functional space. At each boosting iteration, their goal is to find the base hypothesis, selected from some base hypothesis class, that is closest to the Newton descent direction in a Euclidean sense. Typically, the base hypothesis class is fixed to be all binary decision trees up to a given depth. In this work, we study a Heterogeneous Newton Boosting Machine (HNBM) in which the base hypothesis class may vary across boosting iterations. Specifically, at each boosting iteration, the base hypothesis class is chosen, from a fixed set of subclasses, by sampling from a probability distribution. We derive a global linear convergence rate for the HNBM under certain assumptions, and show that it agrees with existing rates for Newton's method when the Newton direction can be perfectly fitted by the base hypothesis at each boosting iteration. We then describe a particular realization of a HNBM, SnapBoost, that, at each boosting iteration, randomly selects between either a decision tree of variable depth or a linear regressor with random Fourier features. We describe how SnapBoost is implemented, with a focus on the training complexity. Finally, we present experimental results, using OpenML and Kaggle datasets, that show that SnapBoost is able to achieve better generalization loss than competing boosting frameworks, without taking significantly longer to tune. Thomas P. Parnell, Andreea Anghel, Malgorzata Lazuka, Nikolas Ioannou, Sebastian Kurella, Peshal Agarwal, Nikolaos Papandreou, Haralampos Pozidis |
NeurIPS | 4 |
| 2020 | Toward a Better Understanding and Evaluation of Tree Structures on Flash SSDsabstractSolid-state drives (SSDs) are extensively used to deploy persistent data stores, as they provide low latency random access, high write throughput, high data density, and low cost. Tree-based data structures are widely used to build persistent data stores, and indeed they lie at the backbone of many of the data management systems used in production and research today. We show that benchmarking a persistent tree-based data structure on an SSD is a complex process, which may easily incur subtle pitfalls that can lead to an inaccurate performance assessment. At a high-level, these pitfalls stem from the interaction of complex software running on complex hardware. On the one hand, tree structures implement internal operations that have non-trivial effects on performance. On the other hand, SSDs employ firmware logic to deal with the idiosyncrasies of the underlying flash memory, which are well known to also lead to complex performance dynamics. We identify seven benchmarking pitfalls using RocksDB and WiredTiger, two widespread implementations of an LSM-Tree and a B+Tree, respectively. We show that such pitfalls can lead to incorrect measurements of key performance indicators, hinder the reproducibility and the representativeness of the results, and lead to suboptimal deployments in production environments. We also provide guidelines on how to avoid these pitfalls to obtain more reliable performance measurements, and to perform more thorough and fair comparisons among different design points. Diego Didona, Nikolas Ioannou, Radu Stoica, Kornilios Kourtis |
Proc. VLDB Endow. | 2 |
| 2019 | Reaping the performance of fast NVM storage with uDepot
Kornilios Kourtis, Nikolas Ioannou, Ioannis Koltsidas |
FAST | 2 |
| 2019 | Understanding the Design Trade-Offs of Hybrid Flash ControllersabstractOver the last few years, NAND flash manufacturers have steadily increased the number of bits stored per cell to achieve significant cost reductions. However, the increased density does not come without drawbacks. All key flash performance metrics, including latency and endurance, significantly degrade as bit density increases. Particularly, sustained write throughput is the worst affected as writes are roughly one order of magnitude slower than reads and further require precursory block erases in the background. As a result, many recent flash controllers operate flash blocks both in single-bit (high endurance and performance) and in multi-bit (high density) mode. In theory, such hybrid controllers are a great way of hiding flash technology limitations. A controller can use a small percentage of the flash blocks in single-bit mode as a cache which allows orders of magnitude higher write bandwidth and endurance in environments where the access patterns of the workload are skewed and bursty. In practice, however, many devices fall short of expectations when write performance varies significantly and utilization increases. We argue that a principled approach is required to understand the design trade-offs of hybrid NAND flash controllers. To this end, we develop a modeling framework for estimating the performance and endurance of hybrid controllers. The modeling framework computes the internal data movement generated by a hybrid controller by relying on advanced analytical models that offer both accurate and fast predictions. The data flow is then translated into higher-level metrics that quantify upper bounds for the overall performance of an SSD such as write throughput, latency, and device endurance. Using our modeling framework, we compare different controller architectures, identify their strong and weak points, and show that there is room to improve the efficiency of the hybrid controllers used today. Radu Stoica, Roman A. Pletka, Nikolas Ioannou, Nikolaos Papandreou, Sasa Tomic, Haralampos Pozidis |
MASCOTS | 3 |
| 2019 | Accelerated ML-Assisted Tumor Detection in High-Resolution Histopathology Images
Nikolas Ioannou, Milos Stanisavljevic, Andreea Anghel, Nikolaos Papandreou, Sonali Andani, Jan Hendrik Rüschoff, Peter Wild, Maria Gabrani, Haralampos Pozidis |
MICCAI (1) | 1 |
| 2019 | SySCD: A System-Aware Parallel Coordinate Descent AlgorithmabstractIn this paper we propose a novel parallel stochastic coordinate descent (SCD) algorithm with convergence guarantees that exhibits strong scalability. We start by studying a state-of-the-art parallel implementation of SCD and identify scalability as well as system-level performance bottlenecks of the respective implementation. We then take a principled approach to develop a new SCD variant which is designed to avoid the identified system bottlenecks, such as limited scaling due to coherence traffic of model sharing across threads, and inefficient CPU cache accesses. Our proposed system-aware parallel coordinate descent algorithm (SySCD) scales to many cores and across numa nodes, and offers a consistent bottom line speedup in training time of up to x12 compared to an optimized asynchronous parallel SCD algorithm and up to x42, compared to state-of-the-art GLM solvers (scikit-learn, Vowpal Wabbit, and H2O) on a range of datasets and multi-core CPU architectures. Nikolas Ioannou, Celestine Dünner, Thomas P. Parnell |
NeurIPS | 1 |
| 2018 | Elevating Commodity Storage with the SALSA Host Translation LayerabstractTo satisfy increasing storage demands in both capacity and performance, industry has turned to multiple storage technologies, including Flash SSDs and SMR disks. These devices employ a translation layer that conceals the idiosyncrasies of their mediums and enables random access. Device translation layers are, however, inherently constrained: resources on the drive are scarce, they cannot be adapted to application requirements, and lack visibility across multiple devices. As a result, performance and durability of many storage devices is severely degraded. In this paper, we present SALSA: a translation layer that executes on the host and allows unmodified applications to better utilize commodity storage. SALSA supports a wide range of single-and multi-device optimizations and, because is implemented in software, can adapt to specific workloads. We describe SALSA's design, and demonstrate its significant benefits using microbenchmarks and case studies based on three applications: MySQL, the Swift object store, and a video server. Nikolas Ioannou, Kornilios Kourtis, Ioannis Koltsidas |
MASCOTS | 1 |
| 2018 | Snap ML: A Hierarchical Framework for Machine LearningabstractWe describe a new software framework for fast training of generalized linear models. The framework, named Snap Machine Learning (Snap ML), combines recent advances in machine learning systems and algorithms in a nested manner to reflect the hierarchical architecture of modern computing systems. We prove theoretically that such a hierarchical system can accelerate training in distributed environments where intra-node communication is cheaper than inter-node communication. Additionally, we provide a review of the implementation of Snap ML in terms of GPU acceleration, pipelining, communication patterns and software architecture, highlighting aspects that were critical for achieving high performance. We evaluate the performance of Snap ML in both single-node and multi-node environments, quantifying the benefit of the hierarchical scheme and the data streaming functionality, and comparing with other widely-used machine learning software frameworks. Finally, we present a logistic regression benchmark on the Criteo Terabyte Click Logs dataset and show that Snap ML achieves the same test loss an order of magnitude faster than any of the previously reported results, including those obtained using TensorFlow and scikit-learn. Celestine Dünner, Thomas P. Parnell, Dimitrios Sarigiannis, Nikolas Ioannou, Andreea Anghel, Gummadi Ravi, Madhusudanan Kandasamy, Haralampos Pozidis |
NeurIPS | 4 |
| 2018 | Management of Next-Generation NAND Flash to Achieve Enterprise-Level Endurance and Latency TargetsabstractDespite its widespread use in consumer devices and enterprise storage systems, NAND flash faces a growing number of challenges. While technology advances have helped to increase the storage density and reduce costs, they have also led to reduced endurance and larger block variations, which cannot be compensated solely by stronger ECC or read-retry schemes but have to be addressed holistically. Our goal is to enable low-cost NAND flash in enterprise storage for cost efficiency. We present novel flash-management approaches that reduce write amplification, achieve better wear leveling, and enhance endurance without sacrificing performance. We introduce block calibration, a technique to determine optimal read-threshold voltage levels that minimize error rates, and novel garbage-collection as well as data-placement schemes that alleviate the effects of block health variability and show how these techniques complement one another and thereby achieve enterprise storage requirements. By combining the proposed schemes, we improve endurance by up to 15× compared to the baseline endurance of NAND flash without using a stronger ECC scheme. The flash-management algorithms presented herein were designed and implemented in simulators, hardware test platforms, and eventually in the flash controllers of production enterprise all-flash arrays. Their effectiveness has been validated across thousands of customer deployments since 2015. Roman A. Pletka, Ioannis Koltsidas, Nikolas Ioannou, Sasa Tomic, Nikolaos Papandreou, Thomas P. Parnell, Haralampos Pozidis, Aaron Fry, Timothy Fisher |
ACM Trans. Storage | 3 |
| 2018 | FlashNet: Flash/Network Stack Co-DesignabstractDuring the past decade, network and storage devices have undergone rapid performance improvements, delivering ultra-low latency and several Gbps of bandwidth. Nevertheless, current network and storage stacks fail to deliver this hardware performance to the applications, often due to the loss of I/O efficiency from stalled CPU performance. While many efforts attempt to address this issue solely on either the network or the storage stack, achieving high-performance for networked-storage applications requires a holistic approach that considers both. In this article, we present FlashNet, a software I/O stack that unifies high-performance network properties with flash storage access and management. FlashNet builds on RDMA principles and abstractions to provide a direct, asynchronous, end-to-end data path between a client and remote flash storage. The key insight behind FlashNet is to co-design the stack’s components (an RDMA controller, a flash controller, and a file system) to enable cross-stack optimizations and maximize I/O efficiency. In micro-benchmarks, FlashNet improves 4kB network I/O operations per second (IOPS by 38.6% to 1.22M, decreases access latency by 43.5% to 50.4μs, and prolongs the flash lifetime by 1.6-5.9× for writes. We illustrate the capabilities of FlashNet by building a Key-Value store and porting a distributed data store that uses RDMA on it. The use of FlashNet’s RDMA API improves the performance of KV store by 2× and requires minimum changes for the ported data store to access remote flash devices. Animesh Trivedi, Nikolas Ioannou, Bernard Metzler, Patrick Stuedi, Jonas Pfefferle, Kornilios Kourtis, Ioannis Koltsidas, Thomas R. Gross |
ACM Trans. Storage | 2 |
| 2017 | FlashNet: flash/network stack co-designabstractDuring the past decade, network and storage devices have undergone rapid performance improvements, delivering ultra-low latency and several Gbps of bandwidth. Nevertheless, current network and storage stacks fail to deliver this hardware performance to the applications, often due to the loss of IO efficiency from stalled CPU performance. While many efforts attempt to address this issue solely on either the network or the storage stack, achieving high-performance for networked-storage applications requires a holistic approach that considers both. Animesh Trivedi, Nikolas Ioannou, Bernard Metzler, Patrick Stuedi, Jonas Pfefferle, Ioannis Koltsidas, Kornilios Kourtis, Thomas R. Gross |
SYSTOR | 2 |
| 2012 | Mixed speculative multithreaded execution modelsabstractThe current trend toward multicore architectures has placed great pressure on programmers and compilers to generate thread-parallel programs. Improved execution performance can no longer be obtained via traditional single-thread instruction level parallelism (ILP), but, instead, via multithreaded execution. One notable technique that facilitates the extraction of parallel threads from sequential applications is thread-level speculation (TLS). This technique allows programmers/compilers to generate threads without checking for inter-thread data and control dependences, which are then transparently enforced by the hardware. Most prior work on TLS has concentrated on thread selection and mechanisms to efficiently support the main TLS operations, such as squashes, data versioning, and commits. This article seeks to enhance TLS functionality by combining it with other speculative multithreaded execution models. The main idea is that TLS already requires extensive hardware support, which when slightly augmented can accommodate other speculative multithreaded techniques. Recognizing that for different applications, or even program phases, the application bottlenecks may be different, it is reasonable to assume that the more versatile a system is, the more efficiently it will be able to execute the given program. Toward this direction, we first show that mixed execution models that combine TLS with Helper Threads (HT), RunAhead execution (RA) and MultiPath execution (MP) perform better than any of the models alone. Based on a simple model that we propose, we show that benefits come from being able to extract additional ILP without harming the TLP extracted by TLS. We then show that by combining all the execution models in a unified one that combines all these speculative multithreaded models, ILP can be further enhanced with only minimal additional cost in hardware. Polychronis Xekalakis, Nikolas Ioannou, Marcelo Cintra |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Autotuning Skeleton-Driven Optimizations for Transactional Worklist ApplicationsabstractSkeleton or pattern-based programming allows parallel programs to be expressed as specialized instances of generic communication and computation patterns. In addition to simplifying the programming task, such well structured programs are also amenable to performance optimizations during code generation and also at runtime. In this paper, we present a new skeleton framework that transparently selects and applies performance optimizations in transactional worklist applications. Using a novel hierarchical autotuning mechanism, it dynamically selects the most suitable set of optimizations for each application and adjusts them accordingly. Our experimental results on the STAMP benchmark suite show that our skeleton autotuning framework can achieve performance improvements of up to 88 percent, with an average of 46 percent, over a baseline version for a 16-core system and up to 115 percent, with an average of 56 percent, for a 32-core system. These performance improvements match or even exceed those obtained by a static exhaustive search of the optimization space. Fabrício Góes, Nikolas Ioannou, Polychronis Xekalakis, Murray Cole, Marcelo Cintra |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | Phase-Based Application-Driven Hierarchical Power Management on the Single-chip Cloud ComputerabstractTo improve energy efficiency processors allow for Dynamic Voltage and Frequency Scaling (DVFS), which enables changing their performance and power consumption on-the-fly. Many-core architectures, such as the Single-chip Cloud Computer (SCC) experimental processor from Intel Labs, have DVFS infrastructures that scale by having many more independent voltage and frequency domains on-die than today's multi-cores. This paper proposes a novel, hierarchical, and transparent client-server power management scheme applicable to such architectures. The scheme tries to minimize energy consumption within a performance window taking into consideration not only the local information for cores within frequency domains but also information that spans multiple frequency and voltage domains. We implement our proposed hierarchical power control using a novel application-driven phase detection and prediction approach for Message Passing Interface (MPI) applications, a natural choice on the SCC with its fast on-chip network and its non-coherent memory hierarchy. This phase predictor operates as the front-end to the hierarchical DVFS controller, providing the necessary DVFS scheduling points. Experimental results with SCC hardware show that our approach provides significant improvement of the Energy Delay Product (EDP) of as much as 27.2%, and 11.4% on average, with an average increase in execution time of 7.7% over a baseline version without DVFS. These improvements come from both improved phase prediction accuracy and more effective DVFS control of the domains, compared to existing approaches. Nikolas Ioannou, Michael Kauschke, Matthias Gries, Marcelo Cintra |
PACT | 1 |
| 2011 | Increasing the energy efficiency of TLS systems using intermediate checkpointingabstractWith the advent of Chip Multiprocessors (CMPs), improving performance relies on the programmers/compilers to expose thread level parallelism to the underlying hardware. However, this is a difficult and error-prone process for the programmers, while state of the art compiler techniques are unable to provide significant benefits for many classes of applications. An alternative is offered by systems that support Thread Level Speculation (TLS), which relieve the programmer and compiler from checking for thread dependences and instead use the hardware to enforce them. Unfortunately, TLS suffers from power inefficency because data misspeculations cause threads to roll back to the beginning of the speculative task. For this reason intermediate check-pointing of TLS threads has been proposed. When a violation does occur, we now have to roll back to a checkpoint before the violating instruction and not to the start of the task. However, previous work omits study of the microarchitectural details and implementation issues that are essential for effective checkpointing. In this paper we study checkpointing on a state-of-the art TLS system. We systematically study the costs associated with checkpointing and analyze the tradeoffs. We also propose changes to the TLS mechanism to allow effective checkpointing. Further, we establish the need for accurately identifying points in execution that are appropriate for checkpointing and analyze various techniques for doing so in terms of both effectiveness and viability. We propose program counter based and hybrid predictors and show that they outperform previous proposals. Placing checkpoints based on dependence predictors results in power improvements while maintaining the performance advantage of TLS. The checkpointing system proposed achieves an energy saving of up to 14%, with an average of 7% over normal TLS execution. Salman Khan 0002, Nikolas Ioannou, Polychronis Xekalakis, Marcelo Cintra |
HiPC | 2 |
| 2011 | Complementing user-level coarse-grain parallelism with implicit speculative parallelismabstractMulti-core and many-core systems are the norm in contemporary processor technology and are expected to remain so for the foreseeable future. Programs using parallel programming primitives like PThreads or OpenMP often exploit coarse-grain parallelism, because it offers a good trade-off between programming effort versus performance gain. Some parallel applications show limited or no scaling beyond a number of cores. Given the abundant number of cores expected in future many-cores, several cores would remain idle in such cases while execution performance stagnates. This paper proposes using cores that do not contribute to performance improvement for running implicit fine-grain speculative threads. In particular, we present a many-core architecture and protocol that allow applications with coarse-grain explicit parallelism to further exploit implicit speculative parallelism within each thread. Implicit speculative parallelism frees the programmer from the additional effort to explicitly partition the work into finer and properly synchronized tasks. Our results show that, for a many-core comprising of 128 cores supporting implicit speculative parallelism in clusters of 2 or 4 cores, performance improves on top of the highest scalability point by 41% on average for the 4-core cluster and by 27% on average for the 2-core cluster. These performance improvements come with an energy consumption that is close to -- and sometimes better than -- the baseline. This approach often leads to better performance and energy efficiency compared to existing alternatives such as Core Fusion and Frequency Boosting. We also investigate the tradeoffs between explicit and implicit threads as input dataset sizes vary. Finally, we present a dynamic mechanism to choose the number of explicit and implicit threads, which performs within 6% of the static oracle selection of threads. Nikolas Ioannou, Marcelo Cintra |
MICRO | 1 |
| 2010 | Profitability-based power allocation for speculative multithreaded systemsabstractWith the shrinking of transistors continuing to follow Moore's Law and the non-scalability of conventional out-of-order processors, multi-core systems are becoming the design choice for industry. Performance extraction is thus largely alleviated from the hardware and placed on the pro-gr ammer/compiler camp, who now have to expose Thread Level Parallelism (TLP) to the underlying system in the form of explicitly parallel applications. Unfortunately, parallel programming is hard and error-prone. The programmer has to parallelize the work, perform the data placement, and deal with thread synchronization. Systems that support speculative multithreaded execution like Thread Level Speculation (TLS), offer an interesting alternative since they relieve the programmer from the burden of parallelizing applications and correctly synchronizing them. Since systems that support speculative multithreading usually treat all threads equally, they are energy-inefficient. This inefficiency stems from the fact that speculation occasionally fails and, thus, power is spent on threads that will have to be discarded. In this paper we propose a power allocation scheme for TLS systems, based on Dynamic Voltage and Frequency Scaling (DVFS), that tries to remedy this inefficiency. More specifically, we propose a profitability-based power allocation scheme, where we ¿steal¿ power fro m non-profitable threads and use it to speed up more useful ones. We evaluate our techniques for a state-of-the-art TLS system and show that, with minimal hardware support, they lead to improvements in ED of up to 39.6% with an average of 21.2%, for a subset of the SPEC 2000 Integer benchmark suite. Polychronis Xekalakis, Nikolas Ioannou, Salman Khan 0002, Marcelo Cintra |
IPDPS | 2 |
| 2009 | Overlapping computation and communication in SMT clusters with commodity interconnectsabstractIn this paper we focus on optimizing the performance in a cluster of Simultaneous Multithreading (SMT) processors connected with a commodity interconnect (e.g. Gbit Ethernet), by applying overlapping of computation with communication. As a test case we consider the parallelized advection equation and discuss the steps that need to be followed to semantically allow overlapping to occur. We propose an implementation based on the concept of Helper Threading that distributes computation and communication in the two sibling threads of an SMT processor, thus creating an asymmetric pair of execution patterns in each hardware context. Our experimental results in an 8-node cluster interconnected with commodity Gbit Ethernet demonstrate that the proposed implementation is able to achieve substantial performance improvements that can exceed 20% in some cases, by efficiently utilizing the available resources of the SMT processors. Georgios I. Goumas, Nikos Anastopoulos, Nectarios Koziris, Nikolas Ioannou |
CLUSTER | 4 |
| 2009 | Combining thread level speculation helper threads and runahead executionabstractWith the current trend toward multicore architectures, improved execution performance can no longer be obtained via traditional single-thread instruction level parallelism (ILP), but, instead, via multithreaded execution.Generating thread-parallel programs is hard and thread-level speculation (TLS) has been suggested as an execution model that can speculatively exploit thread-level parallelism (TLP) even when thread independence cannot be guaranteed by the programmer/compiler. Alternatively, the helper threads (HT) execution model has been proposed where subordinate threads are executed in parallel with a main thread in order to improve the execution efficiency (i.e., ILP) of the latter. Yet another execution model, runahead execution (RA), has also been proposed where subordinate versions of the main thread are dynamically created especially to cope with long-latency operations, again with the aim of improving the execution efficiency of the main thread. Polychronis Xekalakis, Nikolas Ioannou, Marcelo Cintra |
ICS | 2 |