Elena Kakoulli

dblp:88/9350 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
2since 2021 · last 2023
0000-0003-1489-807XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-authorDatabases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Storage systems · 47% Cloud and datacenter computing · 17% Interconnection networks and networks-on-chip · 14%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › storage hierarchy
tiered storage
2.252023
Cost-based Data Prefetching and Scheduling in Big Data Platforms over Tiered Storage Systems · ACM Trans. Database Syst. 2023
Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms · Proc. VLDB Endow. 2021
Automating Distributed Tiered Storage Management in Cluster Computing · Proc. VLDB Endow. 2019
Cloud and datacenter computing
cluster resource management and scheduling
1.222023
Cost-based Data Prefetching and Scheduling in Big Data Platforms over Tiered Storage Systems · ACM Trans. Database Syst. 2023
Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms · Proc. VLDB Endow. 2021
Storage systems › file systems
distributed file system
1.032019
Automating Distributed Tiered Storage Management in Cluster Computing · Proc. VLDB Endow. 2019
OctopusFS in Action: Tiered Storage Management for Data Intensive Computing · Proc. VLDB Endow. 2018
OctopusFS: A Distributed File System with Tiered Storage Management · SIGMOD Conference 2017
Storage systems
data placement
0.722019
Automating Distributed Tiered Storage Management in Cluster Computing · Proc. VLDB Endow. 2019
OctopusFS in Action: Tiered Storage Management for Data Intensive Computing · Proc. VLDB Endow. 2018
Parallel and multicore computing
task scheduling
0.722023
Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms · Proc. VLDB Endow. 2021
Cost-based Data Prefetching and Scheduling in Big Data Platforms over Tiered Storage Systems · ACM Trans. Database Syst. 2023
Memory systems › cache › prefetching
data prefetching
0.712023
Cost-based Data Prefetching and Scheduling in Big Data Platforms over Tiered Storage Systems · ACM Trans. Database Syst. 2023
Memory systems
data locality
0.512021
Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms · Proc. VLDB Endow. 2021
Storage systems
data migration
0.412019
Automating Distributed Tiered Storage Management in Cluster Computing · Proc. VLDB Endow. 2019
Interconnection networks and networks-on-chip
optical network-on-chip
0.312017
Silica-Embedded Silicon Nanophotonic On-Chip Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Interconnection networks and networks-on-chip
congestion control
0.212016
A Holistic Approach Towards Intelligent Hotspot Prevention in Network-on-Chip-Based Multicores · IEEE Trans. Computers 2016
Interconnection networks and networks-on-chip
routing algorithms
0.212016
A Holistic Approach Towards Intelligent Hotspot Prevention in Network-on-Chip-Based Multicores · IEEE Trans. Computers 2016
Cloud and datacenter computing
big data platform
0.112021
Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms · Proc. VLDB Endow. 2021
Cloud and datacenter computing › big data analytics
spark
0.112021
Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms · Proc. VLDB Endow. 2021
Interconnection networks and networks-on-chip
network congestion
0.112012
Intelligent Hotspot Prediction for Network-on-Chip-Based Multicore Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
High-performance computing
cluster computing
0.112019
Automating Distributed Tiered Storage Management in Cluster Computing · Proc. VLDB Endow. 2019
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112018
OctopusFS in Action: Tiered Storage Management for Data Intensive Computing · Proc. VLDB Endow. 2018
Parallel and multicore computing › parallel scheduling
locality-aware scheduling
0.112018
OctopusFS in Action: Tiered Storage Management for Data Intensive Computing · Proc. VLDB Endow. 2018
Storage systems › distributed storage
storage cluster
0.112017
OctopusFS: A Distributed File System with Tiered Storage Management · SIGMOD Conference 2017
Processor architecture and microarchitecture
chip multiprocessor
0.112016
A Holistic Approach Towards Intelligent Hotspot Prevention in Network-on-Chip-Based Multicores · IEEE Trans. Computers 2016

Methods — techniques the papers use, named apart from their topics

minimum cost maximum matching · 1.2bipartite graph pruning · 0.7pruning algorithm · 0.5bipartite graph · 0.5artificial neural network · 0.4machine learning · 0.4incremental learning · 0.4access pattern prediction · 0.4automated data-driven policy · 0.3full-system chip multiprocessor simulation · 0.3
YearPublicationVenuePosition
2023 Cost-based Data Prefetching and Scheduling in Big Data Platforms over Tiered Storage Systems
abstract
The use of storage tiering is becoming popular in data-intensive compute clusters due to the recent advancements in storage technologies. The Hadoop Distributed File System, for example, now supports storing data in memory, SSDs, and HDDs, while OctopusFS and hatS offer fine-grained storage tiering solutions. However, current big data platforms (such as Hadoop and Spark) are not exploiting the presence of storage tiers and the opportunities they present for performance optimizations. Specifically, schedulers and prefetchers will make decisions only based on data locality information and completely ignore the fact that local data are now stored on a variety of storage media with different performance characteristics. This article presents Trident, a scheduling and prefetching framework that is designed to make task assignment, resource scheduling, and prefetching decisions based on both locality and storage tier information. Trident formulates task scheduling as a minimum cost maximum matching problem in a bipartite graph and utilizes two novel pruning algorithms for bounding the size of the graph, while still guaranteeing optimality. In addition, Trident extends YARN’s resource request model and proposes a new storage-tier-aware resource scheduling algorithm. Finally, Trident includes a cost-based data prefetching approach that coordinates with the schedulers for optimizing prefetching operations. Trident is implemented in both Spark and Hadoop and evaluated extensively using a realistic workload derived from Facebook traces as well as an industry-validated benchmark, demonstrating significant benefits in terms of application performance and cluster efficiency.
Herodotos Herodotou, Elena Kakoulli
ACM Trans. Database Syst.2
2021 Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms
abstract
The recent advancements in storage technologies have popularized the use of tiered storage systems in data-intensive compute clusters. The Hadoop Distributed File System (HDFS), for example, now supports storing data in memory, SSDs, and HDDs, while OctopusFS and hatS offer fine-grained storage tiering solutions. However, the task schedulers of big data platforms (such as Hadoop and Spark) will assign tasks to available resources only based on data locality information, and completely ignore the fact that local data is now stored on a variety of storage media with different performance characteristics. This paper presents Trident, a principled task scheduling approach that is designed to make optimal task assignment decisions based on both locality and storage tier information. Trident formulates task scheduling as a minimum cost maximum matching problem in a bipartite graph and uses a standard solver for finding the optimal solution. In addition, Trident utilizes two novel pruning algorithms for bounding the size of the graph, while still guaranteeing optimality. Trident is implemented in both Spark and Hadoop, and evaluated extensively using a realistic workload derived from Facebook traces as well as an industry-validated benchmark, demonstrating significant benefits in terms of application performance and cluster efficiency.
Herodotos Herodotou, Elena Kakoulli
Proc. VLDB Endow.2
2019 Automating Distributed Tiered Storage Management in Cluster Computing
abstract
Data-intensive platforms such as Hadoop and Spark are routinely used to process massive amounts of data residing on distributed file systems like HDFS. Increasing memory sizes and new hardware technologies (e.g., NVRAM, SSDs) have recently led to the introduction of storage tiering in such settings. However, users are now burdened with the additional complexity of managing the multiple storage tiers and the data residing on them while trying to optimize their workloads. In this paper, we develop a general framework for automatically moving data across the available storage tiers in distributed file systems. Moreover, we employ machine learning for tracking and predicting file access patterns, which we use to decide when and which data to move up or down the storage tiers for increasing system performance. Our approach uses incremental learning to dynamically refine the models with new file accesses, allowing them to naturally adjust and adapt to workload changes over time. Our extensive evaluation using realistic workloads derived from Facebook and CMU traces compares our approach with several other policies and showcases significant benefits in terms of both workload performance and cluster efficiency.
Herodotos Herodotou, Elena Kakoulli
Proc. VLDB Endow.2
2018 OctopusFS in Action: Tiered Storage Management for Data Intensive Computing
abstract
The continuous improvements in memory, storage devices, and network technologies of commodity hardware introduce new challenges and opportunities in tiered storage management. Whereas past work is exploiting storage tiers in pairs or for specific applications, OctopusFS---a novel distributed file system that is aware of the underlying storage media---offers a comprehensive solution to managing multiple storage tiers in a distributed setting. OctopusFS contains auto-mated data-driven policies for managing the placement and retrieval of data across the nodes and storage tiers of the cluster. It also exposes the network locations and storage tiers of the data in order to allow higher-level systems to make locality-aware and tier-aware decisions. This demonstration will showcase the web interface of OctopusFS, which enables users to (i) view detailed utilization information for the various storage tiers and nodes, (ii) browse the directory namespace and perform file-related actions, and (iii) execute caching-related operations while observing their performance impact on MapReduce and Spark workloads.
Elena Kakoulli, Nikolaos Karmiris, Herodotos Herodotou
Proc. VLDB Endow.1
2017 OctopusFS: A Distributed File System with Tiered Storage Management
abstract
The ever-growing data storage and I/O demands of modern large-scale data analytics are challenging the current distributed storage systems. A promising trend is to exploit the recent improvements in memory, storage media, and networks for sustaining high performance and low cost. While past work explores using memory or SSDs as local storage or combine local with network-attached storage in cluster computing, this work focuses on managing multiple storage tiers in a distributed setting. We present OctopusFS, a novel distributed file system that is aware of heterogeneous storage media (e.g., memory, SSDs, HDDs, NAS) with different capacities and performance characteristics. The system offers a variety of pluggable policies for automating data management across the storage tiers and cluster nodes. The policies employ multi-objective optimization techniques for making intelligent data management decisions based on the requirements of fault tolerance, data and load balancing, and throughput maximization. At the same time, the storage media are explicitly exposed to users and applications, allowing them to choose the distribution and placement of replicas in the cluster based on their own performance and fault tolerance requirements. Our extensive evaluation shows the immediate benefits of using OctopusFS with data-intensive processing systems, such as Hadoop and Spark, in terms of both increased performance and better cluster utilization.
Elena Kakoulli, Herodotos Herodotou
SIGMOD Conference1
2017 Silica-Embedded Silicon Nanophotonic On-Chip Networks
abstract
On-chip nanophotonics offer high throughput, yet energy-efficient communication, traits that can prove critical to the continuance of multicore chip scalability. In this paper, we investigate and propose silicon nanophotonic components that are embedded entirely in the silica (SiO2) substrate, i.e., reside subsurface, as opposed to die on-surface silicon nanophotonics of prior-art. Among several offered advantages, such silicon-in-silica (SiS) nanophotonic structures empower the implementation of nonobstructive interconnect geometries that deliver an improved power-performance balance, as demonstrated experimentally. First, using exhaustive simulations based on commercial-grade optical software-based tools, we show that such SiS structures are feasible, and derive their geometry characteristics and design parameters. As a second step, utilizing SiS optical channels and filters, we then design two distinct SiS-based nanophotonic network-on-chip (PNoC) mesh-diagonal links topologies as a means of demonstrating our proof of concept. In further pushing the performance envelope, we next develop: 1) an associated contention-aware adaptive routing function and 2) a parallelized photonic channel allocation scheme, with both coupled to SiS-based PNoCs as elements, to respectively replace under-performing routing and flow-control photonic protocols currently utilized. An extensive experimental evaluation, including utilizing traffic benchmarks gathered from full-system chip multiprocessor simulations, shows that our methodology boosts network throughput by up to 59.7%, reduces communication latency by up to 78.7%, while improving the throughput-to-power ratio by up to 31.6% when compared to the state-of-the-art.
Elena Kakoulli, Vassos Soteriou, Charalambos Koutsides, Kyriacos Kalli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 A Holistic Approach Towards Intelligent Hotspot Prevention in Network-on-Chip-Based Multicores
abstract
Traffic hotspots, a severe form of network congestion, can be caused unexpectedly in a network-on-chip (NoC) due to the immanent spatio-temporal unevenness of application traffic. Hotspots reduce the NoC's effective throughput, where in the worst-case scenario, network traffic flows can be frozen indefinitely. To alleviate this problematic phenomenon several adaptive routing algorithms employ online load-balancing schemes, aiming to reduce the possibility of hotspots arising. Since most are not explicitly hotspotagnostic, they cannot completely prevent hotspot formation(s) as their reactive capability to hotspots is merely passive. This paper presents a pro-active Hotspot-Preventive Routing Algorithm (HPRA) which uses the advance knowledge gained from network embedded artificial neural network-based (ANN) hotspot predictors to guide packet routing in mitigating any unforeseen near-future hotspot occurrences. First, these ANN-based predictors are trained offline and during multicore operation they gather online statistical data to predict about-to-be-formed hotspots, promptly informing HPRA to take appropriate hotspot-preventive action(s). Next, in a holistic approach, additional ANN training is performed with data acquired after HPRA interferes, so as to further improve hotspot prediction accuracy; hence, the ANN mechanism does not only predict hotspots, but is also aware of changes that HPRA imposes upon the interconnect infrastructure. Evaluation results, including utilizing real application traffic traces gathered from parallelized workload executions onto a chip multiprocessor architecture, show that HPRA can improve network throughput up to 81 percent when compared with prior-art. Hardware synthesis results affirm the HPRA mechanism's moderate overhead requisites.
Vassos Soteriou, Theocharis Theocharides, Elena Kakoulli
IEEE Trans. Computers3
2015 Design of high-performance, power-efficient optical NoCs using Silica-embedded silicon nanophotonics
abstract
With on-chip electrical interconnects being marred by high energy-to-bandwidth costs, threatening multicore scalability, on-chip nanophotonics, which offer high throughput, yet energy-efficient communication, form an alternative attractive counterpart. In this paper we consider silicon nanophotonic components that are embedded completely within the silica (SiO2) substrate as opposed to prior-art that utilizes die on-surface silicon nanophotonics. As nanophotonic components now reside in the silica substrate's subsurface non-obstructive interconnect geometries offering higher network throughput can be implemented. First, we show using detailed simulations based on commercial optical tools that such Silicon-In-Silica (SiS) structures are feasible, derive their geometry characteristics and design parameters, and then demonstrate our proof of concept by utilizing a hybrid SiS-based photonic mesh-diagonal links network-on-chip topology. In pushing the performance envelope even more, we next develop (1) an associated contention-aware photonic adaptive routing function, and (2) a parallelized photonic channel allocation scheme, that in tandem further reduce message delivery latency. An extensive experimental evaluation, including utilizing traffic benchmarks gathered from full-system chip multiprocessor simulations, shows that our methodology boosts network throughput by up to 30.8%, reduces communication latency by up to 22.5%, and improves the throughput-to-power ratio by up to 23.7% when compared to prior-art.
Elena Kakoulli, Vassos Soteriou, Charalambos Koutsides, Kyriacos Kalli
ICCD1
2015 Designing High-Performance, Power-Efficient NoCs With Embedded Silicon-in-Silica Nanophotonics
abstract
On-chip electrical links exhibit large energy-to-bandwidth costs, whereas on-chip nanophotonics, which attain high throughput, yet energy-efficient communication, have emerged as an alternative interconnect in multicore chips. Here we consider silicon nanophotonic components that are embedded completely within the silica (SiO2) substrate as opposed to existing die on-surface silicon nanophotonics. As nanophotonic components now reside subsurface, within the silica substrate, non-obstructive interconnect geometries offering higher network throughput can be implemented. First, we show using detailed simulations based on commercial tools that such Silicon-in-Silica (SiS) structures are feasible, and then demonstrate our proof of concept by utilizing a SiS-based mesh-interconnected topology with augmented diagonal optical channels that provides both higher effective throughput and throughput-to-power ratio versus prior-art.
Elena Kakoulli, Vassos Soteriou, Charalambos Koutsides, Kyriacos Kalli
NOCS1
2014 Hermes: Architecting a top-performing fault-tolerant routing algorithm for Networks-on-Chips
abstract
Networks-on-Chips (NoCs) are experiencing escalating susceptibility to wear-out and reduced reliability, with the risk of becoming the key point of failure in an entire multicore chip. In this paper we propose Hermes, a highly-robust, distributed fault-tolerant routing algorithm, whose performance degrades gracefully with increasing faulty NoC link counts. Hermes is a deadlock-free hybrid routing algorithm, utilizing load-balanced routing on fault-free paths, while providing pre-reconfigured escape routes in the vicinity of faults. An initial experimental evaluation shows that Hermes improves network throughput by up to 2.2× when compared against the existing state-of-the-art.
Costas Iordanou, Vassos Soteriou, Konstantinos Aisopos, Elena Kakoulli
NOCS4
2012 HPRA: A pro-active Hotspot-Preventive high-performance routing algorithm for Networks-on-Chips
abstract
The inherent spatio-temporal unevenness of traffic flows in Networks-on-Chips (NoCs) can cause unforeseen, and in cases, severe forms of congestion, known as hotspots. Hotspots reduce the NoC's effective throughput, where in the worst case scenario, the entire network can be brought to an unrecoverable halt as a hotspot(s) spreads across the topology. To alleviate this problematic phenomenon several adaptive routing algorithms employ online load-balancing functions, aiming to reduce the possibility of hotspots arising. Most, however, work passively, merely distributing traffic as evenly as possible among alternative network paths, and they cannot guarantee the absence of network congestion as their reactive capability in reducing hotspot formation(s) is limited. In this paper we present a new pro-active Hotspot-Preventive Routing Algorithm (HPRA) which uses the advance knowledge gained from network-embedded Artificial Neural Network-based (ANN) hotspot predictors to guide packet routing across the network in an effort to mitigate any unforeseen near-future occurrences of hotspots. These ANNs are trained offline and during multicore operation they gather online buffer utilization data to predict about-to-be-formed hotspots, promptly informing the HPRA routing algorithm to take appropriate action in preventing hotspot formation(s). Evaluation results across two synthetic traffic patterns, and traffic benchmarks gathered from a chip multiprocessor architecture, show that HPRA can reduce network latency and improve network throughput up to 81% when compared against several existing state-of-the-art congestion-aware routing functions. Hardware synthesis results demonstrate the efficacy of the HPRA mechanism.
Elena Kakoulli, Vassos Soteriou, Theocharis Theocharides
ICCD1
2012 Intelligent Hotspot Prediction for Network-on-Chip-Based Multicore Systems
abstract
Hotspots are network-on-chip (NoC) routers or modules in multicore systems which occasionally receive packetized data from other networked element producers at a rate higher than they can consume it. This adverse phenomenon may greatly reduce the performance of NoCs, especially when wormhole flow-control is employed, as backpressure can cause the buffers of neighboring routers to quickly fill-up leading to a spatial spread in congestion. This can cause the network to saturate prematurely where in the worst scenario the NoC may be rendered unrecoverable. Thus, a hotspot prevention mechanism can be greatly beneficial, as it can potentially enable the interconnection system to adjust its behavior and prevent the rise of potential hotspots, subsequently sustaining NoC performance. The inherent unevenness of traffic patterns in an NoC-based general-purpose multicore system such as a chip multiprocessor, due to the diverse and unpredictable access patterns of applications, produces unexpected hotspots whose appearance cannot be known a priori, as application demands are not predetermined, making hotspot prediction and subsequently prevention difficult. In this paper, we present an artificial neural network-based (ANN) hotspot prediction mechanism that can be potentially used in tandem with a hotspot avoidance or congestion-control mechanism to handle unforeseen hotspot formations efficiently. The ANN uses online statistical data to dynamically monitor the interconnect fabric, and reactively predicts the location of an about to-be-formed hotspot(s), allowing enough time for the multicore system to react to these potential hotspots. Evaluation results indicate that a relatively lightweight ANN-based predictor can forecast hotspot formation(s) with an accuracy ranging from 65% to 92%.
Elena Kakoulli, Vassos Soteriou, Theocharis Theocharides
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1