EDBT 2026 Demo / reviewers in the wild / expert
Thaleia Dimitra Doudali
dblp:207/5379
· DBLP profile ↗
7ranked-venue papers
6as first author
4since 2021 · last 2023
0000-0002-3197-839XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Is Machine Learning Necessary for Cloud Resource Usage Forecasting?abstractRobust forecasts of future resource usage in cloud computing environments enable high efficiency in resource management solutions, such as autoscaling and overcommitment policies. Production-level systems use lightweight combinations of historical information to enable practical deployments. Recently, Machine Learning (ML) models, in particular Long Short Term Memory (LSTM) neural networks, have been proposed by various works, for their improved predictive capabilities. Following this trend, we train LSTM models and observe high levels of prediction accuracy, even on unseen data. Upon meticulous visual inspection of the results, we notice that although the predicted values seem highly accurate, they are nothing but versions of the original data shifted by one time step into the future. Yet, this clear shift seems to be enough to produce a robust forecast, because the values are highly correlated across time. We investigate time series data of various resource usage metrics (CPU, memory, network, disk I/O) across different cloud providers and levels, such as at the physical or virtual machine-level and at the application job-level. We observe that resource utilization displays very small variations in consecutive time steps. This insight can enable very simple solutions, such as data shifts, to be used for cloud resource forecasting and deliver highly accurate predictions. This is the reason why we ask whether complex machine learning models are even necessary to use. We envision that practical resource management systems need to first identify the extent to which simple solutions can be effective, and resort to using machine learning to the extent that enables its practical use. Georgia Christofidi, Konstantinos Papaioannou 0002, Thaleia Dimitra Doudali |
SoCC | 3 |
| 2022 | Coeus: Clustering (A)like Patterns for Practical Machine Intelligent Hybrid Memory ManagementabstractEmerging workloads benefit from massive memory capacities provided by hybrid memory platforms. Recent system-level hybrid memory management solutions integrate machine learning methods to better predict complex data access behaviors. Given the substantial associated learning overheads, such solutions train parallel recurrent neural networks to learn the access patterns at the granularity of a page for a carefully selected page subset. Our observation reveals that the size of this subset varies immensely across workload classes, sizes and patterns. Increasing the granularity at the level of a page group will help reduce the aggregate learning overheads. Yet, unsupervised machine learning clustering methods are not practical to use in this context. Instead, this paper builds Coeus - a page grouping mechanism for machine learning-based hybrid memory management. Coeus is simple, robust and efficient. Coeus leverages data reuse insights to fine-tune the granularity at which patterns are interpreted by the system. As a result, Coeus creates large clusters of pages that share the same access behavior, in a practical way. Coeus reduces by almost 3x the associated learning overheads. In addition, Coeus achieves 3x higher application performance, by the combined effects of applying machine learning to more pages and by performing management operations at better granularity, compared to configurations of existing hybrid memory managers. Thaleia Dimitra Doudali, Ada Gavrilovska |
CCGRID | 1 |
| 2021 | Machine Learning Augmented Hybrid Memory ManagementabstractThe integration of emerging non volatile memory hardware technologies into the main memory substrate, enables massive memory capacities at a reasonable cost in return for slower access speeds. This heterogeneity, along with the greater irregularity in the behavior of emerging workloads, render existing memory management approaches ineffective. This creates a significant gap between the realized vs. achievable performance and efficiency. At the same time, resource management solutions augmented with machine learning show great promise for fine-tuning system configuration knobs and predicting future behaviors. This thesis builds novel system-level mechanisms and reveals new insights towards the practical integration of machine learning in hybrid memory management. The specific contributions of this thesis is a machine learning augmented memory manager, coupled with insightful mechanisms to reduce the associated learning overheads and fine-tune critical operational parameters. The impact of this thesis is realizing an average of 3x application performance improvements and setting the new state-of-the-art in hybrid memory management. Thaleia Dimitra Doudali, Ada Gavrilovska |
HPDC | 1 |
| 2021 | Cori: Dancing to the Right Beat of Periodic Data Movements over Hybrid Memory SystemsabstractEmerging hybrid memory systems that comprise technologies such as Intel's Optane DC Persistent Memory, exhibit disparities in the access speeds and capacity ratios of their heterogeneous memory components. This breaks many assumptions and heuristics designed for traditional DRAM-only platforms. High application performance is feasible via dynamic data movement across memory units, which maximizes the capacity use of DRAM while ensuring efficient use of the aggregate system resources. Newly proposed solutions use performance models and machine intelligence to optimize which and how much data to move dynamically. However, the decision of when to move this data is based on empirical selection of time intervals, or left to the applications. Our experimental evaluation shows that failure to properly conFigure the data movement frequency can lead to 10%-100% performance degradation for a given data movement policy; yet, there is no established methodology on how to properly conFigure this value for a given workload, platform and policy. We propose Cori, a system-level tuning solution that identifies and extracts the necessary application-level data reuse information, and guides the selection of data movement frequency to deliver gains in application performance and system resource efficiency. Experimental evaluation shows that Cori configures data movement frequencies that provide application performance within 3% of the optimal one, and that it can achieve this up to 5 x more quickly than random or brute-force approaches. System-level validation of Cori on a platform with DRAM and Intel's Optane DC PMEM confirms its practicality and tuning efficiency. Thaleia Dimitra Doudali, Daniel Zahka, Ada Gavrilovska |
IPDPS | 1 |
| 2019 | Kleio: A Hybrid Memory Page Scheduler with Machine IntelligenceabstractThe increasing demand of big data analytics for more main memory capacity in datacenters and exascale computing environments is driving the integration of heterogeneous memory technologies. The new technologies exhibit vastly greater differences in access latencies, bandwidth and capacity compared to the traditional NUMA systems. Leveraging this heterogeneity while also delivering application performance enhancements requires intelligent data placement. We present Kleio, a page scheduler with machine intelligence for applications that execute across hybrid memory components. Kleio is a hybrid page scheduler that combines existing, lightweight, history-based data tiering methods for hybrid memory, with novel intelligent placement decisions based on deep neural networks. We contribute new understanding toward the scope of benefits that can be achieved by using intelligent page scheduling in comparison to existing history-based approaches, and towards the choice of the deep learning algorithms and their parameters that are effective for this problem space. Kleio incorporates a new method for prioritizing pages that leads to highest performance boost, while limiting the resulting system resource overheads. Our performance evaluation indicates that Kleio reduces on average 80% of the performance gap between the existing solutions and an oracle with knowledge of future access pattern. Kleio provides hybrid memory systems with fast and effective neural network training and prediction accuracy levels, which bring significant application performance improvements with limited resource overheads, so as to lay the grounds for its practical integration in future systems. Thaleia Dimitra Doudali, Sergey Blagodurov, Abhinav Vishnu, Sudhanva Gurumurthi, Ada Gavrilovska |
HPDC | 1 |
| 2018 | Mnemo: Boosting Memory Cost Efficiency in Hybrid Memory SystemsabstractMnemo is an application profiling tool specialized for data serving and caching workloads, which retrieve data from cloud in-memory key-value stores. The increasing demand to boost application performance via in-memory data retrieval and the resulting spike in the overall system hosting cost, lead to the promise that cheaper but slower memory technologies, such as NVDIMMs (Non Volatile Memory), are going to co-exist with the currently predominant ones, i.e. DRAM. In such future cloud systems where the memory substrate is going to include heterogeneous hardware, Mnemo comes as the necessary memory sizing and data tiering consultant. Mnemo permits quick exploration of the trade-offs between the system cost and application performance, due to the various possible sizings of the hybrid memory system components. Thaleia Dimitra Doudali, Ada Gavrilovska |
SoCC | 1 |
| 2017 | Spaten: A spatio-temporal and textual big data generatorabstractSocial networking users have the ability to check into Points of Interest (POIs) and associate location with their posts or tweets, leading to the creation of Geo-Social Networks (GeoSNs). There are many systems that aim to efficiently store and analyze plain and socially enhanced spatio-temporal data. A proper evaluation of these systems should be done using real data from popular GeoSNs, such as Foursquare, Facebook, etc. However, privacy restrictions prohibit the access to such real data in a large scale. Therefore, evaluations are done using real or synthetic data sets that include either only spatio-textual data (e.g. tweets) or plain spatial data (e.g. GPS traces) that are not socially enhanced. In this paper, we present Spaten, an open-source configurable spatio-temporal and textual data set generator, that extracts GPS traces from realistic routes utilizing Google Maps API, combines them with real POIs and relevant user comments crawled from TripAdvisor and makes the data available for further analysis. The injection of social properties extracted by existing Twitter graphs to the generated data along with further parameterization leads to realistic GeoSN data sets. We create and publicly offer GB-size datasets with millions of check-ins and GPS traces. As a proof of concept, we loaded the generated data into a Big Data enabled NoSQL system, and we evaluated its scalability by performing queries typically found in social networking sites. We hope that Spaten can provide the research community with the ability to generate realistic GeoSN data in a large scale, so as to properly evaluate their work. Thaleia Dimitra Doudali, Ioannis Konstantinou, Nectarios Koziris |
IEEE BigData | 1 |