EDBT 2026 Demo / reviewers in the wild / expert
Andreea Anghel
dblp:57/10370 · also Andreea Simona Anghel
· DBLP profile ↗
20ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0002-6842-9036ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Performance modeling and evaluation · 40% GPUs and heterogeneous computing · 28% Processor architecture and microarchitecture · 12% | |
| Artificial intelligence
2 papers |
Kernel, tree and ensemble methods · 36% Optimization for machine learning · 32% Efficient and distributed learning · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 74% Machine learning and data management · 26% |
Topics — the 14 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
performance prediction |
0.8 | 1 | 2024 | LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services · SC 2024 |
Computational finance and economics › financial fraud detection
anti-money laundering |
0.7 | 1 | 2023 | Realistic Synthetic Financial Transactions for Anti-Money Laundering Models · NeurIPS 2023 |
Computational finance and economics
financial fraud detection |
0.7 | 1 | 2023 | Realistic Synthetic Financial Transactions for Anti-Money Laundering Models · NeurIPS 2023 |
Data integration and cleaning › data generation
synthetic data generation |
0.7 | 1 | 2023 | Realistic Synthetic Financial Transactions for Anti-Money Laundering Models · NeurIPS 2023 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.4 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
Machine learning › Kernel, tree and ensemble methods
gradient boosting |
0.4 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
Machine learning › Optimization for machine learning › second-order optimization
newton method |
0.4 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2018 | Snap ML: A Hierarchical Framework for Machine Learning · NeurIPS 2018 |
Machine learning › Representation and self-supervised learning › representation learning
hierarchical learning |
0.3 | 1 | 2018 | Snap ML: A Hierarchical Framework for Machine Learning · NeurIPS 2018 |
Electronic design automation
design space exploration |
0.3 | 1 | 2018 | Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018 |
Processor architecture and microarchitecture
multicore design |
0.3 | 1 | 2018 | Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018 |
Performance modeling and evaluation
processor performance modeling |
0.3 | 1 | 2018 | Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018 |
Machine learning › Learning theory
generalization |
0.1 | 1 | 2020 | SnapBoost: A Heterogeneous Boosting Machine · NeurIPS 2020 |
High-performance computing › supercomputing
exascale systems |
0.1 | 1 | 2018 | Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018 |
Methods — techniques the papers use, named apart from their topics
machine learning · 2.0agent-based simulation · 2.0predictive modeling · 1.5benchmarking · 1.5hierarchical communication · 0.7data streaming · 0.7GPU acceleration · 0.7random fourier features · 0.4newton descent · 0.4gradient boosting · 0.4hardware performance counters · 0.3analytic modeling · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesabstractAs Large Language Models (LLMs) are rapidly growing in popularity, LLM inference services must be able to serve requests from thousands of users while satisfying performance requirements. The performance of an LLM inference service is largely determined by the hardware onto which it is deployed, but understanding of which hardware will deliver on performance requirements remains challenging. In this work we present LLM-Pilot - a first-of-its-kind system for characterizing and predicting performance of LLM inference services. LLMPilot performs benchmarking of LLM inference services, under a realistic workload, across a variety of GPUs, and optimizes the service configuration for each considered GPU to maximize performance. Finally, using this characterization data, LLM-Pilot learns a predictive model, which can be used to recommend the most cost-effective hardware for a previously unseen LLM. Compared to existing methods, LLM-Pilot can deliver on performance requirements 33% more frequently, whilst reducing costs by 60% on average. Malgorzata Lazuka, Andreea Anghel, Thomas P. Parnell |
SC | 2 |
| 2023 | xCloudServing: Automated ML Serving Across CloudsabstractAs machine learning (ML) models have grown in complexity, so too have the expenses they incur when deployed in the cloud. In order to reduce the costs associated with ML serving, it is necessary to optimize the choice of cloud infrastructure used. Additionally, the chosen infrastructure must be able to deliver on the latency constraints that are typically defined for cloud services. This problem is made more challenging since today's organizations often need to work with more than one cloud provider, and each provider offers its own unique set of interfaces and infrastructure choices. In this work we present xCloudServing - a novel system for consistent and automated deployment of ML inference services across multi-ple cloud providers and regions. We describe the architecture and implementation of xCloudServing, as well as the different optimization algorithms implemented internally. These include established methods from the literature, as well as Niebo - our novel algorithm for minimizing cost whilst satisfying the tail latency constraint. We present simulation results for 5 different ML models over 3 cloud providers and multiple tail latency constraints that indicate that on average, Niebo outperforms state-of-the-art algorithms by 37%. Additionally, we evaluate xCloudServing with live runs and demonstrate that it is robust to nondeterministic effects and exhibits reproducible behavior. Malgorzata Lazuka, Andreea Anghel, Parikshit Ram, Haralampos Pozidis, Thomas P. Parnell |
CLOUD | 2 |
| 2023 | Acceleration of Decision-Tree Ensemble Models on the IBM Telum ProcessorabstractThis paper presents a tensor-based algorithm that leverages a hardware accelerator for inferencing decision-tree-based machine learning models. The algorithm has been integrated in a public software library and is demonstrated on an IBM z16 server, using the Telum processor with the Integrated Accelerator for AI. We describe the architecture and implementation of the algorithm and present experimental results that demonstrate its superior runtime performance compared with popular CPU-based machine learning inference implementations. Nikolaos Papandreou, Jan van Lunteren, Andreea Anghel, Thomas P. Parnell, Martin Petermann, Milos Stanisavljevic, Cédric Lichtenau, Andrew Sica, Dominic Röhm, Elpida Tzortzatos, Haralampos Pozidis |
ISCAS | 3 |
| 2023 | Realistic Synthetic Financial Transactions for Anti-Money Laundering ModelsabstractWith the widespread digitization of finance and the increasing popularity of cryptocurrencies, the sophistication of fraud schemes devised by cybercriminals is growing. Money laundering -- the movement of illicit funds to conceal their origins -- can cross bank and national boundaries, producing complex transaction patterns. The UN estimates 2-5\% of global GDP or \$0.8 - \$2.0 trillion dollars are laundered globally each year. Unfortunately, real data to train machine learning models to detect laundering is generally not available, and previous synthetic data generators have had significant shortcomings. A realistic, standardized, publicly-available benchmark is needed for comparing models and for the advancement of the area.To this end, this paper contributes a synthetic financial transaction dataset generator and a set of synthetically generated AML (Anti-Money Laundering) datasets. We have calibrated this agent-based generator to match real transactions as closely as possible and made the datasets public. We describe the generator in detail and demonstrate how the datasets generated can help compare different machine learning models in terms of their AML abilities. In a key way, using synthetic data in these comparisons can be even better than using real data: the ground truth labels are complete, whilst many laundering transactions in real data are never detected. Erik R. Altman, Jovan Blanusa, Luc von Niederhäusern, Beni Egressy, Andreea Anghel, Kubilay Atasu |
NeurIPS | 5 |
| 2022 | Search-based Methods for Multi-Cloud ConfigurationabstractMulti-cloud computing has become increasingly popular with enterprises looking to avoid vendor lock-in. While most cloud providers offer similar functionality, they may differ significantly in terms of performance and/or cost. A customer looking to benefit from such differences will naturally want to solve the multi-cloud configuration problem: given a workload, which cloud provider should be chosen and how should its nodes be configured in order to minimize runtime or cost? In this work, we consider possible solutions to this multi-cloud optimization problem. We develop and evaluate possible adaptations of state-of-the-art cloud configuration solutions to the multi-cloud domain. Furthermore, we identify an analogy between multi-cloud configuration and the selection-configuration problems that are commonly studied in the automated machine learning (AutoML) field. Inspired by this connection, we utilize popular optimizers from AutoML to solve multi-cloud configuration. Finally, we propose a new algorithm for solving multi-cloud configuration, CloudBandit. It treats the outer problem of cloud provider selection as a best-arm identification problem, in which each arm pull corresponds to running an arbitrary black-box optimizer on the inner problem of node configuration. Our extensive experiments indicate that (a) many state-of-the-art cloud configuration solutions can be adapted to multi-cloud, with best results obtained for adaptations which utilize the hierarchical structure of the multi-cloud configuration domain, (b) hierarchical methods from AutoML can be used for the multi-cloud configuration task and can outperform state-of-the-art cloud configuration solutions and (c) CloudBandit achieves competitive or lower regret relative to other tested algorithms, whilst also identifying configurations that have 65% lower median cost and 20% lower median runtime in production, compared to choosing a random provider and configuration. Malgorzata Lazuka, Thomas P. Parnell, Andreea Anghel, Haralampos Pozidis |
CLOUD | 3 |
| 2020 | SnapBoost: A Heterogeneous Boosting MachineabstractModern gradient boosting software frameworks, such as XGBoost and LightGBM, implement Newton descent in a functional space. At each boosting iteration, their goal is to find the base hypothesis, selected from some base hypothesis class, that is closest to the Newton descent direction in a Euclidean sense. Typically, the base hypothesis class is fixed to be all binary decision trees up to a given depth. In this work, we study a Heterogeneous Newton Boosting Machine (HNBM) in which the base hypothesis class may vary across boosting iterations. Specifically, at each boosting iteration, the base hypothesis class is chosen, from a fixed set of subclasses, by sampling from a probability distribution. We derive a global linear convergence rate for the HNBM under certain assumptions, and show that it agrees with existing rates for Newton's method when the Newton direction can be perfectly fitted by the base hypothesis at each boosting iteration. We then describe a particular realization of a HNBM, SnapBoost, that, at each boosting iteration, randomly selects between either a decision tree of variable depth or a linear regressor with random Fourier features. We describe how SnapBoost is implemented, with a focus on the training complexity. Finally, we present experimental results, using OpenML and Kaggle datasets, that show that SnapBoost is able to achieve better generalization loss than competing boosting frameworks, without taking significantly longer to tune. Thomas P. Parnell, Andreea Anghel, Malgorzata Lazuka, Nikolas Ioannou, Sebastian Kurella, Peshal Agarwal, Nikolaos Papandreou, Haralampos Pozidis |
NeurIPS | 2 |
| 2019 | Accelerated ML-Assisted Tumor Detection in High-Resolution Histopathology Images
Nikolas Ioannou, Milos Stanisavljevic, Andreea Anghel, Nikolaos Papandreou, Sonali Andani, Jan Hendrik Rüschoff, Peter Wild, Maria Gabrani, Haralampos Pozidis |
MICCAI (1) | 3 |
| 2018 | Snap ML: A Hierarchical Framework for Machine LearningabstractWe describe a new software framework for fast training of generalized linear models. The framework, named Snap Machine Learning (Snap ML), combines recent advances in machine learning systems and algorithms in a nested manner to reflect the hierarchical architecture of modern computing systems. We prove theoretically that such a hierarchical system can accelerate training in distributed environments where intra-node communication is cheaper than inter-node communication. Additionally, we provide a review of the implementation of Snap ML in terms of GPU acceleration, pipelining, communication patterns and software architecture, highlighting aspects that were critical for achieving high performance. We evaluate the performance of Snap ML in both single-node and multi-node environments, quantifying the benefit of the hierarchical scheme and the data streaming functionality, and comparing with other widely-used machine learning software frameworks. Finally, we present a logistic regression benchmark on the Criteo Terabyte Click Logs dataset and show that Snap ML achieves the same test loss an order of magnitude faster than any of the previously reported results, including those obtained using TensorFlow and scikit-learn. Celestine Dünner, Thomas P. Parnell, Dimitrios Sarigiannis, Nikolas Ioannou, Andreea Anghel, Gummadi Ravi, Madhusudanan Kandasamy, Haralampos Pozidis |
NeurIPS | 5 |
| 2018 | Predicting cloud performance for HPC applications before deployment
Giovanni Mariani, Andreea Anghel, Rik Jongerius, Gero Dittmann |
Future Gener. Comput. Syst. | 2 |
| 2018 | Analytic Multi-Core Processor Model for Fast Design-Space ExplorationabstractSimulators help computer architects optimize system designs. The limited performance of simulators even of moderate size and detail makes the approach infeasible for design-space exploration of future exascale systems. Analytic models, in contrast, offer very fast turn-around times. In this paper we propose an analytic multi-core processor-performance model that takes as inputs a) a parametric microarchitecture-independent characterization of the target workload, and b) a hardware configuration of the core and the memory hierarchy. The processor-performance model considers instruction-level parallelism (ILP) per type, models single instruction, multiple data (SIMD) features, and considers cache and memory-bandwidth contention between cores. We validate our model by comparing its performance estimates with measurements from hardware performance counters on Intel Xeon and ARM Cortex-A15 systems. We estimate multi-core contention with a maximum error of 11.4 percent. The average single-thread error increases from 25 percent for a state-of-the-art simulator to 59 percent for our model, but the correlation is still 0.8, a high relative accuracy, while we achieve a speedup of several orders of magnitude. With a much higher capacity than simulators and more reliable insights than back-of-the-envelope calculations it makes automated design-space exploration of exascale systems possible, which we show using a real-world case study from radio astronomy. Rik Jongerius, Andreea Anghel, Gero Dittmann, Giovanni Mariani, Erik Vermij, Henk Corporaal |
IEEE Trans. Computers | 2 |
| 2017 | Predicting Cloud Performance for HPC Applications: a User-oriented ApproachabstractCloud computing enables end users to execute high-performance computing applications by renting the required computing power. This pay-for-use approach enables small enterprises and startups to run HPC-related businesses with a significant saving in capital investment and a short time to market. When deploying an application in the cloud, the users may a) fail to understand the interactions of the application with the software layers implementing the cloud system, b) be unaware of some hardware details of the cloud system, and c) fail to understand how sharing part of the cloud system with other users might degrade application performance. These misunderstandings may lead the users to select suboptimal cloud configurations in terms of cost or performance. To aid the users in selecting the optimal cloud configuration for their applications, we suggest that the cloud provider generate a prediction model for the provided system. We propose applying machine-learning techniques to generate this prediction model. First, the cloud provider profiles a set of training applications by means of a hardware-independent profiler and then executes these applications on a set of training cloud configurations to collect actual performance values. The prediction model is trained to learn the dependencies of actual performance data on the application profile and cloud configuration parameters. The advantage of using a hardware-independent profiler is that the cloud users and the cloud provider can analyze applications on different machines and interface with the same prediction model. We validate the proposed methodology for a cloud system implemented with OpenStack. We apply the prediction model to the NAS parallel benchmarks. The resulting relative error is below 15% and the Pareto optimal cloud configurations finally found when maximizing application speed and minimizing execution cost on the prediction model are also at most 15% away from the actual optimal solutions. Giovanni Mariani, Andreea Anghel, Rik Jongerius, Gero Dittmann |
CCGrid | 2 |
| 2017 | Real-time security services for SDN-based datacentersabstractWhile the scale, frequency and impact of the recent cyber- and DoS-attacks have all increased, the traditional security management systems are still supervised by human operators in the decisional loop. To cope with the new breed of machine-driven attacks — particularly those designed to overload the humans in the loop — the next-generation anomaly detection and attack mitigation schema, i.e. the network security management, must improve greatly in speed and accuracy: become machine-driven, too. As infrastructure we propose an FPGA-accelerated Network Function Virtualization that potentially enhances the current multi-Tbps switching fabrics with SDN-based security capabilities of vastly higher performance and scalability. As key novelties, we contribute (i) sub-ms detection lag (ii) of the top 9 Akamai attacks [1] with (iii) a real-time SDN feedback loop between a distributed programmable data plane and a centralized SDN controller, (iv) coupled via a global N:1 mirror. We validate the concept in an actual datacenter network with a new security application that can detect and mitigate real-world dDoS attacks, with lags from 430 us up to 3 ms — several orders of magnitude faster than before. Pál Varga, Georgios Kathareios, Akos Mate, Rolf Clauberg, Andreea Anghel, Peter Orosz, Tamas Tothfalusi, Mitchell Gusat |
CNSM | 5 |
| 2017 | MeSAP: A fast analytic power model for DRAM memoriesabstractThe design of an energy-efficient memory subsystem is one of the key issues that system architects face today. To achieve this goal, architects usually rely on system simulators and trace-based DRAM power models. However, their long execution time makes the approach infeasible for the design-space exploration of next-generation exascale computing systems. Analytic models, in contrast, are orders of magnitude faster. In this paper, we propose a new analytic memory-scheduler-agnostic power model for DRAM, henceforth referred to as MeSAP. Similarly to state-of-the-art trace-based approaches, our analytic model achieves an average error of 20%, while being an order of magnitude faster. Furthermore, we integrate MeSAP into an analytic performance model of general-purpose processors and show its applicability to the design of a computing system targeting scientific image processing applications. Sandeep Poddar, Rik Jongerius, Leandro Fiorin, Giovanni Mariani, Gero Dittmann, Andreea Anghel, Henk Corporaal |
DATE | 6 |
| 2017 | Catch It If You Can: Real-Time Network Anomaly Detection with Low False Alarm RatesabstractUnsupervised anomaly detection (AD) has shown promise against the frequently new cyberattacks. But, as anomalies are not always malicious, such systems generate prodigious false alarm rates. The resulting manual validation workload often overwhelms the IT operators: it slows down the system reaction by orders of magnitude and ultimately thwarts its applicability. Therefore, we propose a real-time network AD system that reduces the manual workload by coupling 2 learning stages. The first stage performs adaptive unsupervised AD using a shallow autoencoder. The second stage uses a custom nearest-neighbor classifier to filter the false positives by modeling the manual classification. We implement a prototype for 10-50Gbps speeds and evaluate it with traffic from a national network operator: we achieve 98.5% true and 1.3% false positive rates, while reducing the human intervention rate by 5x. Georgios Kathareios, Andreea Anghel, Akos Mate, Rolf Clauberg, Mitchell Gusat |
ICMLA | 2 |
| 2017 | A scalable synthetic traffic model of Graph500 for computer networks analysisabstractSummary The Graph500 benchmark attempts to steer the design of High‐Performance Computing systems to maximize the performance under memory‐constricted application workloads. A realistic simulation of such benchmarks for architectural research is challenging due to size and detail limitations. By contrast, synthetic traffic workloads constitute one of the least resource‐consuming methods to evaluate the performance. In this work, we provide a simulation tool for network architects that need to evaluate the suitability of their interconnect for BigData applications. Our development is a low computation‐ and memory‐demanding synthetic traffic model that emulates the behavior of the Graph500 communications and is publicly available in an open‐source network simulator. The characterization of network traffic is inferred from a profile of several executions of the benchmark with different input parameters. We verify the validity of the equations in our model against an execution of the benchmark with a different set of parameters. Furthermore, we identify the impact of the node computation capabilities and network characteristics in the execution time of the model in a Dragonfly network. Pablo Fuentes 0001, Mariano Benito, Enrique Vallejo 0001, José Luis Bosque, Ramón Beivide, Andreea Anghel, Mitchell Gusat, Cyriel Minkenberg, Mateo Valero |
Concurr. Comput. Pract. Exp. | 6 |
| 2017 | Classification of thread profiles for scaling application behavior
Giovanni Mariani, Andreea Anghel, Rik Jongerius, Gero Dittmann |
Parallel Comput. | 2 |
| 2016 | Synthetic Traffic Model of the Graph500 Communications
Pablo Fuentes 0001, Enrique Vallejo 0001, José Luis Bosque, Ramón Beivide, Andreea Anghel, Mitchell Gusat, Cyriel Minkenberg |
ICA3PP | 5 |
| 2015 | Analytic processor model for fast design-space explorationabstractIn this paper, we propose an analytic model that takes as inputs a) a parametric microarchitecture-independent characterization of the target workload, and b) a hardware configuration of the core and the memory hierarchy, and returns as output an estimation of processor-core performance. To validate our technique, we compare our performance estimates with measurements on an Intel® Xeon® system. The average error increases from 21% for a state-of-the-art simulator to 25% for our model, but we achieve a speedup of several orders of magnitude. Thus, the model enables fast designspace exploration and represents a first step towards an analytic exascale system model. Rik Jongerius, Giovanni Mariani, Andreea Anghel, Gero Dittmann, Erik Vermij, Henk Corporaal |
ICCD | 3 |
| 2014 | Holistic power analysis of implementation alternatives for a very large scale synthesis array with phased array stationsabstractThe Square Kilometre Array (SKA) will be the largest radio telescope in the world, generating data at Pb/s rates. Real-time processing will require 1018compute operations per second and system operating costs will be dominated by energy consumption. In this paper we explore design options for the aperture array of the first SKA construction phase and provide lower bounds on their power consumption. We analyze the system's components from the antenna front-end to the central signal processor and identify the main power consumers. We compare ASIC-based and FPGA-based data processing pipelines and show that ASICs can lead to 1.6 to 4 times more power efficiency. Andreea Anghel, Rik Jongerius, Gero Dittmann, Jonas R. M. Weiss, Ronald P. Luijten |
ICASSP | 1 |
| 2014 | Scalable, efficient ASICS for the square kilometre array: From A/D conversion to central correlationabstractThe Square Kilometre Array (SKA) is a future radio telescope, currently being designed by the worldwide radio-astronomy community. During the first of two construction phases, more than 250,000 antennas will be deployed, clustered in aperture-array stations. The antennas will generate 2.5 Pb/s of data, which needs to be processed in real time. For the processing stages from A/D conversion to central correlation, we propose an ASIC solution using only three chip architectures. The architecture is scalable - additional chips support additional antennas or beams - and versatile - it can relocate its receiver band within a range of a few MHz up to 4GHz. This flexibility makes it applicable to both SKA phases 1 and 2. The proposed chips implement an antenna and station processor for 289 antennas with a power consumption on the order of 600W and a correlator, including corner turn, for 911 stations on the order of 90 kW. Martin L. Schmatz, Rik Jongerius, Gero Dittmann, Andreea Anghel, Antonius P. J. Engbersen, Jan van Lunteren, Peter Buchmann |
ICASSP | 4 |