EDBT 2026 Demo / reviewers in the wild / expert
Juan F. Pérez
dblp:61/7535 · also Juan Fernando Pérez
· DBLP profile ↗
28ranked-venue papers
9as first author
6since 2021 · last 2026
0000-0003-4732-1621ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-author · 1 since 2021Computer networks · 5 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Distributed systems · 49% Cloud and datacenter computing · 27% Performance modeling and evaluation · 19% | |
| Computer networks
1 paper |
Optical networks · 75% Wireless networking · 25% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
replication |
1.5 | 5 | 2021 | sPARE: Partial Replication for Multi-Tier Applications in the Cloud · IEEE Trans. Serv. Comput. 2021 Cutting Latency Tail: Analyzing and Validating Replication without Canceling · IEEE Trans. Parallel Distributed Syst. 2017 Power of redundancy: Designing partial replication for multi-tier applications · INFOCOM 2017 |
Distributed systems
fault tolerance |
1.0 | 4 | 2017 | Cutting Latency Tail: Analyzing and Validating Replication without Canceling · IEEE Trans. Parallel Distributed Syst. 2017 Evaluating Replication for Parallel Jobs: An Efficient Approach · IEEE Trans. Parallel Distributed Syst. 2016 Variability-aware request replication for latency curtailment · INFOCOM 2016 |
Cloud and datacenter computing › quality of service
tail latency |
0.4 | 2 | 2021 | Power of redundancy: Designing partial replication for multi-tier applications · INFOCOM 2017 sPARE: Partial Replication for Multi-Tier Applications in the Cloud · IEEE Trans. Serv. Comput. 2021 |
Performance modeling and evaluation › design trade-off analysis
latency-accuracy tradeoff |
0.3 | 1 | 2017 | On the latency-accuracy tradeoff in approximate MapReduce jobs · INFOCOM 2017 |
Performance modeling and evaluation › delay analysis
latency modeling |
0.3 | 1 | 2017 | On the latency-accuracy tradeoff in approximate MapReduce jobs · INFOCOM 2017 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.3 | 1 | 2017 | On the latency-accuracy tradeoff in approximate MapReduce jobs · INFOCOM 2017 |
Distributed systems › replication
partial replication |
0.3 | 1 | 2017 | Power of redundancy: Designing partial replication for multi-tier applications · INFOCOM 2017 |
Performance modeling and evaluation › performance prediction
response-time distribution modeling |
0.3 | 1 | 2017 | Cutting Latency Tail: Analyzing and Validating Replication without Canceling · IEEE Trans. Parallel Distributed Syst. 2017 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.2 | 1 | 2016 | Evaluating Replication for Parallel Jobs: An Efficient Approach · IEEE Trans. Parallel Distributed Syst. 2016 |
Wireless networking
collision resolution |
0.1 | 1 | 2009 | Dimensioning an OBS Switch with Partial Wavelength Conversion and Fiber Delay Lines via a Mean Field Model · INFOCOM 2009 |
Optical networks › optical buffer
fiber delay lines |
0.1 | 1 | 2009 | Dimensioning an OBS Switch with Partial Wavelength Conversion and Fiber Delay Lines via a Mean Field Model · INFOCOM 2009 |
Optical networks › optical switching
optical burst switching |
0.1 | 1 | 2009 | Dimensioning an OBS Switch with Partial Wavelength Conversion and Fiber Delay Lines via a Mean Field Model · INFOCOM 2009 |
Optical networks
wavelength conversion |
0.1 | 1 | 2009 | Dimensioning an OBS Switch with Partial Wavelength Conversion and Fiber Delay Lines via a Mean Field Model · INFOCOM 2009 |
Cloud and datacenter computing
job scheduling |
0.1 | 1 | 2017 | On the latency-accuracy tradeoff in approximate MapReduce jobs · INFOCOM 2017 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.1 | 1 | 2015 | Enhancing reliability and response times via replication in computing clusters · INFOCOM 2015 |
Performance modeling and evaluation › queueing models
mean-field analysis |
0.0 | 1 | 2009 | Dimensioning an OBS Switch with Partial Wavelength Conversion and Fiber Delay Lines via a Mean Field Model · INFOCOM 2009 |
Methods — techniques the papers use, named apart from their topics
stochastic modeling · 1.0token-based arbitration · 0.8iterative searching · 0.5variability-aware replication · 0.3matrix-analytic method · 0.3correlation modeling · 0.3queueing analysis · 0.2approximate modeling · 0.2maximum likelihood estimation · 0.2linear regression · 0.2mean-field analysis · 0.1markovian arrival process · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Differential privacy strategies for data analytics in the banking sectorabstractFinancial institutions face growing pressure to protect customer data while maintaining the ability to obtain business insights through data analysis. To tackle this challenge, Differential Privacy has risen in recent years as a promising solution to guarantee user data privacy. However, the utility of the analytical models trained employing Differential Privacy remains an open question. In this paper, we consider two main approaches: the private-model workflow, employing Centralized Differential Privacy; and the private-data workflow, implementing Local Differential Privacy. We test these two alternatives on an open bank marketing dataset, training classification models and comparing them against a non-private baseline. Findings show a sharp trade-off between privacy and utility in the private-data workflow, which allow for usable models under moderate privacy guarantees (privacy budget in the range 3-5). On the other hand, in the private-model case, adjusting training configurations can help reduce the impact of privacy constraints. Experiments show the sample size, batch size and learning rate to be the most influential hyperparameters, while the noise multiplier and the gradient clipping norm (related to the privacy mechanism) show limited impact. The results provide practical guidance to apply differential privacy in financial environments, supporting both regulatory compliance and effective data use. Daniela Espinosa, Juan F. Pérez, Valérie Gauthier |
Inf. Sci. | 2 |
| 2026 | Key recovery attack on PRESENT using an entropy-based neural distinguisherabstractIn 2019, Gohr introduced neural-distinguishers as a tool to improve differential cryptanalysis using deep learning. Building on this, we propose an Entropy-based Neural Distinguisher for present that requires significantly fewer input bits and network parameters. Our method identifies relevant bit subsets by analyzing the entropy of output differences between ciphertext pairs. We reduce the distinguisher size by: (i) simplifying the network architecture, and (ii) shrinking the input layer via entropy-based bit selection. Our distinguisher achieves accuracy within 1% (resp. 4%) of the state of the art for 7 (resp. 6) rounds of present , using only 28 out of 64 bits and under 10% of the original parameters. Leveraging this efficient setup, we introduce an iterative key recovery method capable of handling 64-bit round keys—unlike Gohr’s 16-bit target. Our approach recovers full keys (64 bits) with 92.8% success (402 out of 433), with all partial recoveries retrieving at least 56 bits and 83.8% retrieving 60 bits or more. Valérie Gauthier, Isabella Martínez, Germán Obando, Juan F. Pérez |
Neural Comput. Appl. | 4 |
| 2025 | DP-TLDM: Differentially Private Tabular Latent Diffusion Model
Chaoyi Zhu, Juan F. Pérez, Marten van Dijk, Lydia Y. Chen |
ARES (1) | 3 |
| 2025 | BatMan-CLR: Making Few-Shots Meta-learners Resilient Against Label Noise
Jeroen Galjaard, Robert Birke, Juan F. Pérez, Lydia Y. Chen |
ECML/PKDD (6) | 3 |
| 2025 | A workflow to systematically design uncertainty-aware visual analytics applicationsabstractAbstract Visual analytics (VA) is a paradigm for insight generation by using visual analysis techniques and automated reasoning by transforming data into hypotheses and visualization to extract new insights. The insights are fed back into the data to enhance it until the desired insight is found. Many applications use this principle to provide meaningful mechanisms to assist decision-makers in achieving their goals. This process can be affected by various uncertainties that can interfere with the user decision-making process. Currently, there are no methodical description and handling tool to include uncertainty in VA systematically. We provide a unified workflow to transform the classic VA cycle into an uncertainty-aware visual analytics (UAVA) cycle consisting of five steps. To prove its usability, three real-world applications represent examples of the UAVA cycle implementation and the described workflow. Robin G. C. Maack, Felix Raith, Juan F. Pérez, Gerik Scheuermann, Christina Gillmann |
Vis. Comput. | 3 |
| 2021 | sPARE: Partial Replication for Multi-Tier Applications in the CloudabstractOffering consistent low latency remains a key challenge for distributed applications, especially when deployed on the cloud where virtual machines (VMs) suffer from capacity variability caused by co-located tenants. Replicating redundant requests was shown to be an effective mechanism to defend application performance from high capacity variability. While the prior art centers on single-tier systems, it still remains an open question how to design replication strategies for distributed multi-tier systems. In this paper, we design a first of its kind PArtial REplication system, sPARE, that replicates and dispatches read-only workloads for distributed multi-tier web applications. The two key components of sPARE are (i) the variability-aware replicator that coordinates the replication levels on all tiers via an iterative searching algorithm, and (ii) the replication-aware arbiter that uses a novel token-based arbitration algorithm (TAD) to dispatch requests in each tier. We evaluate sPARE on web serving and searching applications, i.e., MediaWiki and Solr, the former deployed on our private cloud and the latter on Amazon EC2. Our results based on various interference patterns and traffic loads show that sPARE is able to improve the tail latency of MediaWiki and Solr by a factor of almost 2.7x and 2.9x, respectively. Robert Birke, Juan F. Pérez, Zhan Qiu, Mathias Björkqvist, Lydia Y. Chen |
IEEE Trans. Serv. Comput. | 2 |
| 2019 | Differential Approximation and Sprinting for Multi-Priority Big Data EnginesabstractToday's big data clusters based on the MapReduce paradigm are capable of executing analysis jobs with multiple priorities, providing differential latency guarantees. Traces from production systems show that the latency advantage of high-priority jobs comes at the cost of severe latency degradation of low-priority jobs as well as daunting resource waste caused by repetitive eviction and re-execution of low-priority jobs. We advocate a new resource management design that exploits the idea of differential approximation and sprinting. The unique combination of approximation and sprinting avoids the eviction of low-priority jobs and its consequent latency degradation and resource waste. To this end, we designed, implemented and evaluated DiAS, an extension of the Spark processing engine to support deflate jobs by dropping tasks and to sprint jobs. Our experiments on scenarios with two and three priority classes indicate that DiAS achieves up to 90% and 60% latency reduction for low- and high-priority jobs, respectively. DiAS not only eliminates resource waste but also (surprisingly) lowers energy consumption up to 30% at only a marginal accuracy loss for low-priority jobs. Robert Birke, Isabelly Rocha, Juan F. Pérez, Valerio Schiavoni, Pascal Felber, Lydia Y. Chen |
Middleware | 3 |
| 2017 | Dual Scaling VMs and Queries: Cost-Effective Latency CurtailmentabstractWimpy virtual instances equipped with small numbers of cores and RAM are popular public and private cloud offerings because of their low cost for hosting applications. The challenge is how to run latency-sensitive applications using such instances, which trade off performance for cost. In this study, we analytically and experimentally show that simultaneously scaling resources at coarse granularity and workloads, i.e., submitting multiple query clones to different servers, at fine granularity can overcome the performance disadvantages of wimpy VM instances and achieve stringent latency targets that are even lower than the average execution times of wimpy servers. To such an end, we first derive a closed-form analysis for the latency under any given VM provisioning and query replication level, considering cloning policies that can (not) terminate outstanding clones with (without) an overhead. Validated on trace-driven simulations, our analysis is able to accurately predict the latency and efficiently search for the optimal number of VMs and clones. Secondly, we develop a dual elastic scaler, DuoScale, that dynamically scales VMs and clones according to the workload dynamics so as to achieve the target latency in a cost-effective manner. The effectiveness of DuoScale lies on the observation that the application performance only scales sub-linearly with increasing vertical or horizontal resource provisioning, i.e., resources per VM or number of VMs. We evaluate DuoScale against VM-only scaling strategies via extensive trace-driven simulations as well as experimental results on a cloud test-bed. Our results show that DuoScale is able to achieve the stringent target latency by using clones on wimpy VMs with cost savings up to 50%, compared to scaling brawny VMs that have better performance at a higher unit cost. Juan F. Pérez, Robert Birke, Mathias Björkqvist, Lydia Y. Chen |
ICDCS | 1 |
| 2017 | Power of redundancy: Designing partial replication for multi-tier applicationsabstractReplicating redundant requests has been shown to be an effective mechanism to defend application performance from high capacity variability - the common pitfall in the cloud. While the prior art centers on single-tier systems, it still remains an open question how to design replication strategies for distributed multi-tier systems, where interference from neighboring workloads is entangled with complex tier interdependency. In this paper, we design a first of its kind PArtial REplication system, sPARE, that replicates and dispatches read-only workloads for multi-tier web applications, determining replication factors per tier. The two key components of sPARE are (i) the variability-aware replicator that coordinates the replication levels on all tiers via an iterative searching algorithm, and (ii) the replication-aware arbiter that uses a novel token-based arbitration algorithm (TAD) to dispatch requests in each tier. We evaluate sPARE on web serving and web searching applications, i.e., MediaWiki and Solr, deployed on our private cloud testbed. Our results based on various interference patterns and traffic loads show that sPARE is able to improve the tail latency of MediaWiki and Solr by a factor of almost 2.7x and 2.9x, respectively. Robert Birke, Juan F. Pérez, Zhan Qiu, Mathias Björkqvist, Lydia Y. Chen |
INFOCOM | 2 |
| 2017 | On the latency-accuracy tradeoff in approximate MapReduce jobsabstractTo ensure the scalability of big data analytics, approximate MapReduce platforms emerge to explicitly trade off accuracy for latency. A key step to determine optimal approximation levels is to capture the latency of big data jobs, which is long deemed challenging due to the complex dependency among data inputs and map/reduce tasks. In this paper, we use matrix analytic methods to derive stochastic models that can predict a wide spectrum of latency metrics, e.g., average, tails, and distributions, for approximate MapReduce jobs that are subject to strategies of input sampling and task dropping. In addition to capturing the dependency among waves of map/reduce tasks, our models incorporate two job scheduling policies, namely, exclusive and overlapping, and two task dropping strategies, namely, early and straggler, enabling us to realistically evaluate the potential performance gains of approximate computing. Our numerical analysis shows that the proposed models can guide big data platforms to determine the optimal approximation strategies and degrees of approximation. Juan F. Pérez, Robert Birke, Lydia Y. Chen |
INFOCOM | 1 |
| 2017 | Algorithm 972: jMarkov: An Integrated Framework for Markov Chain ModelingabstractMarkov chains (MC) are a powerful tool for modeling complex stochastic systems. Whereas a number of tools exist for solving different types of MC models, the first step in MC modeling is to define the model parameters. This step is, however, error prone and far from trivial when modeling complex systems. In this article, we introduce jMarkov, a framework for MC modeling that provides the user with the ability to define MC models from the basic rules underlying the system dynamics. From these rules, jMarkov automatically obtains the MC parameters and solves the model to determine steady-state and transient performance measures. The jMarkov framework is composed of four modules: (i) the main module supports MC models with a finite state space; (ii) the jQBD module enables the modeling of Quasi-Birth-and-Death processes, a class of MCs with infinite state space; (iii) the jMDP module offers the capabilities to determine optimal decision rules based on Markov Decision Processes; and (iv) the jPhase module supports the manipulation and inclusion of phase-type variables to represent more general behaviors than that of the standard exponential distribution. In addition, jMarkov is highly extensible, allowing the users to introduce new modeling abstractions and solvers. Juan F. Pérez, Daniel F. Silva, Julio Cesar Goez, Andrés Sarmiento, Raha Akhavan-Tabatabaei, Germán Riaño |
ACM Trans. Math. Softw. | 1 |
| 2017 | Cutting Latency Tail: Analyzing and Validating Replication without CancelingabstractResponse time variability in software applications can severely degrade the quality of the user experience. To reduce this variability, request replication emerges as an effective solution by spawning multiple copies of each request and using the result of the first one to complete. Most previous studies have mainly focused on the mean latency for systems implementing replica cancellation, i.e., all replicas of a request are canceled once the first one finishes. Instead, we develop models to obtain the response-time distribution for systems where replica cancellation may be too expensive or infeasible to implement, as in “fast” systems, such as web services, or in legacy systems. Furthermore, we introduce a novel service model to explicitly consider correlation in the processing times of the request replicas, and design an efficient algorithm to parameterize the model from real data. Extensive evaluations on a MATLAB benchmark and a three-tier web application (MediaWiki) show remarkable accuracy, e.g., 7 (4 percent) average error on the 99th percentile response time for the benchmark (respectively, MediaWiki), the requests of which execute in the order of seconds (respectively, milliseconds). Insights into optimal replication levels are thereby gained from this precise quantitative analysis, under a wide variety of system scenarios. Zhan Qiu, Juan F. Pérez, Robert Birke, Lydia Y. Chen, Peter G. Harrison |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Line: Evaluating Software Applications in Unreliable EnvironmentsabstractCloud computing has paved the way to the flexible deployment of software applications. This flexibility offers service providers a number of options to tailor their deployments to the observed and foreseen customer workloads, without incurring in large capital costs. However, cloud deployments pose novel challenges regarding application reliability and performance. Examples include managing the reliability of deployments that make use of spot instances, or coping with the performance variability caused by multiple tenants in a virtualized environment. In this paper, we introduce Line, a tool for performance and reliability analysis of software applications. Line solves layered queueing network (LQN) models, a popular class of stochastic models in software performance engineering, by setting up and solving an associated system of ordinary differential equations. A key differentiator of Line compared to existing solvers for LQNs is that Line incorporates a model of the environment the application operates in. This enables the modeling of reliability and performance issues such as resource failures, server breakdowns and repairs, slow start-up times, resource interference due to multitenancy, among others. This paper describes the Line tool, its support for performance and reliability modeling, and illustrates its potential by comparing Line predictions against data obtained from a cloud deployment. We also illustrate the applicability of Line with a case study on reliability-aware resource provisioning. Juan F. Pérez, Giuliano Casale |
IEEE Trans. Reliab. | 1 |
| 2016 | Variability-aware request replication for latency curtailmentabstractProcessing time variability is commonplace in distributed systems, where resources display disparate performance due to, e.g., different workload levels, background processes, and contention in virtualized environments. However, it is paramount for service providers to keep variability in response time under control in order to offer responsive services. We investigate how request replication can be used to exploit processing time variability to reduce response times, considering not only mean values but also the tail of the response time distribution. We focus on the distributed setup, where replication is achieved by running copies of requests on multiple servers that otherwise evolve independently, and waiting for the first replica to complete service. We construct models that capture the evolution of a system with replicated requests using approximate methods and observe that highly variable service times offer the best opportunities for replication - reducing the response time tail in particular. Further, the effect of replication is non-uniform over the response time distribution: gains in one metric, e.g., the mean, can be at the cost of another, e.g., the tail percentiles. This is demonstrated in wide range of numerical virtual experiments. It can be seen that capturing service time variability is key to the evaluation of latency tolerance strategies and in their design. Zhan Qiu, Juan F. Pérez, Peter G. Harrison |
INFOCOM | 2 |
| 2016 | Quantifying the Impact of Replication on the Quality-of-Service in Cloud DatabasesabstractCloud databases achieve high availability by automatically replicating data on multiple nodes. However, the overhead caused by the replication process can lead to an increase in the mean and variance of transaction response times, causing unforeseen impacts on the offered quality-of-service (QoS). In this paper, we propose a measurement-driven methodology to predict the impact of replication on Database-as-a-Service (DBaaS) environments. Our methodology uses operational data to parameterize a closed queueing network model of the database cluster together with a Markov model that abstracts the dynamic replication process. Experiments on Amazon RDS show that our methodology predicts response time mean and percentiles with errors of just 1% and 15% respectively, and under operational conditions that are significantly different from the ones used for model parameterization. We show that our modeling approach surpasses standard modeling methods and illustrate the applicability of our methodology for automated DBaaS provisioning. Rasha Osman, Juan F. Pérez, Giuliano Casale |
QRS | 2 |
| 2016 | Tackling Latency via Replication in Distributed SystemsabstractConsistently high reliability and low latency are twin requirements common to many forms of distributed processing; for example, server farms and mirrored storage access. To address them, we consider replication of requests with canceling - i.e. initiate multiple concurrent replicas of a request and use the first successful result returned, canceling all outstanding replicas. This scheme has been studied recently, but mostly for systems with a single central queue, while server farms exploit distributed resources for scalability and robustness. We develop an approximate stochastic model to determine the response-time distribution in a system with distributed queues, and compare its performance against its centralized counterpart. Validation against simulation indicates that our model is accurate for not only the mean response time but also its percentiles, which are particularly relevant for deadline-driven applications. Further, we show that in the distributed set-up, replication with canceling has the potential to reduce response times, even at relatively high utilization. We also find that it offers response times close to those of the centralized system, especially at medium-to-high request reliability. These findings support the use of replication with canceling as an effective mechanism for both fault- and delay-tolerance. Zhan Qiu, Juan F. Pérez, Peter G. Harrison |
ICPE | 2 |
| 2016 | Evaluating Replication for Parallel Jobs: An Efficient ApproachabstractMany modern software applications rely on parallel job processing to exploit large resource pools available in cloud and grid infrastructures. The response time of a parallel job, made of many subtasks, is determined by the last subtask that finishes. Thus, a single laggard subtask or a failure, requiring re-processing, may increase the response time substantially. To overcome these issues, we explore concurrent replication with canceling. This mechanism executes two job replicas concurrently, and retrieves the result of the first replica that completes, immediately canceling the other one. To analyze this mechanism we propose a stochastic model that considers replication at both job-level and task-level. We find that task-level replication achieves a much higher reliability and shorter response times than job-level replication. We also observe that the impact of replication depends on the system utilization, the subtask reliability, and the correlation among replica failures. Based on the model, we propose a resource-provisioning strategy that determines the minimum number of computing nodes needed to achieve a service-level objective (SLO) defined as a response-time percentile. This strategy is evaluated by considering realistic traffic patterns from a parallel cluster, where task-level replication shows the potential to reduce the resource requirements for tight response-time SLOs. Zhan Qiu, Juan F. Pérez |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Evaluating the Effectiveness of Replication for Tail-ToleranceabstractComputing clusters (CC) are a cost-effective high-performance platform for computation-intensive scientific and engineering applications. A key challenge in managing CCs is to consistently achieve low response times. In particular, tail-tolerant methods aim to keep the tail of the response-time distribution short. In this paper we explore concurrent replication with cancelling, a tail-tolerant approach that involves processing requests and their replicas concurrently, retrieving the result from the first replica that completes, and cancelling all other replicas. We propose a stochastic model that considers any number of replicas, general processing and inter-arrival times, and computes the response time distribution. We show that replication can be very effective in keeping the response-time tail short, but these benefits highly depend on the processing-time distribution, as well as on the CC utilization and the statistical characteristics of the arrival process. We also exploit the model to support the selection of the optimal number of replicas, and a resource provisioning strategy that meets service-level objectives on the response-time percentiles. Zhan Qiu, Juan F. Pérez |
CCGRID | 2 |
| 2015 | DICE: Quality-Driven Development of Data-Intensive Cloud ApplicationsabstractModel-driven engineering (MDE) often features quality assurance (QA) techniques to help developers creating software that meets reliability, efficiency, and safety requirements. In this paper, we consider the question of how quality-aware MDE should support data-intensive software systems. This is a difficult challenge, since existing models and QA techniques largely ignore properties of data such as volumes, velocities, or data location. Furthermore, QA requires the ability to characterize the behavior of technologies such as Hadoop/MapReduce, NoSQL, and stream-based processing, which are poorly understood from a modeling standpoint. To foster a community response to these challenges, we present the research agenda of DICE, a quality-aware MDE methodology for data-intensive cloud applications. DICE aims at developing a quality engineering tool chain offering simulation, verification, and architectural optimization for Big Data applications. We overview some key challenges involved in developing these tools and the underpinning models. Giuliano Casale, Danilo Ardagna, Matej Artac, Franck Barbier, Elisabetta Di Nitto, Alexis Henry, Gabriel Iuhasz, Christophe Joubert, José Merseguer, Victor Ion Munteanu, Juan F. Pérez, Dana Petcu, Matteo G. Rossi, Craig Sheridan, Ilias Spais, Daniel Vladuic |
MiSE@ICSE | 11 |
| 2015 | Enhancing reliability and response times via replication in computing clustersabstractComputing clusters have been widely deployed for scientific and engineering applications to support intensive computation and massive data operations. As applications and resources in a cluster are subject to failures, fault-tolerance strategies are commonly adopted, sometimes at the expense of additional delays in job response times, or unnecessarily increasing resource usage. In this paper, we explore concurrent replication with canceling, a fault-tolerance approach where jobs and their replicas are processed concurrently, and the successful completion of either triggers the removals of its replica. We propose a stochastic model to study how this approach affects the cluster service level objectives (SLOs), particularly the offered response time percentiles. In addition to the expected gains in reliability, the proposed model allows us to determine the regions of the utilization where introducing replication with canceling effectively reduces the response times. Moreover, we show how this model can support resource provisioning decisions with reliability and response time guarantees. Zhan Qiu, Juan F. Pérez |
INFOCOM | 2 |
| 2015 | QD-AMVA: Evaluating systems with queue-dependent service requirementsabstractWorkload measurements in enterprise systems often lead to observe a dependence between the number of requests running at a resource and their mean service requirements. However, multiclass performance models that feature these dependences are challenging to analyze, a fact that discourages practitioners from characterizing workload dependences. We here focus on closed multiclass queueing networks and introduce QD-AMVA, the first approximate mean-value analysis (AMVA) algorithm that can efficiently and robustly analyze queue-dependent service times in a multiclass setting. A key feature of QD-AMVA is that it operates on mean values, avoiding the computation of state probabilities. This property is an innovative result for state-dependent models, which increases the computational efficiency and numerical robustness of their evaluation. Extensive validation on random examples, a cloud load-balancing case study and comparison with a fluid method and an existing AMVA approximation prove that QD-AMVA is efficient, robust and easy to apply, thus enhancing the tractability of queue-dependent models. Giuliano Casale, Juan F. Pérez, Weikun Wang |
Perform. Evaluation | 2 |
| 2015 | Beyond the mean in fork-join queues: Efficient approximation for response-time tailsabstractFork-join queues are natural models for various computer and communications systems that involve parallel multitasking and the splitting and resynchronizing of data, such as parallel computing, query processing in distributed databases, and parallel disk access. Job response time in a fork-join queue is a critical performance indicator but its exact analysis is challenging. We introduce a stochastic model for K -node homogeneous fork-join queues ( K ≥ 2 ) that focuses on the difference in length between any node-queue and the shortest one, truncating the state space such that the maximum difference is at most a constant C . Whilst most previous methods focus on the mean response time, our model is also able to evaluate the response time distribution , as well as accommodating phase-type processing times and Markovian arrival processes. In order to tackle scenarios with high loads, which require a large value of C to provide sufficient accuracy, we develop an efficient algorithm using matrix-analytic methods. Tests against simulation show that the proposed model yields accurate results for 2-node fork-join queues. As the model becomes numerically intractable for large values of K , we further propose an approximate approach, based on properties of order statistics and extreme values. The approximation gives a high degree of accuracy on response time tails, and has the advantage of being efficient and scalable, requiring only the analytical results for a single-node and 2-node fork-join queues, which we obtain with the aforementioned matrix-analytic model. Comparison with simulation results shows that our approximation yields good fits for the tails, even in very large cases with general processing and inter-arrival times. Zhan Qiu, Juan F. Pérez, Peter G. Harrison |
Perform. Evaluation | 2 |
| 2015 | Estimating Computational Requirements in Multi-Threaded ApplicationsabstractPerformance models provide effective support for managing quality-of-service (QoS) and costs of enterprise applications. However, expensive high-resolution monitoring would be needed to obtain key model parameters, such as the CPU consumption of individual requests, which are thus more commonly estimated from other measures. However, current estimators are often inaccurate in accounting for scheduling in multi-threaded application servers. To cope with this problem, we propose novel linear regression and maximum likelihood estimators. Our algorithms take as inputs response time and resource queue measurements and return estimates of CPU consumption for individual request types. Results on simulated and real application datasets indicate that our algorithms provide accurate estimates and can scale effectively with the threading levels. Juan F. Pérez, Giuliano Casale, Sergio Pacheco-Sanchez |
IEEE Trans. Software Eng. | 1 |
| 2014 | Assessing the Impact of Concurrent Replication with Canceling in Parallel JobsabstractParallel job processing has become a key feature of many software applications, e.g., in scientific computing. Parallelization allows these applications to exploit large resource pools, such as cloud or grid data centers. However, a job composed of a large number of parallel tasks will suffer a failure if any of its tasks fail, requiring reprocessing and additional delays. In this paper, we explore the effect that the replication of parallel jobs has on the job reliability and response time, as well as on resource utilization. The replication mechanism consists of concurrently processing replicas, at either the job or the task level, retrieving the results of the replica that finishes first, if any, and canceling any remaining replica in process. We propose a stochastic model that explicitly considers parallel job processing, replication at both the job and the task level, and handles general arrival processes. We develop a numerically-efficient algorithm to solve large-scale instances of the model and compute key performance metrics. We observe that the task cancellation mechanism offers an effective way of limiting the increase in resource utilization, allowing the use of replicas that not only increase the job reliability, but have the potential to reduce the response times. Zhan Qiu, Juan F. Pérez |
MASCOTS | 2 |
| 2013 | An Offline Demand Estimation Method for Multi-threaded ApplicationsabstractParameterizing performance models for multi-threaded enterprise applications requires finding the service rates offered by worker threads to the incoming requests. Statistical inference on monitoring data is here helpful to reduce the overheads of application profiling and to infer missing information. While linear regression of utilization data is often used to estimate service rates, it suffers erratic performance and also ignores a large part of application monitoring data, e.g., response times. Yet inference from other metrics, such as response times or queue-length samples, is complicated by the dependence on scheduling policies. To address these issues, we propose novel scheduling-aware estimation approaches for multi-threaded applications based on linear regression and maximum likelihood estimators. The proposed methods estimate demands from samples of the number of requests in execution in the worker threads at the admission instant of a new request. Validation results are presented on simulated and real application datasets for systems with multi-class requests, class switching, and admission control. Juan F. Pérez, Sergio Pacheco-Sanchez, Giuliano Casale |
MASCOTS | 1 |
| 2011 | Quasi-birth-and-death processes with restricted transitions and its applications
Juan F. Pérez, Benny Van Houdt |
Perform. Evaluation | 1 |
| 2010 | A mean field model for an optical switch with a large number of wavelengths and centralized partial conversion
Juan F. Pérez, Benny Van Houdt |
Perform. Evaluation | 1 |
| 2009 | Dimensioning an OBS Switch with Partial Wavelength Conversion and Fiber Delay Lines via a Mean Field ModelabstractIn this paper we introduce a mean field model to analyze an optical switch equipped with both wavelength converters (WCs) and fiber delay lines (FDLs) to resolve contention in OBS networks. Under some very general conditions, that is, a general burst size distribution and any Markovian burst arrival process at each wavelength, this model determines the minimum number of WCs required to achieve a zero loss rate as the number of wavelengths becomes large. The mean field result is exact as the number of wavelengths goes to infinity and turns out to be very accurate for systems with (a few) hundred wavelengths, commonly occurring when using wavelength division multiplexing (WDM). Moreover, we show that if the number of WCs is underdimensioned, (i) periodic system behavior may occur (with the period being the greatest common divisor of the burst lengths) and (ii) increasing the number of WCs may even worsen the loss rate under the often studied minimum horizon allocation policy (as opposed to the minimum gap policy). Finally, we further demonstrate that in terms of the loss rate, including (more) FDLs may have little or no effect on the number of WCs required to achieve a near-zero loss, especially for higher loads. Juan F. Pérez, Benny Van Houdt |
INFOCOM | 1 |