VLDB 2026 Research / reviewers in the wild / expert
Thiago Emmanuel Pereira
dblp:140/9447
· DBLP profile ↗
11ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-2702-6019ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 4 since 2021Systems, architecture and hardware · 4 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KLUE: A Framework for Cost-Effective Experimentation in Emulated Kubernetes ClustersabstractExecuting performance experiments in Kubernetes clusters is a resource- and time-intensive task, especially in large-scale environments typical of production systems. However, such experiments are essential for understanding system behavior and supporting operational decisions that improve efficiency and reliability. This paper introduces KLUE (Kubernetes Lite execUtion Environment), a lightweight framework that enables performance experimentation in emulated Kubernetes clusters. KLUE provides a practical, cost-effective approach for testing and validating configurations, policies, and workloads without the need for extensive physical infrastructure. Using KLUE, we successfully reproduced an experiment originally conducted in a real Kubernetes cluster, achieving a 93.5% reduction in execution cost. We also leveraged the framework to study the impact of multiple application spreading strategies using a 24-hour production-scale trace from a large technology company by spending only 0.14% of the estimated cost required to run the same study on real infrastructure—an analysis that would be economically infeasible in real environments. These results highlight KLUE's potential to accelerate experimentation, reduce costs, and improve decision-making in Kubernetes-based environments, offering a valuable tool for both research and industry settings. Kayky Fidelis, Geraldo Junior, Caetano Albuquerque, Giovanni Farias da Silva, Thiago Emmanuel Pereira, Fábio Morais 0001, Kilian Melcher |
ICPE | 5 |
| 2026 | On the Efficiency and Disruption Trade-Offs of Kubernetes Packing Heuristics
Mariane Santos Zeitouni, Oscar Brito, Matheus Rocha, Thiago Emmanuel Pereira, Gabriel Gomes |
ICPE | 4 |
| 2025 | A Defect Taxonomy for Infrastructure as Code: A Replication StudyabstractBackground: As Infrastructure as Code (IaC) becomes standard practice, ensuring the reliability of IaC scripts is essential. Defect taxonomies are valuable tools for this, offering a common language for issues and enabling systematic tracking. A significant prior study developed such a taxonomy, but based it exclusively on the declarative language Puppet. It remained unknown whether this taxonomy applies to programming language-based IaC (PLIaC) tools like Pulumi, Terraform CDK, and AWS CDK. Aim: We replicated this foundational work to assess the generalizability of the taxonomy across a broader and more diverse landscape. Method: We performed qualitative analysis on 3,364 defect-related commits from 285 open-source PL-IaC repositories (PIPr dataset) to derive a PL-IaC specific defect taxonomy. We then enhanced the ACID tool, originally developed for the prior study, to automatically classify and analyze defect distributions across an expanded dataset- 447 open-source repositories and 94 proprietary projects from VTEX (e-commerce) and Nubank (financial)-incorporating modern PL-IaC tools absent in the original work. Results: Our research confirmed the same eight defect categories identified in the original study, with idempotency and security defects appearing infrequently but persistently across projects. Configuration Data defects maintain high frequency in both open-source and proprietary codebases. Despite differences in project types, overall defect proportions remain similar. Conclusions: Our replication supports the generalizability of the original taxonomy, suggesting IaC development challenges surpass organizational boundaries. Configuration Data defects emerge as a persistent highfrequency problem, while idempotency and security defects remain important concerns despite lower frequency. These patterns appear consistent across open-source and proprietary projects, indicating they are fundamental to the IaC paradigm itself, transcending specific tools or project types. Wendell Oliveira, Filipe Paiva, Thiago Emmanuel Pereira, João Brunet |
ESEM | 3 |
| 2024 | No Clash on Cache: Observations from a Multi-tenant Ecommerce PlatformabstractCaching is a classic technique for improving system performance by reducing client-perceived latency and server load. However, cache management still needs to be improved and is even more difficult in multi-tenant systems. To shed light on these problems and discuss possible solutions, we performed a workload characterization of a multi-tenant cache operated by a large ecommerce platform. In this platform, each one of thousands of tenants operates independently. We found that the workload patterns of the tenants could be very different. Also, the characteristics of the tenants change over time. Based on these findings, we highlight strategies to improve the management of multi-tenant cache systems. Anna Lira, Ruan Alves, Thiago Emmanuel Pereira, Fábio Morais 0001, João Ramalho, Mariana Mendes |
ICPE | 3 |
| 2024 | Prebaking runtime environments to improve the FaaS cold start latency
Daniel Fireman, Thiago Emmanuel Pereira, Luis Mafra, Dalton C. G. Valadares |
Future Gener. Comput. Syst. | 3 |
| 2020 | Prebaking Functions to Warm the Serverless Cold StartabstractFunction-as-service (FaaS) platforms promise a simpler programming model for cloud computing, in which the developers concentrate on writing its applications. In contrast, platform providers take care of resource management and administration. As FaaS users are billed based on the execution of the functions, platform providers have a natural incentive not to keep idle resources running at the platform's expense. However, this strategy may lead to the cold start issue, in which the execution of a function is delayed because there is no ready resource to host the execution. Cold starts can take hundreds of milliseconds to seconds and have been a prohibitive and painful disadvantage for some applications. This work describes and evaluates a technique to start functions, which restores snapshots from previously executed function processes. We developed a prototype of this technique based on the CRIU process checkpoint/restore Linux tool. We evaluate this prototype by running experiments that compare its start-up time against the standard Unix process creation/start-up procedure. We analyze the following three functions: i) a "do-nothing" function, ii) an Image Resizer function, and iii) a function that renders Markdown files. The results attained indicate that the technique can improve the start-up time of function replicas by 40% (in the worst case of a "do-nothing" function) and up to 71% for the Image Resizer one. Further analysis indicates that the runtime initialization is a key factor, and we confirmed it by performing a sensitivity analysis based on synthetically generated functions of different code sizes. These experiments demonstrate that it is critical to decide when to create a snapshot of a function. When one creates the snapshots of warm functions, the speed-up achieved by the prebaking technique is even higher: the speed-up increases from 127.45% to 403.96%, for a small, synthetic function; and for a bigger, synthetic function, this ratio increases from 121.07% to 1932.49%. Daniel Fireman, Thiago Emmanuel Pereira |
Middleware | 3 |
| 2020 | Controlling Garbage Collection and Request Admission to Improve Performance of FaaS ApplicationsabstractRuntime environments like Java's JRE, .NET's CLR, and Ruby's MRI, are popular choices for cloud-based applications and particularly in the Function as a Service (FaaS) serverless computing context. A critical component of runtime environments of these languages is the garbage collector (GC). The GC frees developers from manual memory management, which could potentially ease development and avoid bugs. The benefits of using the GC come with a negative impact on performance; that impact happens because either the GC needs to pause the runtime execution or competes with the running program for computational resources. In this work, we evaluated the usage of a technique - Garbage Collector Control Interceptor (GCI) - that eliminates the negative impact of GC on performance by controlling GC executions and transparently shedding requests while the collections are happening. We executed experiments simulating AWS Lambda's behavior and found that GCI is a viable solution. It benefited the user by improving the response time up to 10.86% at 99.9th percentile and reducing cost by 7.22%, but it also helped the platform provider by improving resource utilization by 14.52%. David Quaresma, Daniel Fireman, Thiago Emmanuel Pereira |
SBAC-PAD | 3 |
| 2018 | Improving Tail Latency of Stateful Cloud Services via GC Control and Load SheddingabstractMost of the modern cloud web services execute on top of runtime environments like .NET's Common Language Runtime or Java Runtime Environment. On the one hand, runtime environments provide several off-the-shelf benefits like code security and cross-platform execution. On the other hand, runtime's features such as just-in-time compilation and automatic memory management add a non-deterministic overhead to the overall service time, increasing the tail of the latency distribution. In this context, the Garbage Collector (GC) is among the leading causes of high tail latency. To tackle this problem, we developed the Garbage Collector Control Interceptor (GCI) - a request interceptor algorithm, which is agnostic regarding the cloud service language, internals, and its incoming load. GCI is wholly decentralized and improves the tail latency of cloud services by making sure that service instances shed the incoming load while cleaning up the runtime heap. We evaluated GCI's effectiveness in a stateful service prototype, varying the number of available instances. Our results showed that using GCI eliminates the impact of the garbage collection on the service latency for small (4 nodes) and large (64 nodes) deployments with no throughput loss. Daniel Fireman, João Brunet, Raquel Lopes 0001, David Quaresma, Thiago Emmanuel Pereira |
CloudCom | 5 |
| 2016 | File system trace replay methods through the lens of metrologyabstractThere are various methods to evaluate the performance of file systems through the replay of file system traces. Despite this diversity, little attention was given on comparing the alternatives, thus bringing some skepticism about the results attained using these methods. In this paper, to fill this understanding gap, we analyze two popular trace replay methods through the lens of metrology. This case study indicates that the evaluated methods provide similar, good precision but are biased in some scenarios. Our results identified limitations in the implementation of the replay tool as well as flaws in the established practices to experiment with trace replayers as the root causes of the measurement bias. After improving the implementation of the trace replayer and discarding inappropriate experimental practices, we were able to reduce the bias, leading to lower measurement uncertainty. Finally, our case study also shows that, in some cases, collecting only the file system activity is not enough to accurately replay the traces; in these cases, collecting resource consumption information, such as the amount of allocated memory, can improve the quality of trace replay methods. Thiago Emmanuel Pereira, Francisco Vilar Brasileiro, Lívia M. R. Sampaio |
MSST | 1 |
| 2016 | A study on the errors and uncertainties of file system trace capture methodsabstractDespite the popularity of trace-based file system performance evaluation, there is currently no accepted methodology to capture and accurately replay file system traces. The less we know about the limitations of trace-based methodologies, the less we can rely on the obtained results. In this paper, we present a case study analyzing the two most popular trace capture methods. The results of the case study indicate that the two evaluated methods provide good precision, but significant bias in some cases. In addition to providing guidelines on how to improve the quality of trace capture, our results allow us to draw important observations about current practice. We show that bias can be corrected by the execution of a calibration procedure, a practice that is mostly absent in the methodology used in the area. Our results also revealed that, to achieve correct calibration, it is crucial to collect information about the background activity and about the operating system layer running in the experimental environment. Finally, our results allow to quantify the overall trace capture uncertainty, providing means for researches to decide on the suitability of trace capture tools for their purposes. Thiago Emmanuel Pereira, Francisco Vilar Brasileiro, Lívia M. R. Sampaio |
SYSTOR | 1 |
| 2013 | On the Accuracy of Trace Replay Methods for File System EvaluationabstractCurrent trace replay methods for file system evaluation fail to represent traced workloads accurately. When using a misrepresented workload one may take wrong conclusions about the system evaluation. For example, a system designer can miss performance problems if the replay of a trace produces an under loaded representation of the real workload. Even worse, one can take wrong design decisions, leading to optimization of untypical workloads. In this study, we captured and replayed traces from standard file systems using methods proposed in the literature, to exemplify the inaccuracy of state-of-art trace replay methods. We also exposed a shortcoming of current methodologies, in a replay of a general purpose workload trace, we observed a difference of up to 100% on request response time, caused by the choice of trace replay method. Thiago Emmanuel Pereira, Lívia M. R. Sampaio, Francisco Vilar Brasileiro |
MASCOTS | 1 |