Thais Aparecida Silva Camacho

dblp:268/5704 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0001-9605-1037ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Using Task Graph Caching to Accelerate TVM Code Generation
abstract
Deep Learning (DL) models are at the core of a growing number of applications, making fast, low-latency execution across diverse device architectures both a critical requirement and a challenge. DL compilers, such as TVM, address this challenge by automatically translating high-level models into optimized low-level code that effectively exploits device architectures. However, search-space-based algorithms face difficulties in exploring the vast optimization sequence space, often resulting in lengthy compilation times that significantly impact the design cycle. This article introduces the Task Graph Caching (TGC) algorithm, 1 which aims to reduce the high compilation time while preserving the quality of the code generated by TVM. In particular, TGC enhances TVM auto-tuning by exploiting the fact that similar DL subgraphs appear both within and across models, thus enabling optimization sequences discovered in past compilations to guide future executions. To achieve this, TGC introduces a cache structure that stores high-performance optimization sequences found in previous TVM executions. This information is then used to seed the population of TVM evolutionary search, avoiding redundant exploration of the optimization space and accelerating convergence. Experimental results on twelve DL models show that TGC can significantly speed up the search for efficient optimization sequences, reducing auto-tuning time by up to 2.89× for Ansor and 3.13× for MetaSchedule on CPU. Moreover, TGC reduces auto-tuning time by up to 3.25× for MetaSchedule on GPU. Furthermore, on average, TGC can maintain the inference time achieved by the default TVM, making it a promising solution for accelerating the compilation of DL models.
Thais Aparecida Silva Camacho, Lucas Fernando Alvarenga e Silva, Márcio Machado Pereira, Guido Araujo
ACM Trans. Archit. Code Optim.1
2023 PB3Opt: Profile-based biased Bayesian optimization to select computing clusters on the cloud
abstract
Summary Given the wide variety of cloud computing resources for creating high‐performance computer clusters and their complex performance relationship with applications, finding the optimal, or near‐optimal, cluster is a complex problem. As a result, several approaches have been proposed to find the optimal, or near‐optimal, cluster for a given high‐performance computing workload, while reducing the search cost. Among the approaches found in the literature, Bayesian optimization is one of the most known and applied. However, it is still possible to increase its performance by integrating it with historical data related to workload behavior. In this context, we suggest the approach, which introduces a bias in the Bayesian optimization expected improvement acquisition function. The new acquisition function uses the ranking of computer clusters of previously explored workloads that have the same behavior as the workload being optimized. Our experimental results show that classifies the behavior of workloads in groups so that the average‐ranking has 88.7% similarity with the ranking of the workload. With this, finds, for almost 95% of workloads, a solution that is less than or equal to 1.2 worse than the optimal computer cluster. In addition, the works well when combined with the paramount iterations technique and is capable of reducing the search cost significantly.
Thais Aparecida Silva Camacho, Vanderson Martins do Rosário, Otávio O. Napoli, Edson Borin
Concurr. Comput. Pract. Exp.1
2021 Smart selection of optimizations in dynamic compilers
abstract
Summary Dynamic compilers perform compilation and generation of target code during runtime, implying that the compilation time is added into the program runtime. Thus, to build a high‐performing dynamic compilation system, it is crucial to be able to generate high‐quality code and, at the same time, have a small compilation cost. In this article, we present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Our approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high‐quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. We evaluate the proposed strategy in two LLVM‐based dynamic binary translators, namely OI‐DBT and HQEMU, and show that these two translators can achieve average speedups of 1.26x and 1.15x in MiBench and Spec Cpu benchmarks, respectively.
Vanderson Martins do Rosário, Anderson Faustino da Silva, Thais Aparecida Silva Camacho, Otávio O. Napoli, Maurício Breternitz, Edson Borin
Concurr. Comput. Pract. Exp.3