Fabrício Góes

dblp:49/3151 · also Luís F. W. Góes, Luís Fabrício Wanderley Góes · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0003-1801-9917ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 4Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique
abstract
This study adapts the Consensual Assessment Technique (CAT) for Large Language Models (LLMs), introducing a novel methodology for poetry evaluation.Using a 90-poem dataset with a ground truth based on publication venue, we demonstrate that this approach allows LLMs to significantly surpass the performance of non-expert human judges.Our method, which leverages forced-choice ranking within small, randomized batches, enabled Claude-3-Opus to achieve a Spearman's Rank Correlation of 0.87 with the ground truth, dramatically outperforming the best human nonexpert evaluation (SRC = 0.38).The LLM assessments also exhibited high inter-rater reliability, underscoring the methodology's robustness.These findings establish that LLMs, when guided by a comparative framework, can be effective and reliable tools for assessing poetry, paving the way for their broader application in other creative domains.
Piotr Sawicki 0001, Marek Grzes, Dan Brown 0001, Fabrício Góes
EMNLP4
2025 Do LLMs Agree on the Creativity Evaluation of Alternative Uses?
Abdullah Al Rabeyah, Fabrício Góes, Marco Volpe 0001, Talles H. Medeiros
ICCC2
2024 CreativeStone: A Creativity Booster for Hearthstone Card Decks
abstract
Digital Collectible Card Games (DCCG) rely on human expertise to create competitive card decks. The most effective decks are shared and become popular among players through online communities. Players then keep enhancing those decks by replacing cards and testing those new deck variations in online arenas. This fine-tuning process is time-consuming and most of the time leads to small improvements in the players win rate. This paper presents CreativeStone, a creativity booster for Hearthstone card decks. Our creativity-guided approach uses a genetic algorithm and the Regent-Dependent Creativity (RDC) metric to identify the core cards of an existing deck, and then improves it towards a more valuable and novel deck by adding cards that are synergic to the core cards, but also different from the ones in the original deck. Our experimental results show that CreativeStone can boost even legendary decks, outperforming handcrafted ones by 21% in win rate.
Celso França, Zisen Zhou, Carolina Fernanda da Silva, Cibele Simões de Oliveira Santos, Lucas Braga Ferreira, Lucas Henrique Pereira, Marcos Pablo Souza de Almeida, Marina Iolanda Oliveira, Thaís Damásio, Vinícius Pacheco, Fabrício Góes
IEEE Trans. Games11
2023 Is GPT-4 Good Enough to Evaluate Jokes?
Fabrício Góes, Piotr Sawicki 0001, Marek Grzes, Marco Volpe 0001, Dan Brown 0001
ICCC1
2023 Pushing GPT's Creativity to Its Limits: Alternative Uses and Torrance Tests
Fabrício Góes, Piotr Sawicki 0001, Marek Grzes, Marco Volpe 0001, Jacob Watson
ICCC1
2023 Bits of Grass: Does GPT already know how to write like Whitman?
Piotr Sawicki 0001, Marek Grzes, Fabrício Góes, Dan Brown 0001, Max Peeperkorn, Aisha Khatun
ICCC3
2023 On the power of special-purpose GPT models to create and evaluate new poetry in old styles
Piotr Sawicki 0001, Marek Grzes, Fabrício Góes, Anna Jordanous, Dan Brown 0001, Simona Paraskevopoulou, Max Peeperkorn, Aisha Khatun
ICCC3
2020 Vectorization-aware loop unrolling with seed forwarding
abstract
Loop unrolling is a widely adopted loop transformation, commonly used for enabling subsequent optimizations. Straight-line-code vectorization (SLP) is an optimization that benefits from unrolling. SLP converts isomorphic instruction sequences into vector code. Since unrolling generates repeatead isomorphic instruction sequences, it enables SLP to vectorize more code. However, most production compilers apply these optimizations independently and uncoordinated. Unrolling is commonly tuned to avoid code bloat, not maximizing the potential for vectorization, leading to missed vectorization opportunities.
Rodrigo Caetano Rocha, Vasileios Porpodas, Pavlos Petoumenos, Fabrício Góes, Zheng Wang 0001, Murray Cole, Hugh Leather
CC4
2019 Super-Node SLP: Optimized Vectorization for Code Sequences Containing Operators and Their Inverse Elements
abstract
SLP Auto-vectorization converts straight-line code into vector code. It scans input code for groups of instructions that can be combined into vectors and replaces them with their corresponding vector instructions. This work introduces Super-Node SLP (SN-SLP), a new SLP-style algorithm, optimized for expressions that include a commutative operator (such as addition) and its corresponding inverse element (subtraction). SN-SLP uses the algebraic properties of commutative operators and their inverse elements to enable additional transformations that extend auto-vectorization to cases difficult for state-of-the-art auto-vectorizing compilers. We implemented SN-SLP in LLVM. Our evaluation on a real system demonstrates considerable performance improvements of benchmark code with no significant change in compilation time.
Vasileios Porpodas, Rodrigo Caetano Rocha, Evgueni Brevnov, Fabrício Góes, Timothy G. Mattson
CGO4
2019 Teaching Parallel Programming to Freshmen in an Undergraduate Computer Science Program
abstract
This Research to Practice Full Paper proposes a teaching approach that introduces parallel programming early in the undergraduate Computer Science curriculum. Experiments were conducted to freshmen in the second course of algorithms and data structures. The strategy for the evaluation of the early education of parallel programming includes the use of OpenMP Application Programming Interface and sorting algorithms. The results indicate that students improved their skills by participating in parallel programing activities introduced at early stages or even at the very beginning of the undergraduate program. Freshmen could hit about 92%, 63% and 44% of easy, medium and hard questions after theoretical and practice activities. This represents an improvement about 19%, 14% and 39% for each respective difficulty level in comparison to the beginning of the study when all freshmen had no knowledge relative to parallel programming. These results aid to demystify parallel programming and to show that freshmen can learn it.
Leonardo B. A. Vasconcelos, Felipe A. L. Soares, Pedro Henrique de Mello Morado Penna, Max V. Machado, Fabrício Góes, Carlos Augusto Paiva da Silva Martins, Henrique Cota de Freitas
FIE5
2019 Automatic parallelization of recursive functions with rewriting rules
Rodrigo Caetano Rocha, Fabrício Góes, Fernando Magno Quintão Pereira
Sci. Comput. Program.2
2018 VW-SLP: auto-vectorization with adaptive vector width
abstract
Auto-vectorization techniques allow the compiler to automatically generate SIMD vector code out of scalar code. SLP is a commonly-used algorithm for converting straight-line code into vector code, which complements the loop-based traditional vectorizers. It works by scanning the input code looking for groups of instructions that can be combined into vectors and replacing them with the corresponding vector instructions. The state-of-the-art SLP algorithm works by attempting to vectorize blocks of code with a fixed vector width and falling back to smaller widths for the whole block upon failure.
Vasileios Porpodas, Rodrigo Caetano Rocha, Fabrício Góes
PACT3
2018 Look-ahead SLP: auto-vectorization in the presence of commutative operations
abstract
Auto-vectorizing compilers automatically generate vector (SIMD) instructions out of scalar code. The state-of-the-art algorithm for straight-line code vectorization is Superword-Level Parallelism (SLP). In this work we identify a major limitation at the core of the SLP algorithm, in the performance-critical step of collecting the vectorization candidate instructions that form the SLP-graph data structure. SLP lacks global knowledge when building its vectorization graph, which negatively affects its local decisions when it encounters commutative instructions. We propose LSLP, an improved algorithm that can plug-in to existing SLP implementations, and can effectively vectorize code with arbitrarily long chains of commutative operations. LSLP relies on short-depth look-ahead for better-informed local decisions. Our evaluation on a real machine shows that LSLP can significantly improve the performance of real-world code with little compilation-time overhead.
Vasileios Porpodas, Rodrigo Caetano Rocha, Fabrício Góes
CGO3
2018 HearthBot: An Autonomous Agent Based on Fuzzy ART Adaptive Neural Networks for the Digital Collectible Card Game HearthStone
abstract
Digital collectible card games, as partially observable games based on alternating turns, such as HearthStone, have been the most played card games in recent years, where the main challenge is the creation of strategies capable of subdue the enemy's moves. From the artificial intelligence perspective, the space of possible strategies is large and dynamic due to the number of cards and actions combinations and also to randomness, which makes the design of efficient autonomous agents a hard problem. This paper presents HearthBot, an autonomous agent that plays HearthStone through an adaptive neural network inspired in the fuzzy adaptive resonance associative map and adaptive resonance theory map. This paper also proposes a new mechanism to categorize and predict information to overcome the overgeneralization problem from those networks. Furthermore, the proposed solution was implemented as a parallel adaptive neural network for a graphics processing unit that achieves a performance compatible with the ones obtained for deep learning methods. HearthBot win rate was evaluated in two experiments playing against a Monte Carlo tree search heuristic with competitive decks on a HearthStone simulator called Metastone. Results show that the proposed solution allows HearthBot to obtain an average win rate performance of 80% against known decks and 70% against unknown decks.
Alysson Ribeiro Da Silva, Fabrício Góes
IEEE Trans. Games2
2017 Creative Flavor Pairing: Using RDC Metric to Generate and Assess Ingredients Combination
Alvaro Amorim, Fabrício Góes, Alysson Ribeiro Da Silva, Celso França
ICCC2
2017 TOAST: Automatic tiling for iterative stencil computations on GPUs
abstract
Summary The stencil pattern is important in many scientific and engineering domains, spurring great interest from researchers and industry. In recent years, various optimizations have been proposed for parallel stencil applications running on graphics processing units (GPUs). In particular, tiling is a technique that can significantly enhance application performance by improving data locality and by reducing the volume of communication between host memory and GPU. In addition, tiling enables stencil applications to process inputs that are larger than the physical GPU memory. However, implementing tiling efficiently is complex, time‐consuming, and error‐prone. In this paper, we propose transparently optimized automatic stencil tiling (TOAST), an automatic tiling mechanism for iterative stencil computations running on GPUs; TOAST has 3 main benefits: (1) It incorporates an optimization model that seeks to maximize data reuse within tiles while respecting the amount of dynamically available GPU memory; (2) it offers a virtualized GPU memory for stencil computations, allowing for large input data; and (3) it performs optimal tiling transparently to the developer of the parallel stencil application. The current implementation of TOAST augments the PSkel framework with an internal solver based on genetic algorithms. Our experimental results show that TOAST improves the performance of iterative stencil applications by up to 13 × compared with their multithreaded (central processing unit–based) optimized versions and up to 48 × compared with a naive tiling approach on GPU. The TOAST mechanism is able to automatically achieve a low percentual overhead of data management compared with actual stencil computation.
Rodrigo Caetano Rocha, Alyson D. Pereira, Fabrício Góes
Concurr. Comput. Pract. Exp.4
2017 CAP Bench: a benchmark suite for performance and energy evaluation of low-power many-core processors
abstract
Summary The constant need for faster and more energy‐efficient processors has been stimulating the development of new architectures, such as low‐power many‐core architectures. Researchers aiming to study these architectures are challenged by peculiar characteristics of some components such as networks‐on‐chip and lack of specific tools to evaluate their performance. In this context, the goal of this paper is to present a benchmark suite to evaluate state‐of‐the‐art low‐power many‐core architectures such as the Kalray MPPA‐256 low‐power processor, which features 256 compute cores in a single chip. The benchmark was designed and used to highlight important aspects and details that need to be considered when developing parallel applications for emerging low‐power many‐core architectures. As a result, this paper demonstrates that the benchmark offers a diverse suite of programs with regard to parallel patterns, job types, communication intensity, and task load strategies suitable for a broad understanding of performance and energy consumption of MPPA‐256 and upcoming many‐core architectures. Copyright © 2016 John Wiley & Sons, Ltd.
Matheus Alcântara Souza, Pedro Henrique de Mello Morado Penna, Matheus M. Queiroz, Alyson D. Pereira, Fabrício Góes, Henrique Cota de Freitas, Márcio Castro 0001, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Concurr. Comput. Pract. Exp.5
2017 HoningStone: Building Creative Combos With Honing Theory for a Digital Card Game
abstract
In recent years, online digital games have left behind the status of entertainment sources to become also professional electronic sports. Worldwide championships offer prizes up to millions of dollars for the best competitors and/or teams among different game categories such as digital collectible card games (DCCG), multiplayer online battle arena, etc. Hearthstone, by Blizzard Entertainment, is a DCCG that has an increasing number of players up to the millions. In this game, individual players compete in one-versus-one matches in alternating turns, until a player is defeated. The greatest challenge in this game is to build a deck of cards and a strategy to combine these cards in order to be competitive against other players without a priori knowledge about their decks and strategies. This is a daunting task that requires deep knowledge of each existing card and great amount of creativity to surprise adversaries in this very adaptive environment. This paper presents a computational system, called HoningStone, that automatically generates creative card combos based on the honing theory of creativity. Our experimental results show that HoningStone can generate combos that are more creative than a greedy randomized algorithm driven by a creativity metric.
Fabrício Góes, Alysson Ribeiro Da Silva, João Saffran, Alvaro Amorim, Celso França, Tiago Zaidan, Bernardo M. P. Olimpio, Lucas R. O. Alves, Hugo Morais, Shirley Luana, Carlos Martins
IEEE Trans. Comput. Intell. AI Games1
2016 Regent-Dependent Creativity: A Domain Independent Metric for the Assessment of Creative Artifacts
Celso França, Fabrício Góes, Alvaro Amorim, Rodrigo Caetano Rocha, Alysson Ribeiro Da Silva
ICCC2
2015 PSkel: A stencil programming framework for CPU-GPU systems
abstract
Summary The use of Graphics Processing Units (GPUs) for high‐performance computing has gained growing momentum in recent years. Unfortunately, GPU‐programming platforms like Compute Unified Device Architecture (CUDA) are complex, user unfriendly, and increase the complexity of developing high‐performance parallel applications. In addition, runtime systems that execute those applications often fail to fully utilize the parallelism of modern CPU‐GPU systems. Typically, parallel kernels run entirely on the most powerful device available, leaving other devices idle. These observations sparked research in two directions: (1) high‐level approaches to software development for GPUs, which strike a balance between performance and ease of programming; and (2) task partitioning to fully utilize the available devices. In this paper, we propose a framework, called PSkel, that provides a single high‐level abstraction for stencil programming on heterogeneous CPU‐GPU systems, while allowing the programmer to partition and assign data and computation to both CPU and GPU. Our current implementation uses parallel skeletons to transparently leverage Intel Threading Building Blocks (Intel Corporation, Santa Clara, CA, USA) and NVIDIA CUDA (Nvidia Corporation, Santa Clara, CA, USA). In our experiments, we observed that parallel applications with task partitioning can improve average performance by up to 76% and 28% compared with CPU‐only and GPU‐only parallel applications, respectively. Copyright © 2015 John Wiley & Sons, Ltd.
Alyson D. Pereira, Luiz Eduardo da Silva Ramos, Fabrício Góes
Concurr. Comput. Pract. Exp.3
2014 Adaptive thread mapping strategies for transactional memory applications
Márcio Castro 0001, Fabrício Góes, Jean-François Méhaut
J. Parallel Distributed Comput.2
2012 Dynamic Thread Mapping Based on Machine Learning for Transactional Memory Applications
Márcio Castro 0001, Fabrício Góes, Luiz Gustavo Fernandes, Jean-François Méhaut
Euro-Par2
2012 Autotuning Skeleton-Driven Optimizations for Transactional Worklist Applications
abstract
Skeleton or pattern-based programming allows parallel programs to be expressed as specialized instances of generic communication and computation patterns. In addition to simplifying the programming task, such well structured programs are also amenable to performance optimizations during code generation and also at runtime. In this paper, we present a new skeleton framework that transparently selects and applies performance optimizations in transactional worklist applications. Using a novel hierarchical autotuning mechanism, it dynamically selects the most suitable set of optimizations for each application and adjusts them accordingly. Our experimental results on the STAMP benchmark suite show that our skeleton autotuning framework can achieve performance improvements of up to 88 percent, with an average of 46 percent, over a baseline version for a 16-core system and up to 115 percent, with an average of 56 percent, for a 32-core system. These performance improvements match or even exceed those obtained by a static exhaustive search of the optimization space.
Fabrício Góes, Nikolas Ioannou, Polychronis Xekalakis, Murray Cole, Marcelo Cintra
IEEE Trans. Parallel Distributed Syst.1
2011 A machine learning-based approach for thread mapping on transactional memory applications
abstract
Thread mapping has been extensively used as a technique to efficiently exploit memory hierarchy on modern chip-multiprocessors. It places threads on cores in order to amortize memory latency and/or to reduce memory contention. However, efficient thread mapping relies upon matching application behavior with system characteristics. Particularly, Software Transactional Memory (STM) applications introduce another dimension due to its runtime system support. Existing STM systems implement several conflict detection and resolution mechanisms, which leads STM applications to behave differently for each combination of these mechanisms. In this paper we propose a machine learning-based approach to automatically infer a suitable thread mapping strategy for transactional memory applications. First, we profile several STM applications from the STAMP benchmark suite considering application, STM system and platform features to build a set of input instances. Then, such data feeds a machine learning algorithm, which produces a decision tree able to predict the most suitable thread mapping strategy for new unobserved instances. Results show that our approach improves performance up to 18.46% compared to the worst case and up to 6.37% over the Linux default thread mapping strategy.
Márcio Castro 0001, Fabrício Góes, Christiane Pousa Ribeiro, Murray Cole, Marcelo Cintra, Jean-François Méhaut
HiPC2
2007 A Configuration Control Mechanism Based on Concurrency Level for a Reconfigurable Consistency Algorithm
abstract
A Reconfigurable Consistency Algorithm (RCA) is an algorithm that guarantees the consistency in Distributed Shared Memory (DSM) Systems. In a RCA, there is a Configuration Control Layer (CCL) that is responsible for selecting the most suitable RCA configuration (behavior) for a specific workload and DSM system. In previous works, we defined an upper bound performance for RCA based on an ideal CCL, which knows apriori the best configuration for each situation. This ideal CCL is based on a set of workloads characteristics that, in most situations, are difficult to extract from the applications (percentage of shared write and read operations and sharing patterns). In this paper we propose, develop and present a heuristical configuration control mechanism for the CCL implementation. This mechanism is based on an easily obtained applications parameter, the concurrency level. Our results show that this configuration control mechanism improves the RCA performance in 15%, on average, compared to other traditional consistency algorithms. Furthermore, the CCL with this mechanism is independent from the workload and DSM system specific characteristics, like sharing patterns and percentage of writes and reads.
Christiane V. Pousa, Fabrício Góes, Carlos Augusto Paiva da Silva Martins
IPDPS2
2007 On the efficacy, efficiency and emergent behavior of task replication in large distributed systems
Walfredo Cirne, Francisco Vilar Brasileiro, Daniel Paranhos da Silva, Fabrício Góes, William Voorsluys
Parallel Comput.4
2006 Dynamically reconfigurable cache architecture using adaptive block allocation policy
abstract
In this paper, we present a dynamically reconfigurable cache architecture using adaptive block allocation policy analyzed by means of simulation. Our main objectives are: to propose a reconfigurable cache architecture and to propose, implement and analyze the performance of an adaptive cache block allocation policy. First, we present a proposal of the reconfigurable cache architecture that can adapt according to the workload. Then we present our adaptive policy and do some performance tests comparing our cache architecture with some set associative configurations. In these tests, we use some traces from BYU Trace Distribution Center of SPEC 2000 Benchmark. Finally, we analyze the results based on some metrics like cache miss ratio, response time, etc.
Milene Barbosa Carvalho, Fabrício Góes, Carlos Augusto Paiva da Silva Martins
IPDPS2
2005 Reconfigurable consistency model for object-based software DSM
abstract
Distributed shared memory (DSM) systems can share a set of objects or virtual memory pages. The data sharing enables the applications to access the data concurrently. But, these concurrently access can generate some inconsistencies in the shared data state. The consistency models are responsible for managing the state of shared data for the applications. The already proposed consistency models are inflexible and cannot adapt to the workload and architectures characteristics. So, they cannot generate the best performance for the workloads and architectures in all the cases. In this work, we propose, present and analyze a reconfigurable consistency model for object based DSMs. We called this consistency model ROCoM (reconfigurable object consistency model). ROCoM behavior was represented using a reconfigurable algorithm (RA) and its analysis was made using a simulation tool. Our results show that ROCoM, on average, had 30% better performance than others consistency models.
Christiane V. Pousa, Fabrício Góes, Carlos Augusto Paiva da Silva Martins
CCGRID2
2005 Reconfigurable Object Consistency Model for Distributed Shared Memory
Christiane V. Pousa, Fabrício Góes, Carlos Augusto Paiva da Silva Martins
ISPA2
2005 AnthillSched: A Scheduling Strategy for Irregular and Iterative I/O-Intensive Parallel Jobs
Fabrício Góes, Pedro Henrique Calais Guerra, Bruno Coutinho, Leonardo Rocha 0001, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Walfredo Cirne
JSSPP1
2004 ClusterSim: a Java-based parallel discrete-event simulation tool for cluster computing
abstract
We present the proposal and implementation of a Java-based parallel discrete-event simulation tool for cluster computing called ClusterSim (cluster simulation tool). The ClusterSim supports visual modeling and simulation of clusters and their workloads for performance analysis. A cluster is composed of single or multiprocessed nodes, parallel job schedulers, network topologies and technologies. A workload is represented by users that submit jobs composed of tasks described by probability' distributions and their internal structure (CPU, I/O and MPI instructions). Our main objectives in This work: to present the proposal and implementations of the software architecture and simulation model of ClusterSim; to verify and validate ClusterSim; to analyze ClusterSim by means of a case study. Our main contributions are: the proposal and implementation of ClusterSim with an hybrid workload model, a graphical environment, the modeling of heterogeneous clusters and a statistical and performance module.
Fabrício Góes, Luiz Eduardo da Silva Ramos, Carlos Augusto Paiva da Silva Martins
CLUSTER1
2004 Reconfigurable Gang Scheduling Algorithm
Fabrício Góes, Carlos Augusto Paiva da Silva Martins
JSSPP1
2002 Performance Analysis of Parallel Programs Using Prober as a Single Aid Tool
abstract
Prober is a functional and performance analysis tool for parallel programs, developed during an undergraduate research project. In this paper we show the new expanded version of Prober, in which some features from different software tools are aggregated. It can be used as a single tool to aid the developer in the performance analysis of parallel programs. Our main goal is to provide a new version of Prober, with additional features. Among them we can highlight: the interpretation of user scripts, a user-level support library, the generation of speedup and efficiency graphics, batch execution and a new user interface. In order to show, verify and analyze the use of the new version of Prober, we did performance tests in a parallel image convolution program. We added performance measuring routines to collect performance data within different internal code segments; built a set of scripts to specify the performance tests; ran the set of scripts in batch mode; used Prober to generate graphics and statistics based on the collected performance data; and analyzed the results using Prober as an aid tool.
Fabrício Góes, Luiz Eduardo da Silva Ramos, Carlos Augusto Paiva da Silva Martins
SBAC-PAD1