Henrique Cota de Freitas

dblp:13/4670 · also Henrique C. Freitas · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
5since 2021 · last 2026
0000-0001-9722-1093ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 CertiCheck-FTA: Interpretable Fraud Detection in Medical Certificates Using Multimodal LLMs
Douglas Silva de Moura, Sandro José Rigo, Henrique Cota de Freitas, Rodrigo da Rosa Righi
CLOSER3
2024 Automated and intelligent system for World Health Organization data forecasting
abstract
Abstract Technological advances and social transformations have enabled the circulation of a large amount of data in the health area. Analyzing this data becomes more critical and more challenging as the volume of data increases. An alternative to performing this analysis is to use data analysis techniques to process input data sets and build consistent databases for input to machine learning algorithms. Thus, it can forecast future scenarios and collaborate with knowledge discovery. In this context, this work aims to develop a parameterizable system with automated decisions capable of collecting and analyzing many indicators provided by the World Health Organization (WHO). After these analysis steps, the system applies machine learning algorithms for predictions of different indicators to process information automatically, finding different knowledge discovery scenarios. Thus, the contribution of this article is an automated and intelligent system for WHO data forecasting. The efficiency of the system's choices and forecasts was proven with experiments in five different areas of health, obtaining assertiveness by up to 99.92%, root‐mean‐square error (RMSE) by up to 0.0286, and Kling‐Gupta Efficiency (KGE) by up to 0.9988, hitting even the most complex cases, as shown in the confusion matrices. Finally, three case studies were carried out to expand the studies and present the potential of the system in different contexts: anemia in children, age‐standardized suicide rates in men, and number of road traffic deaths.
Felipe A. L. Soares, Henrique Cota de Freitas
Expert Syst. J. Knowl. Eng.2
2023 LWMPI: An MPI library for NoC-based lightweight manycore processors with on-chip memory constraints
abstract
Abstract Lightweight manycore processors deliver high performance and energy efficiency by bundling hundreds of low‐power cores, a distributed memory architecture with small local memories and Networks‐on‐Chip in a single die. However, the lack of rich and portable programming models for these processors makes software development a challenging task. Currently, two approaches are employed to address programmability in lightweight manycores: Operating Systems (OSes) and baremetal runtime libraries. The former provides portability but exposes complex Operating System (OS)‐level programming interfaces to developers. The latter focuses on providing rich and high performance interfaces, which are vendor‐specific and yield to non‐portable software. In this work, we address these programmability and portability challenges by combining a rich OS with a well‐known standard for parallel programming. We propose a portable and lightweight Message Passing Interface (MPI) library (LWMPI) designed from scratch to cope with restrictions and intricacies of lightweight manycores. We integrated LWMPI into Nanvix, an open‐source distributed OS that runs on silicon lightweight manycores. The results obtained with a synthetic benchmark and a subset of the CAP Bench applications running on Kalray MPPA‐256 unveil that LWMPI not only delivers a lightweight and richer programming interface but also presents good performance and scalability results.
João Fellipe Uller, João Vicente Souto, Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut
Concurr. Comput. Pract. Exp.5
2021 An open computing language-based parallel Brute Force algorithm for formal concept analysis on heterogeneous architectures
abstract
Abstract Algorithms for the extraction of formal concepts are widely studied in several areas of knowledge, such as finance, health, and statistics. However, these algorithms require high‐performance processing due to their combinatorial characteristics. In this work, an Open computing language (OpenCL)‐based Brute Force algorithm is proposed and evaluated for formal concept extraction on heterogeneous architectures (CPU+GPU and CPU+FPGA). The CPU+GPU architecture presents higher performance and scalability than other architectures when our Brute Force algorithm processes high dimensional contexts with many objects and attributes. Our parallel approach shows performance results up to 18× better than a smarter sequential algorithm called Data‐Peeler. Moreover, our Brute Force algorithm running on CPU+GPU architecture has greater energy efficiency, reaching at least 1.79× more operations per energy consumption than other algorithms on different architectures explored in this work.
João P. P. Novais, Lucas Andrade Maciel, Matheus Alcântara Souza, Mark A. J. Song, Henrique Cota de Freitas
Concurr. Comput. Pract. Exp.5
2021 Inter-kernel communication facility of a distributed operating system for NoC-based lightweight manycores
Pedro Henrique de Mello Morado Penna, João Vicente Souto, João Fellipe Uller, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut
J. Parallel Distributed Comput.5
2019 Teaching Parallel Programming to Freshmen in an Undergraduate Computer Science Program
abstract
This Research to Practice Full Paper proposes a teaching approach that introduces parallel programming early in the undergraduate Computer Science curriculum. Experiments were conducted to freshmen in the second course of algorithms and data structures. The strategy for the evaluation of the early education of parallel programming includes the use of OpenMP Application Programming Interface and sorting algorithms. The results indicate that students improved their skills by participating in parallel programing activities introduced at early stages or even at the very beginning of the undergraduate program. Freshmen could hit about 92%, 63% and 44% of easy, medium and hard questions after theoretical and practice activities. This represents an improvement about 19%, 14% and 39% for each respective difficulty level in comparison to the beginning of the study when all freshmen had no knowledge relative to parallel programming. These results aid to demystify parallel programming and to show that freshmen can learn it.
Leonardo B. A. Vasconcelos, Felipe A. L. Soares, Pedro Henrique de Mello Morado Penna, Max V. Machado, Fabrício Góes, Carlos Augusto Paiva da Silva Martins, Henrique Cota de Freitas
FIE7
2019 A comprehensive performance evaluation of the BinLPT workload-aware loop scheduler
abstract
Summary Workload‐aware loop schedulers were introduced to deliver better performance than classical loop scheduling strategies. However, they presented limitations such as inflexible built‐in workload estimators and suboptimal chunk scheduling. Targeting these challenges, we proposed previously a workload‐aware scheduling strategy called BinLPT, which relies on three features: (i) user‐supplied estimations of the workload of the loop; (ii) a greedy heuristic that adaptively partitions the iteration space in several chunks; and (iii) a scheduling scheme based on the Longest Processing Time (LPT) rule and on‐demand technique. In this paper, we present two new contributions to the state‐of‐the‐art. First, we introduce a multiloop support feature to BinLPT, which enables the reuse of estimations across loops. Based on this feature, we integrated BinLPT into a real‐world elastodynamics application, and we evaluated it running on a supercomputer. Second, we present an evaluation of BinLPT using simulations as well as synthetic and application kernels. We carried out this analysis on a large‐scale NUMA machine under a variety of workloads. Our results revealed that BinLPT is better at balancing the workloads of the loop iterations and this behavior improves as the algorithmic complexity of the loop increases. Overall, BinLPT delivers up to 37.15% and 9.11% better performance than well‐known loop scheduling strategies, for the application kernels and the elastodynamics simulation, respectively.
Pedro Henrique de Mello Morado Penna, Antônio Tadeu A. Gomes, Márcio Castro 0001, Patricia Della Méa Plentz, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut
Concurr. Comput. Pract. Exp.5
2018 Design Space Exploration of Energy Efficient NoC-and Cache-Based Many-Core Architecture
abstract
Performance of parallel scientific applications on many-core processor architectures is a challenge that increases every day, especially when energy efficiency is concerned. To achieve this, it is necessary to explore architectures with high processing power composed by a network-on-chip to integrate many processing cores and other components. In this context, this paper presents a design space exploration over NoC-based manycore processor architectures with distributed and shared caches, using full-system simulations. We evaluate bottlenecks in such architectures with regard to energy efficiency, using different parallel scientific applications and considering aspects from caches and NoCs jointly. Five applications from NAS Parallel Benchmarks were executed over the proposed architectures, which vary in number of cores; in L2 cache size; and in 12 types of NoC topologies. A clustered topology was set up, in which we obtain performance gains up to 30.56% and reduction in energy consumption up to 38.53%, when compared to a traditional one.
Matheus Alcântara Souza, Henrique Cota de Freitas, Jean-François Méhaut
SBAC-PAD2
2018 Energy Efficient Parallel K-Means Clustering for an Intel® Hybrid Multi-Chip Package
abstract
FPGA devices have been proving to be good candidates to accelerate applications from different research topics. For instance, machine learning applications such as K-Means clustering usually relies on large amount of data to be processed, and, despite the performance offered by other architectures, FPGAs can offer better energy efficiency. With that in mind, Intel has launched a platform that integrates a multicore and an FPGA in the same package, enabling low latency and coherent fine-grained data offload. In this paper, we present a parallel implementation of the K-Means clustering algorithm, for this novel platform, using OpenCL language, and compared it against other platforms. We found that the CPU+FPGA platform was more energy efficient than the CPU-only approach from 70.71% to 85.92%, with Standard and Tiny input sizes respectively, and up to 68.21% of performance improvement was obtained with Tiny input size. Furthermore, it was up to 7.2×more energy efficient than an Intel® Xeon Phi ™, 21.5×than a cluster of Raspberry Pi boards, and 3.8×than the low-power MPPA-256 architecture, when the Standard input size was used.
Matheus Alcântara Souza, Lucas Andrade Maciel, Pedro Henrique de Mello Morado Penna, Henrique Cota de Freitas
SBAC-PAD4
2017 Design methodology for workload-aware loop scheduling strategies based on genetic algorithm and simulation
abstract
Summary In high‐performance computing, the application's workload must be evenly balanced among threads to deliver cutting‐edge performance and scalability. In OpenMP, the load balancing problem arises when scheduling loop iterations to threads. In this context, several scheduling strategies have been proposed, but they do not take into account the input workload of the application and thus turn out to be suboptimal. In this work, we introduce a design methodology to propose, study, and assess the performance of workload‐aware loop scheduling strategies. In this methodology, a genetic algorithm is employed to explore the state space solution of the problem itself and to guide the design of new loop scheduling strategies, and a simulator is used to evaluate their performance. As a proof of concept, we show how the proposed methodology was used to propose and study a new workload‐aware loop scheduling strategy named smart round‐robin (SRR). We implemented this strategy into GNU Compiler Collection's OpenMP runtime. We carry out several experiments to validate the simulator and to evaluate the performance of SRR. Our experimental results show that SRR may deliver up to 37.89%and 14.10%better performance than OpenMP's dynamic loop scheduling strategy in the simulated environment and in a real‐world application kernel, respectively. Copyright © 2016 John Wiley & Sons, Ltd.
Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut
Concurr. Comput. Pract. Exp.3
2017 CAP Bench: a benchmark suite for performance and energy evaluation of low-power many-core processors
abstract
Summary The constant need for faster and more energy‐efficient processors has been stimulating the development of new architectures, such as low‐power many‐core architectures. Researchers aiming to study these architectures are challenged by peculiar characteristics of some components such as networks‐on‐chip and lack of specific tools to evaluate their performance. In this context, the goal of this paper is to present a benchmark suite to evaluate state‐of‐the‐art low‐power many‐core architectures such as the Kalray MPPA‐256 low‐power processor, which features 256 compute cores in a single chip. The benchmark was designed and used to highlight important aspects and details that need to be considered when developing parallel applications for emerging low‐power many‐core architectures. As a result, this paper demonstrates that the benchmark offers a diverse suite of programs with regard to parallel patterns, job types, communication intensity, and task load strategies suitable for a broad understanding of performance and energy consumption of MPPA‐256 and upcoming many‐core architectures. Copyright © 2016 John Wiley & Sons, Ltd.
Matheus Alcântara Souza, Pedro Henrique de Mello Morado Penna, Matheus M. Queiroz, Alyson D. Pereira, Fabrício Góes, Henrique Cota de Freitas, Márcio Castro 0001, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Concurr. Comput. Pract. Exp.6
2015 On the energy efficiency and performance of irregular application executions on multicore, NUMA and manycore platforms
Emilio Francesquini, Márcio Castro 0001, Pedro Henrique de Mello Morado Penna, Fabrice Dupros, Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
J. Parallel Distributed Comput.5
2013 Method for teaching parallelism on heterogeneous many-core processors using research projects
abstract
Parallel programming and parallel architectures are necessary to achieve scalability and performance. It is difficult to evaluate when to teach parallelism and how to change the paradigm from serial to parallel algorithm in traditional curricula. Currently, there are efforts to introduce parallel programming since there are multi-core processors. However, there is a new chip generation called many-core processor. For instance, one processor chip can be built with 1,000 processing cores. Moreover, this type of processor is designed to achieve scalability and performance based on heterogeneous cores. How to teach parallelism to undergraduate and graduate students? Human resources are necessary to design and program parallel architectures based on this next generation of many-core processor. Therefore, the main goal of this paper is to show an experience based on research projects. The idea is to join students from different courses and levels, e.g. Computer Science, Information Systems, Computer Engineering, and Graduate in Informatics. All of them working together in order to understand all characteristics of heterogeneous many-core processors based on integrated environment composed of computer clusters and simulation. The proposed method focuses on projects convergence to teach how to extract characteristics from benchmark traces in order to simulate many-core processors based on Networks-on-Chip. Consequently, students can understand parallel heterogeneous architectures and how to program them. Main results present the number of students interested in this research field along last three years and several scientific papers published. It is important to highlight that papers have students' participation and two papers are related to education. Results reinforce the contribution of the proposed method since we have several benefits including continuous and cooperative work along more years.
Henrique Cota de Freitas
FIE1
2012 Introducing parallel programming to traditional undergraduate courses
abstract
Parallel programming is an important issue for current multi-core processors and necessary for new generations of many-core architectures. This includes processors, computers, and clusters. However, the introduction of parallel programming in undergraduate courses demands new efforts to prepare students for this new reality. This paper describes an experiment on a traditional Computer Science course during a two-year period. The main focus is the question of when to introduce parallel programming models in order to improve the quality of learning. The goal is to propose a method of introducing parallel programming based on OpenMP (a shared-variable model) and MPI (a message-passing model). Results show that when the OpenMP model is introduced before the MPI model the best results are achieved. The main contribution of this paper is the proposed method that correlates several concepts such as concurrency, parallelism, speedup, and scalability to improve student motivation and learning.
Henrique Cota de Freitas
FIE1
2012 Parallel and distributed kmeans to identify the translation initiation site of proteins
abstract
Prediction of the translation initiation site is of vital importance in bioinformatics since through this process it is possible to understand the organic formation and metabolic behavior of living organisms. Sequential algorithms are not always a viable solution due to the fact that mRNA databases are normally very large, resulting in long processing times. Applying parallel and distributed computing resources to such databases could help reduce this time. The objective of this article is to present a class balancing solution for the translation initiation site process using parallel and distributed computing resources in a hybrid model. The results reveal a speedup of up to 23 times compared to sequential methods and performance rates for accuracy, precision, sensitivity, specificity and adjusted accuracy of 91.15%, 39.83%, 89.11%, 88.93% and 89.02%, respectively, for the Homo sapiens database. For the Drosophila melanogaster database, the speedup was 18.33 times and accuracy, precision, sensitivity, specificity and adjusted accuracy were 95.22%, 43.01%, 90.83%, 90.47% and 90.64%, respectively. Both sets of results are considered important. Thus, the solution presented in this article demonstrated itself viable for the problem in question.
Laerte M. Rodrigues, Luis E. Zárate, Cristiane Nobre, Henrique Cota de Freitas
SMC4
2010 Impact of Parallel Workloads on NoC Architecture Design
abstract
Due to the multi-core processors, the importance of parallel workloads has increased considerably. However, many-core chips demand new interconnection strategies, since traditional crossbars or buses, common for current multi-core processors, have problems related to wires and scalability. For this reason, Networks-on-Chip (NoCs) have been developed in order to support the performance and parallelism focused on several workloads. Although a Network-on-Chip is a good option, most designs consist of a large number of routers. These routers are responsible for forwarding packets, and consequently, for supporting message-passing workloads. In this context, the NoC performance is a problem. Therefore, the main goal of this paper is to evaluate the impact of well-known parallel workloads on NoC architecture design. In order to achieve high performance, the results point out to parallel workloads with small packets and cluster-based NoCs with circuit switching and adaptable topologies.
Henrique Cota de Freitas, Lucas Mello Schnorr, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
PDP1
2010 Parallel Shared-Memory Workloads Performance on Asymmetric Multi-core Architectures
abstract
Putting performance asymmetric cores inside the same processor can be a good alternative to obtain high performance per area, throughput and single-threaded performance. However, the impact of running parallel applications on this type of machine is not clear, since most of previous work focused on multi-programmed and server workloads where there is low or no dependence between threads. In this work, we analyze the impact of running parallel shared-memory programs on heterogeneous multi-core setups using six parallel applications with diverse parallelization schemes. Moreover, we show that, in some cases, with a high number of cores, it is better to put one complex core than several simple ones. The impact of sharing the address space between asymmetric cores with private caches was also investigated and the number of invalidations per write access was not greater than a comparable homogeneous configuration.
Felipe Lopes Madruga, Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux
PDP2
2009 On the design of reconfigurable crossbar switch for adaptable on-chip topologies in programmable NoC routers
abstract
Research works have focused on high-performance on-chip interconnections with low cost and energy consumption for the next generation of many-core processors. In the same way, parallel applications will explore thread level parallelism and message-passing communication through a Network-on-Chip (NoC) to perform a high data throughput. Due to the dynamic changing of communication patterns, the goal of this paper is to present the design of Reconfigurable Crossbar Switch for NoC Routers (RCS-NR) capable of adapting topologies on demand. RCS-NR has an optimized architecture, similar area and lower energy consumption (up to 87.41%) relative to a traditional crossbar switch. Furthermore, RCS-NR-based NoC router has a similar area, higher throughput, and it is up to 98.76% more efficient in energy consumption than a conventional NoC.
Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux
ACM Great Lakes Symposium on VLSI1
2009 Design of Interleaved Multithreading for Network Processors on Chip
abstract
Thread level parallelism and multi-core processors are current alternatives to increase performance of general-purpose applications. In the same way, networks-on-ohip (NoCs) are the main alternatives for supporting packet throughput for the next generations of many-core processors. NPoC (network processor on chip) is a proposal to increase the performance of programmable NoC routers and multi-cluster NoC architectures using interleaved multithreading (IMT) technique. Therefore, the main goal of this paper is to present the design impact of interleaved multithreading for network processors on chip focusing on area and performance feasibility. Results show that NPoC-based router has an acceptable and similar area relative to a conventional NoC, and higher performance up to 7.1% than the same NPoC version without IMT.
Henrique Cota de Freitas, Felipe Lopes Madruga, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
ISCAS1
2009 Performance Evaluation of NoC Architectures for Parallel Workloads
abstract
Network-on-Chip is the state-of-the-art approach to interconnect many processing cores in the next generation of general-purpose processors. In this context, the problem is to choose NoC architectures capable of achieving high performance for parallel programs. Therefore, the main goal of this paper is to evaluate the performance of three NoC architectures using well-known parallel workloads.
Henrique Cota de Freitas, Marco A. Z. Alves, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
NOCS1
2008 NOC architecture design for multi-cluster chips
abstract
For the next generation of multi-core processors, the on-chip interconnection networks must be efficient to achieve high data throughput and performance. Moreover, these interconnections must be flexible and scalable in order to provide parallel on-demand computing. For this reason, the goal of this paper is to present design decisions of a multi-cluster NoC (MCNoC) architecture in order to support collective communication patterns through topology reconfiguration on an FPGA-based multi-cluster chip. The MCNoCpsilas results show a small area occupation, low power consumption and high performance.
Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux, Tatiana Gadelha Serra dos Santos
FPL1
2007 Evaluating Network-on-Chip for Homogeneous Embedded Multiprocessors in FPGAs
abstract
This paper presents performance and area evaluation of a homogeneous multiprocessor communication system based on network-on-chip (NoC) in FPGA platforms. Two homogenous chip multiprocessor proposals were designed and compared for Xilinx FPGAs using MicroBlaze processors: one based on NoC and the other based on shared memory/bus. One of the main findings is the communication performance evaluation of NoC for parallel computing applications. The comparison results show that an efficient implementation of NoC on FPGA can improve communication speed by up to seven times with low area overhead, according to the data size and the number of processors connected to the network.
Henrique Cota de Freitas, Dalton Martini Colombo, Fernanda Lima Kastensmidt, Philippe Olivier Alexandre Navaux
ISCAS1
2006 Reconfigurable crossbar switch architecture for network processors
abstract
This paper presents the proposal and development of a reconfigurable crossbar switch (RCS) architecture for network processors. Its main purpose is to increase the performance, and flexibility for environments with multiprocessors and computer clusters. The results include VHDL simulation of RCS and the use of it in a broadcast function implementation, found in message passing support middleware.
Henrique Cota de Freitas, Milene Barbosa Carvalho, Alexandre Marques Amaral, Amanda R. M. Diniz, Carlos Augusto Paiva da Silva Martins, Luiz Eduardo da Silva Ramos
ISCAS1