VLDB 2026 Research / reviewers in the wild / expert
Rodolfo Azevedo
dblp:73/1601 · also Rodolfo Jardim de Azevedo
· DBLP profile ↗
47ranked-venue papers
1as first author
8since 2021 · last 2024
0000-0002-8803-0401ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 since 2021Software engineering, systems software and programming languages · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learn by Example in a Modern Embedded System CourseabstractLearning by example is a widely used technique in programming education and is explored in embedded systems courses. This demonstration introduces an active learning methodology whereby students begin with exercises featuring example codes and then proceed to develop their own code base. This personally developed code is subject to an automated validation system that ensures its correctness, both in functionality and code quality. Consequently, students acquire a code base of their own creation, serving as a reference for devising solutions to real-world problems. Rafael Corsi Ferrão, Igor dos Santos Montagner, Rodolfo Azevedo |
ITiCSE (2) | 3 |
| 2024 | Embedded-check a Code Quality Tool for Automatic Firmware VerificationabstractDeveloping embedded microcontroller code is a complex task, especially for undergrad students new to this area. These students often make high-level conceptual mistakes beyond the scope of commercial standards like MISRA-C. These conceptual errors need to be checked manually through code feedback, a process that is time-consuming, error-prone, and does not scale well with an increasing number of students and/or assignments. In this paper, we present an embedded-check an automated tool that can detect common and critical errors students make when learning to code firmware. A set of 13 rules (baremetal and FreeRTOS) was devised based on our experience from several years of teaching Embedded systems. To validate our tool, we compared its results with manual code review of N=99 projects from the last 3 course offerings. We furthered our analysis by running our tool on N=1132 coding lab submissions that did not receive manual feedback and were used as part of classroom activities. We found that the top-3 errors flagged in the projects were already present when students completed the lab activities. We found that (i) our tool also identified all issues discovered during manual code feedback, (ii) our tool detected issues in 86% of student submissions, whereas manual code feedback only flagged 28% of the submissions as problematic, and (iii) 94.3% of students made some code quality error on individual assignments. Within this results, we believe that our tool can have a significant impact when used both as an formative assessment tool to support learning and as a learning analytics tool to improve teaching. Rafael Corsi Ferrão, Igor dos Santos Montagner, Mariana Silva, Craig B. Zilles, Rodolfo Azevedo |
ITiCSE (1) | 5 |
| 2023 | Using Logging-on-Write to Improve Non-Volatile Memory Checkpoints via Processing-in-MemoryabstractNVM architectures must keep consistent data in case of failures, a property called crash consistency. A common way to do so is by checkpoint mechanisms. However, most of the strategies developed have performance and usability problems. Among main limitations are non-software-transparent strategies, the addition of logging operations in the critical execution path, and increased writes to NVM, resulting in significant bandwidth usage between processor and memory. DONUTS solves these problems by a hardware mechanism that provides crash consistency via checkpoints integrated into cache replacement policy using Processing-in-Memory to perform logging operations. Its approach reduces writes from the processor's memory controller and the NVM external bandwidth usage but generates unnecessary log entries. This paper expands DONUTS to a multi-core scenario evaluating two processing-in-memory strategies. The first logs during read operations, and the second uses a new lazy strategy to log data exclusively on write operations when these operations are effectively needed. Finally, we compare runtime performance, log rate, energy consumption, and memory space. Results show that our new logging-on-write strategy maintained the DONUTS runtime performance but reduced energy consumption in multi-core applications by around 33% due to an average reduction of 42% in log operations. Also, the new strategy generates a checkpoint size 5x smaller than the previous system, maximizing the use of NVM. Compared to other systems, DONUTS presented an average overhead of 1% to 3% against up to 7% of previous software-transparent better-performing projects. Kleber Kruger, Ricardo Pannain, Rodolfo Azevedo |
SBAC-PAD | 3 |
| 2023 | DONUTS: An efficient method for checkpointing in non-volatile memoriesabstractSummary Non‐volatile memory (NVM) is an emerging technology being explored as an alternative to DRAM main memory in computing systems because of its persistence, higher storage density, lower energy consumption, and access latency close to DRAM. However, persistent memory systems must ensure data consistency on system failures, a property known as crash consistency. One of the main challenges in these systems is creating efficient checkpointing mechanisms in terms of performance and usability. Thus, it is necessary to remove persistence from the critical execution path and reduce the excessive number of writes to NVM caused by logging operations, resulting in increased memory bandwidth usage. Another limitation is that most proposed mechanisms restrict application source code to programming interfaces based on transactional models, typed as software‐based approaches. These factors make it challenging to adopt NVM systems. This article presents a software‐transparent mechanism based on dynamic epochs with logging operations via processing‐in‐memory and checkpoints integrated into the cache replacement policy. Compared to the previous best‐performing system, our strategy reduces 50.6% of writes to NVM. Furthermore, it does not increase the average memory bandwidth usage, providing crash consistency with less than 2% runtime overhead. Kleber Kruger, Ricardo Pannain, Rodolfo Azevedo |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | A Computational Thinking Course for Pre-Service TeachersabstractComputational Thinking (CT) is a set of thinking processes used by computer scientists to formulate problems and describe solutions. In recent years, CT has been largely explored by the Computing Education community. Due to its potential to contribute to problem solving and analytical thinking, CT could also be a relevant subject in K-12 education. For such, there is a need to qualify educators in CT. In this paper, we describe our experience in a large-scale course on Computational Thinking for pre-service teachers. Our goal is to qualify pre-service teachers in CT, both in the comprehension of its definition and how it can be incorporated in classroom. Our experience indicates that the participants had a positive awareness of CT concepts, as well as an adequate understanding of the pedagogical approaches to teach the subject. They also developed positive attitudes towards the field of Computer Science. Eduardo C. Oliveira, Ronaldo Celso Messias Correia, Rodolfo Azevedo, Simone Telles, Alessandra Alaniz Macedo, Roberto A. Bittencourt |
EDUCON | 3 |
| 2022 | How much C can students learn in one week? Experiences teaching C in advanced CS coursesabstractIn this paper Innovative Practice Full Paper we analyze the results of the application of a Concept Inventory for introductory C programming when teaching C for students in advanced CS courses and propose improvements based on misconceptions found and student perceptions. Even though Python is now the most used language for introductory courses, many advanced CS courses, such as Operating Systems and Em-bedded Computing, still benefit from using the C language. Thus, instructors now face a new challenge: teaching C to students who are already proficient in a higher-level language. Since this must be done in addition to their existing course, it is necessary to (i) be fast; and (ii) assure that students learned enough to be successful in the courses. An approach described in previous work was to concentrate all courses that benefit from using C (Computer Systems, Embedded Computing and Programming Challenges) in a single semester and joining their classes for the first week of the semester. This resulted in promising preliminary results, but lacked a more rigorous assessment in order to draw more meaningful conclusions. Concept Inventories (CI) can be used to both assess how many students have accurate knowledge of the subject but also to identify common pitfalls and misconceptions. In this work we combine three years of experiences in teaching C in advanced CS courses with data from the application of an Introductory C programming CI. We have received 109 responses from three consecutive offerings, detailing student performance in seven key areas. Satisfactory performance was obtained in Parameter use and scope, Iteration, Recursion, Structures and Pointers. Variable Scope and Boolean Expressions were the most challenging areas. We analyze the most common misconceptions and relate them to the material and students’ previous experiences. A questionnaire answered by students (N=56) at the end of the activity is analyzed in order to understand the level of affinity of students with the C language before the course and their perceptions about it, we noticed that most students have not had contact with the language before and feel motivated to learn C. To verify the knowledge in the language on a practical project, we analyzed the first big delivery of code that they have to do after two weeks the end of the course , through the analysis of the students codes (N=58) we can measure that many students use more advanced resources of the language that were not seen in the course or taught officially by other courses, indicating that they were able to learn enough of the language to go it alone in more complex matters. Rafael Corsi Ferrão, Igor dos Santos Montagner, Ricardo Caceffo, Rodolfo Azevedo |
FIE | 4 |
| 2022 | Prof5: A RISC-V profiler toolabstractRISC-V is supported by a series of design and simulation tools that enable simple instruction set customization and rapid exploration of application-specific accelerators. Evaluating the performance and energy impact of specific design choices and optimizations on applications remains, however, challenging. Traditional RT- or Gate-level simulation, while fairly precise, is complex and slow, and is, therefore, typically limited to small fractions of code. Functional simulation, while faster, is typically imprecise and lacks the detailed information presented by traditional profilers. We introduce Prof5, a profiler for RISC-V designs that combines functional simulation with precise energy and timing models calibrated from RTL simulation and power analysis. Prof5 is based on the Spike simulator and provides detailed, function-level timing and energy statistics that can be used to guide design and optimization choices, and enable rapid design-space exploration. Prof5 can furthermore aid the user in creating new timing and energy models for custom designs and architecture variations. Energy and timing estimation with Prof5 is 8000x faster than traditional synthesis-based analysis with an average of 95% accuracy for an embedded RISC-V processor. Jonathas Silveira, Lucas Castro, Rodrigo Zeli, Daniel Lazari, Marcelo Guedes, Rodolfo Azevedo, Lucas Francisco Wanner |
SBAC-PAD | 7 |
| 2022 | Optically connected memory for disaggregated data centers
Mauricio G. Palma, Maarten Hattink, Ruth Rubio-Noriega, Lois Orosa 0001, Onur Mutlu, Keren Bergman, Rodolfo Azevedo |
J. Parallel Distributed Comput. | 8 |
| 2020 | Optically Connected Memory for Disaggregated Data CentersabstractRecent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.07 pJ energy-per-bit consumption and (2) OCM performs up to 5.5x faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes. Alexander Gazman, Maarten Hattink, Mauricio G. Palma, Meisam Bahadori, Ruth Rubio-Noriega, Lois Orosa 0001, Madeleine Glick, Onur Mutlu, Keren Bergman, Rodolfo Azevedo |
SBAC-PAD | 11 |
| 2020 | ADeLe: A description language for approximate hardware
Isaías B. Felzmann, Matheus Martins Susin, Liana Dessandre Duenha, Rodolfo Azevedo, Lucas Francisco Wanner |
Future Gener. Comput. Syst. | 4 |
| 2020 | AxRAM: A lightweight implicit interface for approximate data access
João Fabrício Filho, Isaías B. Felzmann, Rodolfo Azevedo, Lucas Francisco Wanner |
Future Gener. Comput. Syst. | 3 |
| 2019 | Identifying and Validating Java Misconceptions Toward a CS1 Concept InventoryabstractA misconception is a common misunderstanding that students may have about a specific topic. The identification, documentation, and validation of misconceptions is a long and time-consuming work, usually carried out using iterative cycles of students answering open-ended questionnaires, interviews with instructors and students, exam analysis, and discussion with experts. A comprehensive list of validated misconceptions in some subject can be used to build formal evaluation methods like the Concept Inventory (CI), a multiple-choice questionnaire that is usually performed as pre-post tests in order to assess any change in student understanding. In CS1, validated misconceptions were identified and documented in C and Python programming languages. Although there are studies related to misconceptions in the Java language, these misconceptions lack the formality, comprehensiveness, and robustness of their C and Python counterparts. On this work, we propose a methodology to adapt the validated misconceptions in C and Python to Java. Initially, through the analysis of an initial list of 33 misconceptions in C and 28 in Python, we identified and documented in an antipattern format 31 possible misconceptions in Java. We then developed a final term exam, composed of 7 open-ended questions, in which each question was designed to address some of the misconceptions covered in the course (N=27). Through the analysis of the exams answers (N = 69 students), it was possible to validate 22 of the misconceptions (81%). Also, 6 new misconceptions were identified, leading to a total of 28 valid misconceptions in Java. Ricardo Caceffo, Pablo Frank-Bolton, Rodolfo Azevedo |
ITiCSE | 4 |
| 2019 | Towards a Transprecision Polymorphic Floating-Point Unit for Mixed-Precision ComputingabstractMixed-precision is a paradigm that tries to combine computations with different levels of precision to compose results. This approach has been used extensively to optimize scientific applications and has shown speed and energy gains, without causing any relevant precision loss. However, to exploit mixed precision opportunities most applications need to be recompiled to use different instructions and types. Thus, in this work, we present a new floating-point unit design, able to automatically decide when an instruction should be executed using less precision, without recompilation or user direct intervention. Our proposal takes advantage of ad-hoc polymorphism to perform computations with different data types, dynamically selecting a proper instruction on demand, and can also be configured according to overall precision requirements. Our simulated results show that, for some double-precision benchmarks, we are able to execute more than 90% of all floating-point operations in half-precision, without affecting its accuracy and resulting in a precision error below 1%. In addition, this new technology may increase instruction level parallelism and cut down the necessity for type casting operations. Alisson Carvalho, Rodolfo Azevedo |
SBAC-PAD | 2 |
| 2019 | AVPP: Address-first Value-next Predictor with Value Prefetching for Improving the Efficiency of Load Value PredictionabstractValue prediction improves instruction level parallelism in superscalar processors by breaking true data dependencies. Although this technique can significantly improve overall performance, most of the state-of-the-art value prediction approaches require high hardware cost, which is the main obstacle for its wide adoption in current processors. To tackle this issue, we revisit load value prediction as an efficient alternative to the classical approaches that predict all instructions. By speculating only on loads, the pressure over shared resources (e.g., the Physical Register File) and the predictor size can be substantially reduced (e.g., more than 90% reduction compared to recent works). We observe that existing value predictors cannot achieve very high performance when speculating only on load instructions. To solve this problem, we propose a new, accurate and low-cost mechanism for predicting the values of load instructions: the Address-first Value-next Predictor with Value Prefetching (AVPP). The key idea of our predictor is to predict the load address first (which, we find, is much more predictable than the value) and to use a small non-speculative Value Table (VT)—indexed by the predicted address—to predict the value next. To increase the coverage of AVPP, we aim to increase the hit rate of the VT by predicting also the load address of a future instance of the same load instruction and prefetching its value in the VT. We show that AVPP is relatively easy to implement, requiring only 2.5% of the area of a 32KB L1 data cache. We compare our mechanism with five state-of-the-art value prediction techniques, evaluated within the context of load value prediction, in a relatively narrow out-of-order processor. On average, our AVPP predictor achieves 11.2% speedup and 3.7% of energy savings over the baseline processor, outperforming all the state-of-the-art predictors in 16 of the 23 benchmarks we evaluate. We evaluate AVPP implemented together with different prefetching techniques, showing additive performance gains (20% average speedup). In addition, we propose a new taxonomy to classify different value predictor policies regarding predictor update, predictor availability, and in-flight pending updates. We evaluate these policies in detail. Lois Orosa 0001, Rodolfo Azevedo, Onur Mutlu |
ACM Trans. Archit. Code Optim. | 2 |
| 2018 | ADeLe: Rapid Architectural Simulation for Approximate HardwareabstractRecent research has introduced approximate hardware units that produce incorrect outputs deterministically or probabilistically for some small subset of inputs but allow significantly higher throughput or lower power than their errorfree counterparts. The integration, validation, and evaluation of these approximate units in architectures and processors, however, remains challenging. In this paper, we introduce ADeLe, a high-level language for the description, configuration, and integration of approximate hardware units into processors. ADeLe reduces the design effort for approximate hardware by modeling approximations at a high level of abstraction and automatically injecting them into a processor model for architectural simulation. Approximations in ADeLe may modify or completely replace the functional behavior of instructions according to user-defined policies. Instructions may be approximated deterministically or probabilistically (e.g., based on operating voltage and frequency). To allow for controlled testing, approximations may be enabled and disabled from software. Energy is automatically accounted based on customizable models that consider the potential power savings of the approximations that are enabled in the system. ADeLe provides designers with a generic and flexible verification framework, allowing them to easily evaluate the energy-quality trade-offs of their designs in applications. We demonstrate the language and corresponding framework by introducing different approximation techniques into a processor model, on top of which we run selected applications. We demonstrate ADeLe using 6 approximate designs with 4 image processing and 2 floating point applications. Our experiments show how ADeLe may be used to generate approximate CPUs and to evaluate energy-quality trade-offs for different applications with reduced effort. Isaías B. Felzmann, Matheus Martins Susin, Liana Dessandre Duenha, Rodolfo Azevedo, Lucas Francisco Wanner |
SBAC-PAD | 4 |
| 2018 | Exploring Active Learning Approaches to Computer Science ClassesabstractWe present our experience in a Computer Science (CS) introductory course, where three teaching practices were implemented and compared: lectured-based learning, problem-based learning, and peer instruction. We chose Information Systems, a first-term undergraduate course, for this study. It overviews a variety of topics in CS, such as algorithms, data structures and programming logic. We initially conducted interviews with previous instructors, who assisted in the collection of data, requirements, and needs pertaining to both students and instructors. We also carried out a survey among students enrolled in the program, in order to identify suggestions on how the classes could become more dynamic and motivating. In sequence, the experiment was designed to format and evaluate classes in the chosen paradigms. We focused on assessing and analyzing how the students' motivation and learning process were affected, as well as how difficult it was for instructors to prepare classes and how much time they expended in doing so. Results indicate that a paradigm shift from traditional teaching is not only expected by students and instructor; it is well received, and had a positive influence on the students' learning and motivation. We also found, however, that the proposed changes brought on an unwelcome overhead for the instructors, as additional time and effort are required to implement such practices. Ricardo Caceffo, Guilherme Gama, Rodolfo Azevedo |
SIGCSE | 3 |
| 2018 | Dark-Silicon Aware Design Space Exploration
Ricardo Santos 0002, Liana Dessandre Duenha, Matheus Sousa, Luiz Augusto Tedesco, João Carlos Melgarejo, Tony Santos, Rodolfo Azevedo, Edward D. Moreno |
J. Parallel Distributed Comput. | 8 |
| 2017 | Exploiting performance, dynamic power and energy scaling in full-system simulatorsabstractSummary Energy consumption constraints have become a critical issue in Multiprocessor Systems on Chip (MPSoC) designs. Whereas processor performance comes with a high power cost, there is an increasing interest in exploring the trade‐off between power and performance, taking into account the target application domain. Dynamic Voltage and Frequency Scaling (DVFS) techniques adaptively scale frequency or voltage level of CPU allowing it to reach just enough performance to process the system workload while meeting throughput constraints, and thereby, reducing the energy consumption. To explore this wide design space for energy efficiency and performance, hardware and software components, a system‐level simulation infrastructure must provide features to evaluate power savings mechanisms in early stages of the design. This paper presents an extension work of a framework for MPSoCs designs to support DVFS in MPSoCs simulators and evaluates three DVFS mechanisms. Our experiments show that applying DVFS in the system can save power and energy consumption, with negligible loss of performance. Copyright © 2016 John Wiley & Sons, Ltd. Liana Dessandre Duenha, Guilherme A. Madalozzo, Fernando Gehm Moraes, Rodolfo Azevedo |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Temporal frequent value localityabstractFrequent value locality is a type of locality based on the observation that a small set of values is accessed very frequently. Several works have exploited it to construct different architectural schemes, such as memory and cache designs or bus and network optimizations. Although these previous works consider different criteria to establish what is a frequent value and what is not, they assume that these frequent values are constant for the whole program, or they analyze the dynamic properties at a granularity too high to take advantage of them. In this paper, we observe that the frequent values change dynamically during the program execution, presenting what we call temporal frequent value locality, and its potential depends on the granularity of the observation. Furthermore, we instrument and analyze the dynamic patterns of the SPEC2006 benchmarks using two realistic schemes, and we compare them with other classical frequent value locality proposals. The results show that temporal frequent value locality has the potential to be used successfully in architectural optimizations. We also simulate our scheme for main memory bandwidth compression, achieving an increase of 63% (up to 285%) in the effective bandwidth, the double of the best tested state-of-the-art compression algorithm. Lois Orosa 0001, Rodolfo Azevedo |
ASAP | 2 |
| 2016 | On the Dark Silicon Automatic Evaluation on Multicore ProcessorsabstractThe advent of Dark Silicon as result of the limit on Dennard scaling forced modern processor designs to reduce the chip area that can work on maximum clock frequency. This effect reduced the free gains from Moore's law. This work introduces a less conservative dark silicon estimate based on chip components power density and technological process, so that designers could explore architectural resources to mitigate it. We implemented our dark silicon estimation tool on top of MultiExplorer and evaluated on a set of Intel Pentium and AMD K8/10 multicore processors built on transistor technologies from 90nm down to 32nm. Our contributions are twofold: (1) Our experiments have shown dark silicon estimates up to 8.26% of the chip area compared to a baseline 90nm real processor, we also evaluated clock frequency behavior based on Dennard scaling and obtained up to 15.65% dark silicon on chip area. (2) We designed and showed that a dark silicon aware Design Space Exploration (DSE) strategy can minimize chip dark area while increasing performance at design time. Our results on DSE found dark silicon free multicore platforms while providing 3.6 speedup. Tony Santos, Liana Dessandre Duenha, Ricardo Santos 0002, Edward D. Moreno, Rodolfo Azevedo |
SBAC-PAD | 6 |
| 2016 | Developing a Computer Science Concept Inventory for Introductory ProgrammingabstractA Concept Inventory (CI) is a set of multiple choice questions used to reveal student's misconceptions related to some topic. Each available choice (besides the correct choice) is a distractor that is carefully developed to address a specific misunderstanding, a student wrong thought. In computer science introductory programming courses, the development of CIs is still beginning, with many topics requiring further study and analysis. We identify, through analysis of open-ended exams and instructor interviews, introductory programming course misconceptions related to function parameter use and scope, variables, recursion, iteration, structures, pointers and boolean expressions. We categorize these misconceptions and define high-quality distractors founded in words used by students in their responses to exam questions. We discuss the difficulty of assessing introductory programming misconceptions independent of the syntax of a language and we present a detailed discussion of two pilot CIs related to parameters: an open-ended question (to help identify new misunderstandings) and a multiple choice question with suggested distractors that we identified. Ricardo Caceffo, Steven A. Wolfman, Kellogg S. Booth, Rodolfo Azevedo |
SIGCSE | 4 |
| 2016 | MPSoCBench: A benchmark for high-level evaluation of multiprocessor system-on-chip tools and methodologies
Liana Dessandre Duenha, Guilherme A. Madalozzo, Thiago Santiago, Fernando Gehm Moraes, Rodolfo Azevedo |
J. Parallel Distributed Comput. | 5 |
| 2015 | MultiExplorer: A tool set for multicore system-on-chip design explorationabstractThis paper proposes MultiExplorer, a new toolset for MPSoCs modelling, experimentation, and design space exploration, by combining fast high-abstraction simulation and low-level physical estimates (power, area, and timing). The MultiExplorer infrastructure takes a range of high and low-level parameters to improve accuracy in the design of a multiprocessor system on a chip. Our toolset results show a viable alternative to explore multiprocessor scalability (1-64 cores) on affordable simulation times. Rodrigo Devigo, Liana Dessandre Duenha, Rodolfo Azevedo, Ricardo Santos 0002 |
ASAP | 3 |
| 2015 | SHRINK: reducing the ISA complexity via instruction recyclingabstractMicroprocessor manufacturers typically keep old instruction sets in modern processors to ensure backward compatibility with legacy software. The introduction of newer extensions to the ISA increases the design complexity of microprocessor front-ends, exacerbates the consumption of precious on-chip resources (e.g., silicon area and energy), and demands more efforts for hardware verification and debugging. We analyzed several x86 applications and operating systems deployed between 1995 and 2012 and observed that many instructions stop being used over time, and more than 500 instructions were never used in these applications. We also investigate the impact of including these unused instructions in the design of the x86 decoders and propose SHRINK, a mechanism to remove old instructions without breaking backward compatibility with legacy code. SHRINK allows us to remove 40% of the instructions from the x86 ISA and improve the critical path, area, and power consumption of the instruction decoder, respectively, by 23%, 48%, and 49%, on average. Bruno Cardoso Lopes, Rafael Auler, Edson Borin, Rodolfo Azevedo |
ISCA | 5 |
| 2014 | Wear-out analysis of Error Correction Techniques in Phase-Change MemoryabstractPhase-Change Memory (PCM) is new memory technology and a possible replacement for DRAM, whose scaling limitations require new lithography technologies. Despite being promising, PCM has limited endurance (its cells withstand roughly 108bit-flips before failing), which prompted the adoption of Error Correction Techniques (ECTs). However, previous lifetime analyses of ECTs did not consider the difference between the bit-flip frequencies of data and code bits, which may lead to inaccurate wear-out analyses for the ECTs. In this work, we improve the wear-out analysis of PCM by modeling and analyzing the bit-flip probabilities of five ECTs. Our models also enable an accurate estimation of energy consumption and analysis of the endurance-energy trade-off for each ECT. Caio Hoffman, Rodolfo Azevedo, Guido Araujo |
DATE | 3 |
| 2014 | Cloud-based OpenMP Parallelization Using a MapReduce RuntimeabstractHarnessing the flexibility and scaling features of the cloud can open up opportunities to address some relevant research problems in scientific computing. Nevertheless, cloudbased parallel programming models need to address some relevant issues, namely communication overhead, workload balance and fault tolerance. Programming models, which work well in multicore machines (e.g. OpenMP), still do not offer a smooth transition path to the cloud, which could bridge the gap from a local prototype execution to a cloud production run. On the other hand, cloud-based execution models, like MapReduce, are very effective in performing regular fault-tolerant computation on large distributed workloads. In this paper we propose OpenMR, an execution model based on OpenMP semantics and MapReduce, which eases the task of programming parallel applications in the cloud. Specifically, this work addresses the problem of performing loop parallelization in a distributed environment, through the mapping of loop iterations to MapReduce nodes. By doing so, the cloud programming interface becomes the programming language itself, freeing the developer from the task of distributing workload and data, while enabling fault-tolerance and workload balancing. To assess the validity of the proposal, we modified benchmarks from the SPEC OMP2012 and Rodinia suites to fit the proposed model, developed I/O-bound synthetic benchmarks and validated them using Amazon AWS services. We compare the results to the execution of OpenMP in an SMP architecture, and show that OpenMR exhibits good scalability under a simple programming model. Rodolfo Wottrich, Rodolfo Azevedo, Guido Araujo |
SBAC-PAD | 2 |
| 2013 | Zombie memory: extending memory lifetime by reviving dead blocksabstractZombie is an endurance management framework that enables a variety of error correction mechanisms to extend the lifetimes of memories that suffer from bit failures caused by wearout, such as phase-change memory (PCM). Zombie supports both single-level cell (SLC) and multi-level cell (MLC) variants. It extends the lifetime of blocks in working memory pages (primary blocks) by pairing them with spare blocks, i.e., working blocks in pages that have been disabled due to exhaustion of a single block's error correction resources, which would be 'dead' otherwise. Spare blocks adaptively provide error correction resources to primary blocks as failures accumulate over time. This reduces the waste caused by early block failures, making working blocks in discarded pages a useful resource. Even though we use PCM as the target technology, Zombie applies to any memory technology that suffers stuck-at cell failures. Rodolfo Azevedo, John D. Davis, Karin Strauss, Parikshit Gopalan, Mark S. Manasse, Sergey Yekhanin |
ISCA | 1 |
| 2013 | An automatic energy consumption characterization of processors using ArchC
Marcelo Guedes, Rafael Auler, Liana Dessandre Duenha, Edson Borin, Rodolfo Azevedo |
J. Syst. Archit. | 5 |
| 2012 | An ArchC approach for automatic energy consumption characterization of processorsabstractThis paper presents acSynth, an ArchC framework for energy characterization and simulation. Based on Tiwari's Method, a subject processor is characterized in an affordable time and the information is fed into acSynth to bring architecture level power analysis. The framework provides power reports and energy profiling. The experimental results show the characterization flow for the Plasma processor, a MIPS-I HDL description. The acSynth can provide power analysis at 35.1 million instructions per second in simulation with small accuracy diversion and without loss of generality. The system allows the execution of large tests in minutes, which would otherwise take years in a standard HDL methodology. Marcelo Guedes, Rafael Auler, Edson Borin, Rodolfo Azevedo |
RSP | 4 |
| 2012 | Energy-Performance Tradeoffs in Software Transactional MemoryabstractTransactional memory (TM) is a new synchronization mechanism devised to simplify parallel programming, thereby helping programmers to unleash the power of current multicore processors. Although software implementations of TM (STM) have been extensively analyzed in terms of runtime performance, little attention has been paid to an equally important constraint faced by nearly all computer systems: energy consumption. In this work we conduct a comprehensive study of energy and runtime tradeoff sin software transactional memory systems. We characterize the behavior of three state-of-the-art lock-based STM algorithms, along with three different conflict resolution schemes. As a result of this characterization, we propose a DVFS-based technique that can be integrated into the resolution policies so as to improve the energy-delay product (EDP). Experimental results show that our DVFS-enhanced policies are indeed beneficial for applications with high contention levels. Improvements of up to 59% in EDP can be observed in this scenario, with an average EDP reduction of 16% across the STAMP workloads. Alexandro Baldassin, João P. L. de Carvalho, Leonardo A. G. Garcia, Rodolfo Azevedo |
SBAC-PAD | 4 |
| 2012 | A transactional runtime system for the Cell/BE architecture
Alexandro Baldassin, Felipe Goldstein, Rodolfo Azevedo |
J. Parallel Distributed Comput. | 3 |
| 2010 | STM versus lock-based systems: an energy consumption perspectiveabstractThe shift towards multicore processors and the well-known drawbacks imposed by lock-based synchronization have forced researchers to devise new alternatives for building concurrent software, of which transactional memory is a promising one. This work presents a comprehensive study on the energy consumption of a state-of-the-art STM (Software Transactional Memory) implementation using STAMP, a representative set of transactional workloads, comparing it to its lock-based counterpart. Our results show that STM can be up to 22x (~3x on average) more energy-inefficient when compared to locks. This work is a novel step towards a better understanding of the energy behavior of STM systems. Felipe Klein, Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo, Rodolfo Azevedo |
ISLPED | 6 |
| 2009 | SPARC16: A New Compression Approach for the SPARC ArchitectureabstractRISC processors can be used to face the ever increasing demand for performance required by embedded systems. Nevertheless, this solution comes with the cost of poor code density. Alternative encodings for instruction sets, such as MIPS16 and Thumb, represent an effective approach to deal with this drawback. This article proposes to apply a new encoding to the SPARCv8 architecture. Through extensive analysis of a program mix from the Mibench and Mediabench benchmark suites, we suggest a new 16-bit instruction set, easily translated to its 32-bit counterpart during execution time. Using the aforementioned program mix to infer how code could be represented in the proposed 16-bit ISA, compression ratios as low as 56% can be obtained. We also evaluated the cache behavior and showed reductions of 42% on cache misses that can increase performance up to 28% (for patricia program with 2KB cache). Leonardo Luiz Ecco, Bruno Cardoso Lopes, Eduardo C. Xavier, Ricardo Pannain, Paulo Centoducatte, Rodolfo Azevedo |
SBAC-PAD | 6 |
| 2009 | A Multi-Model Engine for High-Level Power Estimation Accuracy OptimizationabstractRegister transfer level (RTL) power macromodeling is a mature research topic with a variety of equation and table-based approaches. Despite its maturity, macromodeling is not yet widely accepted as a de facto industrial standard for power estimation at the RT level. Each approach has many variants depending upon the parameters chosen to capture power variation. Every macromodeling technique has some intrinsic limitation affecting either its performance or its accuracy. Therefore, alternative macromodeling methods can be envisaged as part of a power modeling toolkit from which multiple models for a given component could be exploited so as to reduce the estimation errors resulting from conventional single-model approaches. This paper describes two different approaches for a new multi-model power estimation engine. The first one selects the macromodeling technique that leads to the least estimation error, for a given system component, depending on the properties of its input-vector stream. A proper selection function is built after component characterization and used during estimation. Though simple, this approach has revealed a substantial improvement in estimation accuracy. The second one builds a power estimate function that captures the correlation between individual macromodel estimates and input-stream properties. Experimental results show that our multi-model engine improves the robustness of power analysis with negligible usage overhead. Accuracy becomes seven times better on average, as compared to conventional single-model estimators, while the overall maximum estimation error is divided by 9. Felipe Klein, Roberto Leao, Guido Araujo, Luiz Cláudio Villar dos Santos, Rodolfo Azevedo |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2008 | A Software Transactional Memory System for an Asymmetric Processor ArchitectureabstractDue to the advent of multi-core processors and the consequent need for better concurrent programming abstractions, new synchronization paradigms have emerged. A promising one, known as software transactional memory (STM), aims to use transactions as the key synchronization mechanism to ease program development as well as increase its performance. Many (if not all) of the current STM implementations target homogeneous architectures. In this paper we describe an implementation of an STM system for an asymmetric architecture, the Cell BE. We evaluated our Transactional Software Cache (TSC) mechanism using a well-known micro-benchmark (IntSet) and the Genome application from STAMP. The results show that an STM implementation for the Cell architecture is feasible if the shared-memory programming model is adopted. When compared to a conventional lock-based implementation, the STM version of Genome obtained a performance gain of 84% and 24% with large and small input sets, respectively. Felipe Goldstein, Alexandro Baldassin, Paulo Centoducatte, Rodolfo Azevedo, Leonardo A. G. Garcia |
SBAC-PAD | 4 |
| 2007 | A multi-model power estimation engine for accuracy optimizationabstractRTL power macromodeling is a mature research topic with a variety of equation and table-based approaches. Despite its maturity, macromodeling is not yet widely accepted as an industrial de facto standard for power estimation at the RT level. Each approach has many variants depending upon the parameters chosen to capture power variation. Every macromodeling technique has some intrinsic limitation affecting either its performance or its accuracy. Therefore, alternative macromodeling methods can be envisaged as part of a power modeling toolkit from which the most suitable method for a given component should be automatically selected. Thispaper describes a new multi-model power estimation engine that selects the macromodeling technique leading to the least estimation error for a given system component depending on the properties of its input-vector stream. A proper selection function is built after component characterization and used during estimation. Experimental results show that our multi-model engine improves the robustness of power analysis with negligible usage overhead. Accuracy becomes 3 times better on average, as compared to conventional single-model estimators, while the overall maximum estimation error is divided by 8. Felipe Klein, Guido Araujo, Rodolfo Azevedo, Roberto Leao, Luiz Cláudio Villar dos Santos |
ISLPED | 3 |
| 2006 | 2D-VLIW: An Architecture Based on the Geometry of ComputationabstractThis work proposes a new architecture and execution model called 2D-VLIW. This architecture adopts an execution model based on large pieces of computation running over a matrix of functional units connected by a set of local register spread across the matrix. Experiments using the Mediabench and SPECint00 programs and the Trimaran compiler show performance gains ranging from 5% to 63%, when comparing the proposal to an EPIC architecture with the same number of registers and functional units. It also show that the g72-enc program running on a 2D-VLIW3times3 matrix had a speedup of 1.37 over a 2times2 matrix while the same program over the EPIC processor with 9 functional units had a speedup of 1.12 over an EPIC processor with 4 functional units. For some internal procedures from Mediabench and SPECint programs, the average 2D-VLIW OPC (operations per cycle) was up to 10 times greater than for the equivalent EPIC processor Ricardo Santos 0002, Rodolfo Azevedo, Guido Araujo |
ASAP | 2 |
| 2006 | Exploiting dynamic reconfiguration techniques: the 2D-VLIW approachabstractFast reconfiguration is a mandatory feature for re-configurable computing architectures. Research in this area has been increasingly focusing on new reconfiguration techniques that can sustain the architecture performance and to allow the simultaneous execution, at the same stage, of configuration and computation tasks. In this context, this paper presents a new dynamic reconfiguration technique, based on a configuration cache, that tackles this challenge by configuring and executing operations on functional units during the execution stage. This approach is implemented in a pipelined reconfigurable multiple-issue architecture called 2D-VLIW. Our dynamic reconfiguration technique takes advantage of the 2D-VLIW pipelined execution by starting reconfiguration concurrently to activities like reading operand registers and executing operations. Ricardo Santos 0002, Rodolfo Azevedo, Guido Araujo |
IPDPS | 2 |
| 2005 | Processor Centric Specification and Modelling of MPSoCs
Cristiano C. de Araújo, Edna Barros, Rodolfo Azevedo, Guido Araujo |
FDL | 3 |
| 2004 | Multi-profile based code compressionabstractCode compression has been shown to be an effective technique to reduce code size in memory constrained embedded systems. It has also been used as a way to increase cache hit ratio, thus reducing power consumption and improving performance. This paper proposes an approach to mix static/dynamic instruction profiling in dictionary construction, so as to best exploit trade-offs in compression ratio/performance. Compressed instructions are stored as variable-size indices into fixed-size codewords, eliminating compressed code misalignments. Experimental results, using the Leon (SPARCv8) processor and a program mix from MiBench and Mediabench, show that our approach halves the number of cache accesses and power consumption while produces compression ratios as low as 56%. Eduardo Braulio Wanderley Netto, Rodolfo Azevedo, Paulo Centoducatte, Guido Araujo |
DAC | 2 |
| 2004 | Modeling and Simulating Memory Hierarchies in a Platform-Based Design MethodologyabstractThis paper presents an environment based on SystemC for architecture specification of programmable systems. Making use of the new architecture description language ArchC, able to capture the processor description as well as the memory subsystem configuration, this environment offers support for system-level specification, intended for platform-based design. As a case study, it is presented the memory architecture exploration for a simple image processing application, yet a more robust environment evaluation is performed through the execution of some real-world benchmarks. Pablo Viana, Edna Barros, Sandro Rigo, Rodolfo Azevedo, Guido Araujo |
DATE | 4 |
| 2004 | Optimizations for Compiled Simulation Using Instruction Type InformationabstractThe design of new architectures can be simplified with the use of retargetable instruction set simulation tools, which can validate the design decisions in the design exploration cycle with high flexibility and reduced cost. The growing system complexity makes the traditional approach inefficient for today's architectures. Compiled simulation techniques make use of a priori knowledge to accelerate the simulation, with the highest efficiency achieved by employing static scheduling techniques. This paper presents our approach to the static scheduling compiled simulation technique that is 90% faster than the best published performance results. It also introduces two novel optimization techniques based on instruction type information that further increase the simulation speed by more than 100%. The so-called fast static compiled simulation (FSCS) technique applicability will be demonstrated by the use of the SPARC and MIPS architectures. Marcus Bartholomeu, Rodolfo Azevedo, Sandro Rigo, Guido Araujo |
SBAC-PAD | 2 |
| 2004 | Multi-Profile Instruction Based CompressionabstractCode compression has been used to minimize the memory area requirement of embedded systems. Recently, performance improvement and energy consumption reduction are observed as a by-product of compression. In this paper we propose a novel technique for efficiently exploring the trade-offs involved in code compression. Our multiprofile approach to build dictionaries combines the best features of both static and dynamic program behaviors. The experiments with Mediabench and MiBench suites and the Leon (SPARCv8) processor reveal a compression ratio as low as 71% while performance speed-up reaches 1.5. Eduardo Braulio Wanderley Netto, Rodolfo Azevedo, Paulo Centoducatte, Guido Araujo |
SBAC-PAD | 2 |
| 2004 | ArchC: A SystemC-Based Architecture Description LanguageabstractThis paper presents an architecture description language (ADL) called ArchC, which is an open-source SystemC-based language that is specialized for processor architecture description. Its main goal is to provide enough information, at the right level of abstraction, in order to allow users to explore and verify new architectures, by automatically generating software tools like simulators and co-verification interfaces. ArchC's key features are a storage-based co-verification mechanism that automatically checks the consistency of a refined ArchC model against a reference (functional) description, memory hierarchy modeling capability, the possibility of integration with other SystemC IPs and the automatic generation of high-level SystemC simulators. We have used ArchC to synthesize both functional and cycle-based simulators for the MIPS, Intel 8051 and SPARC V8 processors, as well as functional models of modern architectures like TMS320C62x, XScale and PowerPC. Sandro Rigo, Guido Araujo, Marcus Bartholomeu, Rodolfo Azevedo |
SBAC-PAD | 4 |
| 2003 | Exploring Memory Hierarchy with ArchCabstractWe present the cache configuration exploration of a programmable system, in order to find the best matching between the architecture and a given application. Here, programmable systems composed by processor and memories may be rapidly simulated making use of ArchC, an architecture description language (ADL) based on SystemC. Initially designed to model processor architectures, ArchC was extended to support a more detailed description of the memory subsystem, allowing the design space exploration of the whole programmable system. As an example, it is shown an image processing application, running on a SPARC-V8 processor-based architecture, which had its memory organization adjusted to minimize cache misses. Pablo Viana, Edna Barros, Sandro Rigo, Rodolfo Azevedo, Guido Araujo |
SBAC-PAD | 4 |
| 2001 | Tailoring pipeline bypassing and functional unit mapping to application in clustered VLIW architecturesabstractIn this paper we describe a design exploration methodology for clustered VLIW architectures. The central idea of this work is a set of three techniques aimed at reducing the cost of expensive inter-cluster copy operations. Instruction scheduling is performed using a list-scheduling algorithm that stores operand chains into the same register file. Functional units are assigned to clusters based on the application inter-cluster communication pattern. Finally, a careful insertion of pipeline bypasses is used to increase the number of data-dependencies that can be satisfied by pipeline register operands. Experimental results, using the SPEC95 benchmark and the IMPACT compiler, reveal a substantial reduction in the number of copies between clusters. Marcio Buss, Rodolfo Azevedo, Paulo Centoducatte, Guido Araujo |
CASES | 2 |
| 2000 | Expression-tree-based algorithms for code compression on embedded RISC architecturesabstractReducing program size has become an important goal in the design of modern embedded systems targeted to mass production. This problem has driven efforts aimed at designing processors with shorter instruction formats (e.g., ARM Thumb and MIPS16) or able to execute compressed code (e.g., IBM PowerPC 405), This paper proposes three code compression algorithms for embedded RISC architectures. In all algorithms, the encoded symbols are extracted from program expression trees. The algorithms differ on the granularity of the encoded symbol, which are selected from whole trees, parts of trees, or single instructions. Dictionary-based decompression engines are proposed for each compression algorithm. Experimental results, based on SPEC CINT95 programs running on the MIPS R4000 processor, reveal an average compression ratio of 53.6% (31.5%) if the area of the decompression engine is (not) considered. Guido Araujo, Paulo Centoducatte, Rodolfo Azevedo, Ricardo Pannain |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |