EDBT 2026 Demo / reviewers in the wild / expert
Mateus B. Rutzig
dblp:13/3791 · also Mateus Beck Rutzig
· DBLP profile ↗
26ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-2836-2009ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-aware DVFS-driven workload provisioning in heterogeneous cloud FaaS architecturesabstractAbstract Cloud Warehouses are evolving with diverse computational resources, including CPUs, GPUs, and accelerators, catering to a multitude of tenant applications. While this heterogeneity promises improved performance and energy efficiency, harnessing its full potential poses challenges due to dynamic workload characteristics and variable application demands. To address this, scheduling approaches combined with optimization techniques like Dynamic Voltage and Frequency Scaling (DVFS) are crucial. However, integrating these approaches effectively can be complex, potentially leading to conflicts and diminished benefits. This research proposes two frameworks, EAPECloud and EAPECloud-DVFS, designed for energy-aware collaborative provisioning in heterogeneous CPU-GPU cloud nodes. The first approach reduces energy consumption by selecting and maintaining a static combination of the best scheduler and V-F pair for most workloads. The second approach goes further by dynamically adjusting the V-F pair of each device using DVFS techniques while selecting the optimal scheduler. While the static approach delivers strong results in most cases, the dynamic strategy achieves even greater energy savings, albeit with an additional convergence time to determine the optimal V-F pair. Although each framework has distinct advantages and use cases, our findings demonstrate that both approaches effectively reduce energy consumption in heterogeneous environments, with EAPECloud-DVFS achieving up to a 126.33% performance improvement compared to the Linux CPU Governor, highlighting its efficiency and applicability in real-time systems. Lucas Rister Machado, Gregory de Moraes Rossato, Antonio Carlos Schneider Beck, Michael G. Jordan, Mateus B. Rutzig |
J. Supercomput. | 5 |
| 2024 | Exploiting Virtual Layers and Pruning for FPGA-Based Adaptive Traffic ClassificationabstractTraffic classification is crucial to many network administration tasks, from resource management to QoS monitoring. While DNNs are the state-of-the-art method for classifying traffic, they are computationally intensive, posing challenges to their adoption within the networking infrastructure. One popular alternative is exploiting the FPGAs in the Smart Network Interface Cards (SmartNICs) to speed up the DNN inference processing. In this context, there are two main obstacles. First, modern networks experience high-volume and very volatile traffic flow, making the design of in-network accelerators difficult. Second, the use of FPGA-enabled SmartNICs involves reconfiguration when changing the classification task, which leads to significant time and energy overheads. In this work, we propose two different but complementary solutions to the aforementioned challenges: the use of pruning, which dynamically removes parts of the DNN to speed up its processing at the cost of controlled accuracy drops, therefore adapting the inference processing to the constantly changing traffic; and Hardware Virtual Layers (HWVL), which eliminate the need for FPGA reconfigurations for seamless and almost instantaneous task switching. Both approaches are combined in the Spyke Framework, improving throughput in up to 1.46× and reducing the energy per inference in up to 1.37× compared to a state-of-the-art FPGA accelerator. Julio Costella Vicenzi, Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
DSD | 4 |
| 2023 | Adaptive Inference on Reconfigurable SmartNICs for Traffic Classification
Julio Costella Vicenzi, Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
AINA (2) | 4 |
| 2023 | Pruning and Early-Exit Co-Optimization for CNN Acceleration on FPGAsabstractThe challenge of processing heavy-load ML tasks, particularly CNN-based ones at resource-constrained IoT devices, has encouraged the use of edge servers. The edge offers performance levels higher than the end devices and better latency and security levels than the Cloud. On top of that, the rising complexity of ML applications, the ever-increasing number of connected devices, and the current demands for energy efficiency require optimizing such CNN models. Pruning and early-exit are notable optimizations that have been successfully used to alleviate the computational cost of inference. However, these optimizations have not yet been exploited simultaneously: while pruning is usually applied at design time, which involves retraining the CNN before deployment, early-exit is inherently dynamic. In this work, we propose AdaPEx, a framework that exploits the intrinsic reconfigurable FPGA capabilities so both can be cooperatively employed. AdaPEx first explores the trade-off between pruning and early-exit at design-time, creating a design space never exploited in the state-of-the-art. Then, AdaPEx applies FPGA reconfiguration as a means to enable the combined use of pruning and early-exit dynamically. At runtime, this allows matching the inference processing to the current edge conditions and a user-configurable accuracy threshold. In a smart IoT application, AdaPEx processes up to 1.32× more inferences and improves EDP by up to 2.55× over the state-of-the-art FPGA-based FINN accelerator. Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Jerónimo Castrillón, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2023 | MVSym: Efficient symbiotic exploitation of HLS-kernel multi-versioning for collaborative CPU-FPGA cloud systems
Michael G. Jordan, Bernardo Neuhaus Lignati, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
Integr. | 4 |
| 2023 | Energy-aware fully-adaptive resource provisioning in collaborative CPU-FPGA cloud environments
Michael G. Jordan, Guilherme Korol, Tiago Knorst, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 4 |
| 2022 | AdaFlow: A Framework for Adaptive Dataflow CNN Acceleration on FPGAsabstractTo meet latency and privacy requirements, resource-hungry deep learning applications have been migrating to the Edge, where IoT devices can offload the inference processing to local Edge servers. Since FPGAs have successfully accelerated an increasing number of deep learning applications (especially CNN-based ones), they emerge as an effective alternative for Edge platforms. However, Edge applications may present highly unpredictable workloads, requiring runtime adaptability in the inference processing. Although some works apply model switching on CPU and GPU platforms by exploiting different pruning rates at runtime, so the inference can adapt according to some quality-performance trade-off, FPGA-based accelerators refrain from this approach since they are synthesized to specific CNN models. In this context, this work enables model switching on FPGAs by adding to the well-known FINN accelerator an extra level of adaptability (i.e., flexibility) and support to the dynamic use of pruning via fast model switch on flexible accelerators, at the cost of some extra logic, or via FPGA reconfigurations of fixed accelerators. From that, we developed AdaFlow: a framework that automatically builds, at design time, a library from these new available versions (flexible and fixed, pruned or not) that will be used, at runtime, to dynamically select a given version according to a user-configurable accuracy threshold and current workload conditions. We have evaluated AdaFlow under a smart Edge surveillance application with two CNN models and two datasets, showing that AdaFlow processes, on average, 1.3× more inferences and increases, on average, 1.4× the power efficiency over state-of-the-art statically deployed dataflow accelerators. Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2022 | ConfAx: Exploiting Approximate Computing for Configurable FPGA CNN Acceleration at the EdgeabstractThe number of CNN-based applications executing at the Edge has been considerably increasing. Considering that CNNs are recognized error-resilient and the varied Edge conditions, we exploit hardware-level Approximate Computing to optimize FPGA-based CNN accelerators without any model retraining or other modifications. Given that, we propose ConfAx, a fully configurable multi-target Framework that navigates the accuracy-performance-resource trade-off to deploy different versions of CNN FPGA approximate accelerators. With an Edge case study (video surveillance), we show that ConfAx reduces power (up to $1.65\times)$ and energy (up to $ 1.44\times$) over a state-of-the-art accelerator at minor accuracy penalties (0.88% on average). Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
ISCAS | 3 |
| 2022 | SIS-ASTROS: An Integrated Simulation System for the Artillery Saturation Rocket System (ASTROS)
Cesar Tadeu Pozzer, João Baptista dos Santos Martins, Lisandra M. Fontoura, Luís A. Lima Silva, Mateus B. Rutzig, Raul Ceretta Nunes, Edison Pignaton de Freitas |
SIMULTECH | 5 |
| 2021 | Exploiting HLS-Generated Multi-Version Kernels to Improve CPU-FPGA Cloud SystemsabstractCloud Warehouses have been exploiting CPU-FPGA collaborative execution environments, where multiple clients share the same infrastructure to achieve to maximize resource utilization with the highest possible energy efficiency and scalability. However, the resource provisioning is challenging in these environments, since kernels may be dispatched to both CPU and FPGA concurrently in a highly variant scenario, in terms of available resources and workload characteristics. In this work, we propose MultiVers, a framework that leverages automatic HLS generation to enable further gains in such CPU-FPGA collaborative systems. MultiVers exploits the automatic generation from HLS to build libraries containing multiple versions of each incoming kernel request, greatly enlarging the available design space exploration passive of optimization by the allocation strategies in the cloud provider. Multivers makes both kernel multiversioning and allocation strategy to work symbiotically, allowing fine-tuning in terms of resource usage, performance, energy, or any combination of these parameters. We show the efficiency of MultiVers by using real-world cloud request scenarios with a diversity of benchmarks, achieving average improvements on makespan and energy of up to 4.62x and 19.04x, respectively, over traditional allocation strategies executing non-optimized kernels. Bernardo Neuhaus Lignati, Michael G. Jordan, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
ASP-DAC | 4 |
| 2021 | FAIR: Fully-Adaptive Framework for Improving Resource Provisioning in Collaborative CPU-FPGA Cloud EnvironmentsabstractCloud Warehouses have been exploiting CPU-FPGA collaborative environments to accelerate multi-tenant applications to achieve scalability and maximize resource utilization. However, resource provisioning is challenging in these environments since kernels may be dispatched to CPU and FPGA concurrently in a scenario with highly variant workloads and demands. The provisioning complexity is further aggravated due to diverse CPU and FPGA architectures being used at Cloud Warehouses (e.g., different FPGA/CPU devices between nodes). That means that the resource manager needs to consider the workload to be allocated and the characteristics of the Cloud infrastructure, which can be non-uniform. This paper shows that efficient resource provisioning in CPU-FPGA cloud environments requires different strategies depending on the demand, architecture, and workload. To provide the best use of resources in this complex environment, we propose FAIR, a Fully-Adaptive approach for Improving Resource provisioning in Collaborative CPU-FPGA Cloud. FAIR is end user-transparent and, in contrast to existing approaches, exploits the benefits of multiple provisioning strategies by dynamically selecting the most appropriate depending on the warehouse needs, workload properties, and target architecture. Over a varied set of scenarios, FAIR significantly improves the performance and energy efficiency of the environment compared to the use of fixed single strategies. On average, FAIR provides 32% performance improvements over the use of the best fixed single strategy. Compared to an Oracle that always selects the best energy strategies, FAIR achieves only 3% energy degradation. Michael G. Jordan, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
SBAC-PAD | 3 |
| 2021 | Synergistically Exploiting CNN Pruning and HLS Versioning for Adaptive Inference on Multi-FPGAs at the EdgeabstractFPGAs, because of their energy efficiency, reconfigurability, and easily tunable HLS designs, have been used to accelerate an increasing number of machine learning, especially CNN-based, applications. As a representative example, IoT Edge applications, which require low latency processing of resource-hungry CNNs, offload the inferences from resource-limited IoT end nodes to Edge servers featuring FPGAs. However, the ever-increasing number of end nodes pressures these FPGA-based servers with new performance and adaptability challenges. While some works have exploited CNN optimizations to alleviate inferences’ computation and memory burdens, others have exploited HLS to tune accelerators for statically defined optimization goals. However, these works have not tackled both CNN and HLS optimizations altogether; neither have they provided any adaptability at runtime, where the workload’s characteristics are unpredictable. In this context, we propose a hybrid two-step approach that, first, creates new optimization opportunities at design-time through the automatic training of CNN model variants (obtained via pruning) and the automatic generation of versions of convolutional accelerators (obtained during HLS synthesis); and, second, synergistically exploits these created CNN and HLS optimization opportunities to deliver a fully dynamic Multi-FPGA system that adapts its resources in a fully automatic or user-configurable manner. We implement this two-step approach as the AdaServ Framework and show, through a smart video surveillance Edge application as a case study, that it adapts to the always-changing Edge conditions: AdaServ processes at least 3.37× more inferences (using the automatic approach) and is at least 6.68× more energy-efficient (user-configurable approach) than original convolutional accelerators and CNN Models (VGG-16 and AlexNet). We also show that AdaServ achieves better results than solutions dynamically changing only the CNN model or HLS version, highlighting the importance of exploring both; and that it is always better than the best statically chosen CNN model and HLS version, showing the need for dynamic adaptability. Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | MCEA: A Resource-Aware Multicore CGRA Architecture for the EdgeabstractModern IoT edge devices must address the unpredictability of applications with strict power and temperature constraints. In this scenario, heterogeneous multicore architectures have been driving many solutions due to their high energy efficiency and ability to exploit Task-Level Parallelism. However, while their performance is highly dependent on the quality of the scheduling, their adaptability and generality get restricted when they use fixed-size hardware accelerators. Considering that, this work proposes MCEA, a transparent and power-adaptive multicore reconfigurable architecture. MCEA dynamically adapts the hardware to the workload rather than migrating applications; and predicatively sizes its reconfigurable accelerators without prior knowledge of the applications' behaviors. For that, MCEA uses a synergistic and online profiling system with power gating, achieving performance levels near of homogeneous architectures with fixed and oversized reconfigurable fabric (within 99% on average) while presenting energy efficiency levels similar to heterogeneous architectures statically tuned to a specific workload (within 99% on average). Therefore, MCEA improves Energy-Delay Product in 1.55x and 1.21x when compared to their heterogeneous and homogeneous counterparts, and in 4.72x when compared to a multicore with OoO processors only. We also show that MCEA outperforms a state-of-the-art reconfigurable architecture for the edge under the same power envelope. Guilherme Korol, Michael G. Jordan, Marcelo Brandalero, Michael Hübner 0001, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
FPL | 5 |
| 2019 | Boosting SIMD Benefits through a Run-time and Energy Efficient DLP DetectionabstractData Level Parallelism has been improving performance-energy tradeoff of current processors by coupling SIMD engines, such as Intel AVX and ARM NEON. Special libraries and compilers are used to support DLP execution on such engines. However, timing overhead on hand coding is inevitable since most software developers are not skilled to extract DLP using unfamiliar libraries. In addition, DLP detection through compiler, besides breaking software compatibility, is limited to static code analysis, which compromises performance gains. In this work, we propose a runtime DLP detection named as Dynamic SIMD Assembler, which transparently identifies vectorizable code regions to execute in the ARM NEON engine. Due to its dynamic fashion, DSA keeps software compatibility and avoids timing overhead on software developing process. Results have shown that DSA outperforms ARM NEON auto-vectorization compiler by 32% since it covers wider vectorized regions, such as Dynamic Range, Sentinel and Conditional Loops. In addition, DSA outperforms hand-vectorized code using ARM library by 26% reducing 45% of energy consumption with no penalties over software development time. Michael G. Jordan, Tiago Knorst, Julio Costella Vicenzi, Mateus B. Rutzig |
DATE | 4 |
| 2018 | Semi-Autonomous Navigation for Virtual Tactical Simulations in the Military Domain
Juliana Rubenich Brondani, Luís A. Lima Silva, Mateus B. Rutzig, Cesar Tadeu Pozzer, Raul Ceretta Nunes, João Baptista dos Santos Martins, Edison Pignaton de Freitas |
SIMULTECH | 3 |
| 2017 | Improving EDP in multi-core embedded systems through multidimensional frequency scalingabstractEnergy saving management in multi-core embedded environments has been a challenge for designers. To achieve energy efficiency, most studies consider dynamic frequency scaling on one hardware component only, such as processor or memory - which will most likely also affect performance. This work proposes the use of frequency scaling considering the three most important hardware components altogether: processors, L2 cache, and RAM; seeking for the best set of frequencies for each one of them to improve the Energy-Delay Product (EDP), depending on the application's behavior. Therefore, this work addresses multidimensional frequency scaling for multi-core embedded systems. By evaluating different frequency levels, we show that the EDP can be improved in up to 46.4% when compared to the standard way that the frequencies are configured. Wagner dos Santos Marques, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Mateus B. Rutzig, Fábio D. Rossi |
ISCAS | 5 |
| 2017 | A framework to automatically generate heterogeneous organization reconfigurable multiprocessingabstractHeterogeneous MPSoCs are vastly used in current embedded systems but they are highly dependent on special compilers. Dynamic reconfigurable systems are an alternative to overcome such drawback due to their adaptability. However, when such architectures are considered, one must concern about which hardware blocks should be heterogeneous and their degree of heterogeneity. In this work, we propose a framework that automatically generates heterogeneous reconfigurable multiprocessors that exploit the ideal ILP of parallel applications to improve performance/watt. Our generated system achieves, on average, 32% of performance improvements with 33% of energy savings over its manual generated counterpart, with equivalent chip area. Josimar Sfreddo, Rafael Fao de Moura, Michael G. Jordan, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck, Mateus B. Rutzig |
ISCAS | 6 |
| 2016 | A reconfigurable heterogeneous multicore with a homogeneous ISA
Jeckson Dellagostin Souza, Luigi Carro, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2013 | A transparent and energy aware reconfigurable multiprocessor platform for simultaneous ILP and TLP exploitationabstractAs the number of embedded applications increases, companies are launching new platforms within short periods of time to efficiently execute software with the lowest possible energy consumption. However, for each new platform deployment, new tool chains, with additional libraries, debuggers and compilers must come along, breaking binary compatibility. This strategy implies in high hardware and software redesign costs. In this scenario, we propose the exploitation of Custom Reconfigurable Arrays for Multiprocessor Systems (CReAMS). CReAMS is composed of multiple adaptive reconfigurable processors that simultaneously exploit Instruction and Thread Level Parallelism. It works in a transparent fashion, so binary compatibility is maintained, with no need to change the software development process or environment. We also show that CReAMS delivers higher performance per watt in comparison to a 4-issue Superscalar processor, when the same power budget is considered for both designs. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
DATE | 1 |
| 2013 | A run-time adaptive multiprocessor systemabstractBecause of the continuous increase in the number and complexity of embedded applications, new platforms have been launched within shorter periods of time to fulfill their performance requirements with the lowest energy consumption possible. However, for each new platform deployment, new tool chains, with additional libraries, debuggers and compilers must come along, breaking binary compatibility. This strategy implies in high hardware and software redesign costs. In this scenario, we propose the exploitation of custom reconfigurable arrays for multiprocessor systems. The proposed approach is composed of multiple adaptive reconfigurable processors that simultaneously exploit Instruction and Thread Level Parallelism. It works in a transparent fashion, so binary compatibility is maintained, with no need to change the software development process or environment. Results show that our proposal delivers higher performance per watt in comparison to a 4-issue Superscalar processor, when the same power budget is considered. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
ISCAS | 1 |
| 2013 | Towards a multiple-ISA embedded system
Jair Fajardo Junior, Mateus B. Rutzig, Luigi Carro, Antonio Carlos Schneider Beck |
J. Syst. Archit. | 2 |
| 2009 | A low cost and adaptable routing network for reconfigurable systemsabstractNowadays, scalability, parallelism and fault-tolerance are key features to take advantage of last silicon technology advances, and that is why reconfigurable architectures are in the spotlight. However, one of the major problems in designing reconfigurable and parallel processing elements concerns the design of a cost-effective interconnection network. This way, considering that Multistage Interconnection Network (MIN) has been successfully used in several computer system levels and applications in the past, in this work we propose the use of a MIN, at the word level, on a coarse-grained reconfigurable architecture. More precisely, this work presents a novel parallel self-placement and routing mechanism for MIN on the circuit-switching mode. We take into account one-to-one as well as multicast (one-to-many) permutations. Our approach is scalable and it is targeted to be used in run-time environments where dynamic routing among functional units is required. In addition, our algorithm is embedded in the switch structure, and it is independent of the interstage interconnection pattern. Our approach can handle blocking and non-blocking networks, symmetrical or asymmetrical topologies. As case study, we use the proposed technique in a dynamic reconfigurable system, showing a major area reduction of 30% without performance overhead. Ricardo S. Ferreira 0001, Marcone Laure, Antonio Carlos Schneider Beck, Thiago Lo, Mateus B. Rutzig, Luigi Carro |
IPDPS | 5 |
| 2008 | Transparent Reconfigurable Acceleration for Heterogeneous Embedded ApplicationsabstractEmbedded systems are becoming increasingly complex. Besides the additional processing capabilities, they are characterized by high diversity of computational models coexisting in a single device. Although reconfigurable architectures have already shown to be a potential solution for such systems, they just present significant speedups of very specific dataflow oriented kernels. Furthermore, reconfigurable fabric is still withheld by the need of special tools and compilers, clearly not sustaining backward software compatibility. In this paper, we propose a new technique to optimize both dataflow and control-flow oriented code in a totally transparent process, without the need of any modification in the source or binary codes. For that, we have developed a Binary Translation algorithm implemented in hardware, which works in parallel to a MIPS processor. The proposed mechanism is responsible for transforming sequences of instructions at runtime to be executed on a dynamic coarse-grain reconfigurable array, supporting speculative execution. Executing the MIBench suite, we show performance improvements of up to 2.5 times, while reducing 1.7 times the required energy, using trivial hardware resources. Antonio Carlos Schneider Beck, Mateus B. Rutzig, Georgi Gaydadjiev, Luigi Carro |
DATE | 2 |
| 2008 | Reducing interconnection cost in coarse-grained dynamic computing through multistage networkabstractCoarse-grained reconfigurable architectures appear as a scalable solution to embedded system design, with a reduced reconfiguration time, memory footprint, as well as placement and routing complexity. To ensure high performance, data must be efficiently delivered to the reconfigurable matrix. For that, several architectures propose the use of fully interconnected local networks, as crossbar or large multiplexers. However, these interconnections are very area consuming. Therefore, in order to reduce the interconnection complexity without losing performance, this work proposes to use Multistage Interconnection Networks. As a case study, we have implemented the proposed approach in a tightly coupled reconfigurable array, which works together with a MIPS processor. Simulation results over the Mibench Benchmark set show savings of up to 26% of the total area, with a decrease of only 1% on the average performance. Ricardo S. Ferreira 0001, Marcone Laure, Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
FPL | 3 |
| 2008 | Balancing reconfigurable data path resources according to application requirementsabstractProcessor architectures are changing mainly due to the excessive power dissipation and the future break of Moore's law. Thus, new alternatives are necessary to sustain the performance increase of the processors, while still allowing low energy computations. Reconfigurable systems are strongly emerging as one of these solutions. However, because they are very area consuming and deal with a large number of applications with diverse behaviors, new tools must be developed to automatically handle this new problem. This way, in this work we present a tool aimed to balance the reconfigurable area occupied with the performance required by a given application, calculating the exact size and shape of a reconfigurable data path. Using as case study a tightly coupled reconfigurable array and the Mibench Benchmark set, we show that the solution found by the proposed tool saves four times area in comparison with the non-optimized version of the reconfigurable logic, with a decrease of only 5.8% on average of its original performance. This way, we open new applications for reconfigurable devices as low cost accelerators. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
IPDPS | 1 |
| 2008 | Binary translation process to optimize nanowire arrays usageabstractThe last years of semiconductor research have allowed the construction of nanowires. These atomic structures promise to have more than one order of magnitude higher density when compared to 22 nm CMOS, with less power dissipation, but unfortunately with much lower switching speed. Although there are several works that deal with the problems related to specific fabrication issues, the best use of these structures from a design perspective is still an open field of research. The design solution should take into account available parallelism (to cope with these slower than CMOS devices) and reliability. Moreover, one has to tackle the software compatibility problem, in the sense that nanowire circuits are being considered as accelerators, and not a CMOS replacement. In this paper we propose the use of a nanowire array together with a binary translation mechanism that allows the coupling of the nanowire array to a regular microprocessor, and we show how can one expect high performance and low dissipation, while still considering the intrinsic reliability issue of nanowires. Eduardo Luis Rhod, Mateus B. Rutzig, Luigi Carro |
ISCAS | 2 |