EDBT 2026 Demo / reviewers in the wild / expert
César A. M. Marcon
dblp:13/6510 · also César Augusto Missio Marcon
· DBLP profile ↗
50ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-7811-7896ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cognitive Banking: An Adaptive and Decentralized Decision-Making Model Based on Digital Assets for Multi-Agent Systems
Edilson Filho, César A. M. Marcon, Laurent Vercouter, Jarbas Silveira, Walter Abrahão dos Santos |
COMPSAC | 2 |
| 2026 | DEMC: A dynamic multi-ECC memory controller with per-block adaptation
Marco P. Stefani, Felipe G. A. e Silva, Jarbas Silveira, Giovanna Borba, Luthero Vargas, Edson I. Moreno, César A. M. Marcon |
Integr. | 7 |
| 2025 | Accelerating Machine Learning using RISC-V Vector Extension in a Manycore PlatformabstractThis work addresses the acceleration of convolutional neural network (CNN) inference in manycore architectures using coarse and fine-grain parallelism. Prior approaches focus on dedicated accelerators or modified NoCs, limiting flexibility. This work proposes integrating a RISC-V processor extended with the vector extensions (RVV) as general-purpose processing elements in a NoC-based manycore. The implementation applies depthwise convolution mapped across PEs and uses autovectorization provided by the compiler. Experiments on a $4 \times 4$ manycore running the first AlexNet layer achieved up to 5.70x speedup and reduced execution cycles by 82.45% compared to a scalar single-core baseline. Willian Analdo Nunes, Antônio Vinicius Corrêa Dos Santos, César A. M. Marcon, Fernando Gehm Moraes |
VLSI-SoC | 3 |
| 2025 | Conjunctive Merge Instruction to Accelerate Sparse Matrix - Dense Vector MultiplicationabstractSparse linear algebra is essential in many domains due to reduced computation and efficient memory usage. However, the irregularity of sparse data poses challenges for conventional software and hardware. While specialized accelerators offer performance gains, they lack general-purpose flexibility and rely on processor communication, creating bottlenecks. This work addresses these issues by proposing a tiling strategy to improve vector register usage and extending the RISC-V Vector (RVV) ISA with a custom merge instruction. Experiments using the gem5 simulator show that the tiled vector version achieved speedups of up to $1.30 \times(95 \%)$ sparsity) and $1.72 \times(65 \%)$. In contrast, the version with merge instructions reached up to $1.81 \times$ and $6.04 \times$, respectively, over a baseline implementation. Manuel Osterno, César A. M. Marcon, Jarbas Silveira, Fernando Gehm Moraes, Jardel Silveira |
VLSI-SoC | 2 |
| 2025 | Securet3d: An Adaptive, Secure, and Fault-Tolerant Aware Routing Algorithm for Vertically-Partially Connected 3D-NoCabstractMultiprocessor systems-on-chip (MPSoCs) based on 3-D networks-on-chip (3D-NoCs) are crucial architectures for robust parallel computing, efficiently sharing resources across complex applications. To ensure the secure operation of these systems, it is essential to implement adaptive, fault-tolerant mechanisms capable of protecting sensitive data. This work proposes the Securet3d routing algorithm, which establishes secure data paths in fault-tolerant 3D-NoCs. Our approach enhances the Reflect3d algorithm by introducing a detailed scheme for mapping secure paths and improving the system’s ability to withstand faults. To validate its effectiveness, we compare Securet3d with three other fault-tolerant routing algorithms for vertically-partially connected 3D-NoCs. All algorithms were implemented in SystemVerilog and evaluated through simulation using ModelSim and hardware synthesis with Cadence’s Genus tool. Experimental results show that Securet3d reduces latency and enhances cost-effectiveness compared with other approaches. When implemented with a 28-nm technology library, Securet3d demonstrates minimal area and energy overhead, indicating scalability and efficiency. Under denial-of-service (DoS) attacks, Securet3d maintains basically unaltered average packet latencies on 70, 90, and 29 clock cycles for uniform random, bit-complement, and shuffle traffic, significantly lower than those of other algorithms without including security mechanisms (5763, 4632, and 3712 clock cycles in average, respectively). These results highlight the superior security, scalability, and adaptability of Securet3d for complex communication systems. Alexandre Almeida da Silva, Lucas Nogueira, Alexandre A. P. Coelho, Jarbas Silveira, César A. M. Marcon |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Optimizing the Cost of Maintenance Scheduling for Railway Lines Using a Hybrid Evolutionary Algorithm
João Pedro Augusto Costa, Omar Andrés Carmona Cortes, César A. M. Marcon |
HIS (4) | 3 |
| 2023 | A Triple Burst Error Correction Based on Region Selection CodeabstractThe evolution of microelectronics boosts more scalable and complex circuit designs, providing high processing speed and greater storage capacity. However, reliability issues have grown significantly as electronic devices scale down, increasing the fault rate, mainly in critical applications exposed to radiation. Memories are sensitive to charged particles, which can corrupt data due to the transient effects. Error correction codes (ECCs) are highly applied to mitigate data failures, increasing memory reliability. The matrix region selection code (MRSC) is an ECC designed to correct a high rate of adjacent errors in memory but less effectively for nonadjacent errors. However, MRSC has a 2-D structure that makes it challenging to implement in memory where one address is accessed at a time. This article introduces the triple burst error correction based on region selection code (TBEC-RSC), an ECC that uses MRSC concepts, converting the MRSC format to a 1-D structure. TBEC-RSC was implemented and evaluated in a 16-bit data version; however, the code is easily extensible to the higher base-2 data words (e.g., 64 bits). Experimental results showed that TBEC-RSC corrects 100% of triple burst errors and more than 40% of 8-bit burst errors. Felipe G. A. e Silva, Alan Cadore Pinheiro, Jarbas Silveira, César A. M. Marcon |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | Fast Transform Decision Scheme for VVC Intra-Frame Prediction Using Decision TreesabstractThis paper presents a fast transform decision scheme for Versatile Video Coding (VVC) intra-frame prediction using decision trees. VVC introduces several novel coding tools to improve the coding efficiency of the intra-frame prediction at the cost of a high computational effort, including a new transform coding process using Multiple Transform Selection (MTS) for primary transform and Low-Frequency Non-Separable Transform (LFNST) for secondary transform. We developed an efficient complexity reduction scheme composed of two solutions based on decision tree classifiers to avoid the MTS and LFNST evaluations in the costly Rate-Distortion optimization (RDO) process. Experimental results showed that the proposed scheme provides 11% of encoding timesaving with a negligible impact on the coding efficiency. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ISCAS | 3 |
| 2022 | OPCoSA: an Optimized Product Code for space applications
David C. C. Freitas, Jarbas Silveira, César A. M. Marcon, Lirida A. B. Naviner, João Cesar M. Mota |
Integr. | 3 |
| 2022 | Configurable Fast Block Partitioning for VVC Intra Coding Using Light Gradient Boosting MachineabstractThis article presents a configurable fast block partitioning decision for Versatile Video Coding (VVC) intra-frame prediction using Light Gradient Boosting Machine (LGBM). VVC further improves the coding efficiency by introducing a Quadtree with nested Multi-Type Tree (QTMT), enabling five split types allowing square and rectangular Coding Unit (CU) sizes. However, this improvement in the coding efficiency comes at the cost of a high computational burden since several combinations of block sizes and prediction modes are evaluated through the costly Rate-Distortion Optimization (RDO) process. In this article, we propose a partitioning decision using LGBM classifiers to avoid the exhaustive RDO process and skip the evaluation of split types that are unlikely to be chosen as the best one. For this purpose, five classifiers (one for each split type) were offline trained with an efficient training process and using effective features of texture, coding, and context information. The proposed solution is highly configurable and can provide several operation points with different tradeoffs between timesaving and coding efficiency, according to the application requirements. Considering five operation points, the configurable solution can reduce the encoding time from 35.22% to 61.34%, with coding efficiency losses from 0.46% to 2.43%. Compared to the state-of-the-art, our solution is able to outperform the related works in terms of combined rate-distortion and timesaving. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Analysis of VVC Intra Prediction Block Partitioning StructureabstractThis paper presents an encoding time and encoding efficiency analysis of the Quadtree with nested Multi-type Tree (QTMT) structure in the Versatile Video Coding (VVC) intra-frame prediction. The QTMT structure enables VVC to improve the compression performance compared to its predecessor standard at the cost of a higher encoding complexity. The intra-frame prediction time raised about 26 times compared to the HEVC reference software, and most of this time is related to the new block partitioning structure. Thus, this paper provides a detailed description of the VVC block partitioning structure and an in-depth analysis of the QTMT structure regarding coding time and coding efficiency. Based on the presented analyses, this paper can guide outcoming works focusing on the block partitioning of the VVC intra-frame prediction. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
VCIP | 3 |
| 2021 | Learning-Based Complexity Reduction Scheme for VVC Intra-Frame PredictionabstractThis paper presents a learning-based complexity reduction scheme for Versatile Video Coding (VVC) intra-frame prediction. VVC introduces several novel coding tools to improve the coding efficiency of the intra-frame prediction at the cost of a high computational effort. Thus, we developed an efficient complexity reduction scheme composed of three solutions based on machine learning and statistical analysis to reduce the number of intra prediction modes evaluated in the costly Rate-Distortion Optimization (RDO) process. Experimental results demonstrated that the proposed solution provides 18.32% encoding timesaving with a negligible impact on the coding efficiency. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
VCIP | 3 |
| 2021 | Using curved angular intra-frame prediction to improve video coding efficiency
Ramon Fernandes, Gustavo Sanchez, Rodrigo Cataldo, Luciano Volcan Agostini, César A. M. Marcon |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Performance analysis of VVC intra coding
Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | LPC: An Error Correction Code for Mitigating Faults in 3D MemoriesabstractThe radiation sensitivity of memory cells increases dramatically as CMOS manufacture technology scales down; therefore, the reliability of memories has become a challenge. 3D technology has gained attention for having several advantages compared to the 2D counterpart, such as high integration density, high performance, low power, and high communication speed. Although several studies are targeting 3D memories, the effects on reliability using this technology have received little attention. This work introduces Line Product Code (LPC), a modified product code-based Error Correction Code (ECC) that uses both Hamming and parity in both rows and columns to implement reliable 3D memories. We implemented two lightweight LPC-based decoding algorithms in interleaved (LPCa-I) and non-interleaved (LPCa) versions, which allowed us to analyze LPC through a set of simulation cases that considers four severity levels of error incidence. The experimental results showed the effectiveness of the LPC-based algorithms, reaching correction rates of up 2.3 times higher compared to other Hamming-based algorithms. David C. C. Freitas, David F. M. Mota, César A. M. Marcon, Jarbas Silveira, João Cesar M. Mota |
IEEE Trans. Computers | 3 |
| 2021 | Subutai: Speeding Up Legacy Parallel Applications Through Data SynchronizationabstractThe decrease of the performance gain dictated by Moore's Law boosted the development of manycore architectures to replace single-core architectures. These new architectures must employ parallel applications and distribute its workload over a multitude of cores to reach the desired performance. Parallel applications are harder to develop than sequential ones since the developer must guarantee data integrity using synchronization primitives. While multiple novel solutions have been proposed to speed up parallel applications through handling one type of data synchronization primitive, exceptionally few works support multiple types of synchronization primitives and legacy code. This article proposes Subutai, a hardware/software co-design solution for accelerating multiple synchronization primitives without modifying the application source code. By providing a new user library, while retaining an existing synchronization API, legacy and novel applications can benefit from our solution. Our experimental evaluation, which provides a POSIX Threads implementation, demonstrates Subutai speeds up to 2.71× and 4.61× the execution of single- and multiple-application executions, respectively. Rodrigo Cataldo, Ramon Fernandes, Kevin J. M. Martin, Jarbas Silveira, Gustavo Sanchez, Martha Johanna Sepúlveda, César A. M. Marcon, Jean-Philippe Diguet |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2020 | Complexity Analysis Of VVC Intra CodingabstractVersatile Video Coding (VVC) is the next-generation of video coding standards, which was developed to double the coding efficiency over its predecessor High-Efficiency Video Coding (HEVC). Several new coding tools have been investigated and adopted in the VVC Test Model (VTM), whose current version can improve the intra coding efficiency by 24% at the cost of a much higher coding complexity than the HEVC Test Model (HM). Thus, this paper provides a detailed VVC intra coding complexity analysis, which can support upcoming works for finding the most timeconsuming tool that could be simplified to achieve a real-time encoder design. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ICIP | 3 |
| 2020 | Fast Partitioning Decision Scheme for Versatile Video Coding Intra-Frame PredictionabstractThis paper presents a fast partitioning decision scheme for Versatile Video Coding (VVC) intra prediction. VVC adopts the tree coding block structure named Quadtree with nested Multi-Type (QTMT), which significantly improves the coding efficiency at the cost of a high computational effort, limiting its adoption for real applications. Thus, we developed a scheme that encompasses two strategies exploring the features of the current block and the encoding context through the selected intra prediction mode to skip unnecessary evaluations of binary and ternary partitions. Experimental results, obtained with VVC Test Model 5.0 and considering the Common Test Conditions (CTC) under All-Intra (AI) encoder configuration, demonstrated that the proposed scheme achieves 31.41% of coding time saving, on average, with negligible coding efficiency loss. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ISCAS | 3 |
| 2020 | Open-Source NoC-Based Many-Core for Evaluating Hardware Trojan Detection MethodsabstractIn many-cores based on Network-on-Chip (NoC), several applications execute simultaneously, sharing computation, communication and memory resources. This resource sharing leads to security and trust problems. Hardware Trojans (HTs) may steal sensitive information, degrade system performance, and in extreme cases, induce physical damages. Methods available in the literature to prevent attacks include firewalls, denial-of-service detection, dedicated routing algorithms, cryptography, task migration, and secure zones. The goal of this paper is to add an HT in an NoC, able to execute three types of attacks: packet duplication, block applications, and misrouting. The paper qualitatively evaluates the attacks' effect against methods available in the literature, and its effects showed in an NoC-based many-core. The resulting system is an open-source NoC-based many-core for researchers to evaluate new methods against HT attacks. Iaçanã I. Weber, Geaninne Marchezan, Luciano L. Caimi, César A. M. Marcon, Fernando Gehm Moraes |
ISCAS | 4 |
| 2020 | PCoSA: A product error correction code for use in memory devices targeting space applications
David C. C. Freitas, David F. M. Mota, Roger C. Goerl, César A. M. Marcon, Fabian Vargas 0001, Jarbas Silveira, João Cesar M. Mota |
Integr. | 4 |
| 2020 | Fast 3D-HEVC Depth Map Encoding Using Machine LearningabstractThis paper presents a fast depth map encoding for 3D-High Efficiency Video Coding (3D-HEVC) based on static decision trees. We used data mining and machine learning to correlate the encoder context attributes, building the static decision trees. Each decision tree defines that a depth map Coding Unit (CU) must be or not be split into smaller blocks, considering the encoding context through the evaluation of the encoder attributes. Specialized decision trees for I-frames, P-frames and B-frames define the partitioning of 64 × 64, 32 × 32, and 16 × 16 CUs. We trained the decision trees using data extracted from the 3D-HEVC Test Model considering all-intra and random-access configurations, and we evaluated the proposed approach considering the common test conditions. The experimental results demonstrated that this approach can halve the 3D-HEVC encoder computational effort with less than 0.24% of BD-rate increase on the average for all-intra configuration. When running on random-access configuration, our solution is able to reduce up to 58% the complete 3D-HEVC encoder computational effort with a BD-rate drop of only 0.13%. These results surpass all related works regarding computational effort reduction and BD-rate. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Using Smart Routing for Secure and Dependable NoC-Based MPSoCsabstractThe Internet-of-Things (IoT) boosted the building of computational systems that share computation, communication and storage resources for uncountable types of applications. MultiProcessor System-on-Chip (MPSoC) is a fundamental component of such systems offering large parallelism degree in an ocean of processors and memories connected through one or more Network-on-Chips (NoCs). Therefore, a massive quantity of sensitive information of several applications can share computation and communication resources of the MPSoCs demanding security mechanisms and policies. Besides, the advances of CMOS technologies increases the quantity of static and dynamic faults, requiring a dependable and resilient target architecture, which can be partially fulfilled by an effective and efficient NoC design. This work addresses fault tolerance and security at NoC level with SDR, a routing algorithm that includes the concept of security zones in the MPSoC while providing support for dependable routing avoiding faulty links. The proposed routing algorithm prioritizes communication paths deemed secure in 2D mesh NoCs with deadlock freedom. Experimental results employing realistic workload scenarios based on the NASA Numeric Aerodynamic Simulation (NAS) Parallel Benchmark (NPB) and a fault model for 65nm and 22nm CMOS fabrication technologies demonstrates the scalability, security, and dependability of SDR. Ramon Fernandes, César A. M. Marcon, Rodrigo Cataldo, Martha Johanna Sepúlveda |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | TITAN: Tile Timing-Aware Balancing Algorithm for Speeding Up the 3D-HEVC Intra CodingabstractThis paper presents the Tile Timing-Aware balancing algorithm (TITAN) for speeding up the 3D-High Efficiency Video Coding (3D-HEVC) video encoding. The design of the TITAN algorithm is based on the premise that the encoding time of tiles partitioning of neighbor frames are similar. Therefore, based on the encoding time of the last-encoded frame, TITAN controls the tiles boundaries aiming to maximize the balance among tiles. Our software evaluation with TITAN implemented in 3D-HEVC Test Model 16.0 achieved an average of 6.2% higher speedup compared to the uniform-sized tiles. This is the first work in the literature proposing to speed up the 3D-HEVC video encoding by balancing the tiles workload. Mário Saldanha, Gustavo Sanchez, Bruno Zatt, César A. M. Marcon, Luciano Volcan Agostini |
ISCAS | 4 |
| 2019 | Performance Analysis of Depth Intra-Coding in 3D-HEVCabstractThe depth maps intra-frame prediction of 3D High Efficiency Video Coding (3D-HEVC) inherits all texture encoding techniques provided by HEVC and provides new coding tools for depth map predictions. These tools comprise algorithms, such as bipartition modes, intra-picture skip, and DC-only. This paper details these tools and shows how they work together with the original HEVC algorithms in the depth map intra-frame prediction for allowing high-efficiency encoding. Besides, this paper analyzes the encoding time and the encoding mode distribution of the intra-frame prediction tools over different quantization scenarios. We aim to provide support for upcoming works on depth map encoding, including complexity reduction and control, real-time embedded systems implementations, and even the development of improved tools to encode depth maps. Gustavo Sanchez, Jarbas Silveira, Luciano Volcan Agostini, César A. M. Marcon |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applicationsabstractParallel applications are essential for efficiently using the computational power of a Multiprocessor System-on-Chip (MPSoC). Unfortunately, these applications do not scale effortlessly with the number of cores because of synchronization operations that take away valuable computational time and restrict the parallelization gains. Moreover, synchronization is also a bottleneck due to sequential access to shared memory. We address this issue and introduce "Subutai", a hardware/software (HW/SW) architecture designed to distribute essential synchronization mechanisms over the Network-on-Chip (NoC). It includes Network Interfaces (NIs), drivers and a custom library of a NoC-based MPSoC architecture that speeds up the essential synchronization primitives of any legacy parallel application. Besides, we provide a fast simulation tool for parallel applications and a HW architecture of the NI. Experimental results with PARSEC benchmark show an average application speedup of 2.05 compared to the same architecture running legacy SW solutions for 36% overhead of HW architecture. Rodrigo Cataldo, Ramon Fernandes, Kevin J. M. Martin, Martha Johanna Sepúlveda, Altamiro Amadeu Susin, César A. M. Marcon, Jean-Philippe Diguet |
DAC | 6 |
| 2018 | Fast 3D-Hevc Depth Maps Intra-Frame Prediction Using Data MiningabstractThis paper presents a fast 3D-High Efficiency Video Coding (3D-HEVC) depth maps intra-frame prediction based on static Coding Unit (CU) splitting decisions trees. This coding approach uses data mining to extract the correlation among the encoder context attributes and to define a split decision tree for each CU level of the depth maps encoding. The decision trees were trained using the information extracted from 3D-HEVC Test Model (3D-HTM) and using the Common Test Conditions (CTC). Each decision tree defines if the current CU must be split into smaller sizes, considering the encoding context through the evaluation of some current encoder attributes. The proposed solution reaches a complexity reduction of 59.0% for depth maps coding with a negligible impact of 0.18% in the encoding efficiency of synthesized views. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ICASSP | 3 |
| 2018 | DCDM-Intra: Dynamically Configurable 3D-HEVC Depth Maps Intra-Frame Prediction AlgorithmabstractThis work proposes the Dynamically Configurable 3D-HEVC Depth Maps Intra-Frame Prediction (DCDM-Intra), which explores the fact that Rough Mode Decision (RMD) was inherited from texture coding without taking advantages of the depth maps simplicity. DCDM-Intra classifies the HEVC Intra-frame Prediction Modes (IPMs) according to their BD-rate impact when encoding depth maps. Then, the application or the user can dynamically define the number of IPMs supported by the depth maps intra-prediction, according to the system status or application requirements, removing the IPMs that less affect the encoding efficiency. DCDM-Intra allows 35 distinct operation points with different encoding efficiency impacts. Experimental results demonstrate that the exclusion of 20% of IPMs causes a BD-rate increase of 0.03%, and the removal of almost 51% of IPMs rises 0.11% the BD-rate. Gustavo Sanchez, Ramon Fernandes, Luciano Volcan Agostini, César A. M. Marcon |
ICIP | 4 |
| 2018 | High Efficient Architecture for 3D-HEVC DMM-1 Decoder Targeting 1080p VideosabstractThis paper presents an efficient hardware design for the Depth Modeling Mode 1 (DMM-1) decoder of the 3D-High Efficiency Video Coding (3D-HEVC). The designed architecture uses a lossless wedgelet memory compression technique to reduce the used memory, and a well-balanced parallelism level to allow the desired throughput at the minimum possible power consumption and area usage. The architecture was synthesized for the 65nm ST standard cells technology, using 4,047 gates and consuming 0.95mW. It is capable of processing 1080p videos at 30 frames per second, decoding all allowed block sizes and wedgelets patterns. Besides, the proposed architecture saves 10.7% of area and 29.1% of power when compared with a version without memory compression. At the best of the author's knowledge, this is the first work with a dedicated hardware design targeting the DMM-1 decoder. Gustavo Sanchez, Luciano Volcan Agostini, César A. M. Marcon |
ISCAS | 3 |
| 2018 | An Extensible Code for Correcting Multiple Cell Upset in Memory Arrays
Felipe G. A. e Silva, Jardel Silveira, Jarbas Silveira, César A. M. Marcon, Fabian Vargas 0001, Otávio Alcântara de Lima Júnior |
J. Electron. Test. | 4 |
| 2018 | A reduced computational effort mode-level scheme for 3D-HEVC depth maps intra-frame prediction
Gustavo Sanchez, Luciano Volcan Agostini, César A. M. Marcon |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Complexity reduction by modes reduction in RD-list for intra-frame prediction in 3D-HEVC depth mapsabstractThis paper presents a complexity reduction technique for 3D High Efficiency Video Coding (3D-HEVC) using Rate-Distortion (RD) list reduction in the depth maps intra-frame prediction. In the 3D-HEVC standardization, the HEVC intra-frame prediction was inherited from texture to depth maps. However, considering depth maps simpler behavior than texture, the intra-frame prediction contains an unnecessary complexity since there is a high number of modes evaluated by their RD-cost. This paper evaluates the usage of fewer best-ranked modes in RD-list. As a case study, 23.9% of complexity reduction was achieved by reducing the Rough Mode Decision (RMD) selection and the Most Probable Modes (MPM) algorithm to the selected configuration with small impact in BD-rate. Gustavo Sanchez, Luciano Volcan Agostini, César A. M. Marcon |
ISCAS | 3 |
| 2016 | Efficient traffic balancing for NoC routing latency minimizationabstractModern technologies of integrated circuits allow billions of transistors arranged into a single chip, enabling to implement complex systems, which need a scalable and parallel communication architecture. Network-on-Chip (NoC) is a natural candidate to fulfill such communication requirements, providing high performance when the communication demands are balanced. This work proposes a new static balancing method that uses the application's traffic pattern for NoC latency reduction. This method allows the generation of a deterministic routing algorithm with simplistic implementation and low latency. Experimental results compare four balancing methods, showing the improvement of the proposed static balancing concerning the average NoC latency. Joao Marcelo Ferreira, Jarbas Silveira, Jardel Silveira, Rodrigo Cataldo, Thais Webber, Fernando Gehm Moraes, César A. M. Marcon |
ISCAS | 7 |
| 2016 | DMNI: A specialized network interface for NoC-based MPSoCsabstractCurrent proposals of NoC-based MPSoC adopt an NI (Network Interface) interconnected to a DMA (Direct Memory Access) module to enable the communication between processors through the NoC. The adoption of both modules decouples computation from communication, and a standard interface at the NI provides an abstract way for designers to connect new cores. However, this architecture is inherited from bus-based architectures and can be optimized, by removing unnecessary interfaces, signals, and registers. This paper presents a specialized communication interface for NoC-based MPSoCs, called DMNI (Direct Memory Network Interface). The DMNI merges the functionalities of the DMA and the NI into a single component, directly connecting the NoC router with the processor memory. To avoid stalls in the communication, the design of the DMNI supports simultaneous packet reception and transmission. A simplified and generic programming interface exposes the DMNI services to the software layer. Results show a reduction in the silicon area and performance improvement in the packet transmission. Marcelo Ruaro, Felipe B. Lazzarotto, César A. M. Marcon, Fernando Gehm Moraes |
ISCAS | 3 |
| 2015 | Preprocessing of Scenarios for Fast and Efficient Routing Reconfiguration in Fault-Tolerant NoCsabstractNewest processes of CMOS manufacturing allow integrating billions of transistors in a single chip. This huge integration enables to perform complex circuits, which require an energy efficient communication architecture with high scalability and parallelism degree, such as a Network-on-Chip (NoC). However, these technologies are very close to physical limitations implying the susceptibility increase of faults on manufacture and at runtime. Therefore, it is essential to provide a method for efficient fault recovery, enabling the NoC operation even in the presence of faults on routers or links, and still ensure deadlock-free routing even for irregular topologies. A preprocessing approach of the most probable fault scenarios enables to anticipate the computation of deadlock-free routings, reducing the time necessary to interrupt the system operation in a fault event. This work describes a preprocessing technique of fault scenarios based on forecasting fault tendency, which employs a fault threshold circuit and a high-level software that identifies the most relevant fault scenarios. We propose methods for dissimilarity analysis of scenarios based on measurements of cross-correlation of link fault matrices. At runtime, the preprocessing technique employs analytic metrics of average distance routing and links load for fast search of sound fault scenarios. Finally, we use RTL simulation with synthetic traffic to prove the quality of our approach. Jarbas Silveira, César A. M. Marcon, Paulo Cortez 0002, Giovanni Cordeiro Barroso, Joao Marcelo Ferreira, Rafael Mota |
PDP | 2 |
| 2014 | A monitored NoC with runtime path adaptationabstractNetworks-on-chip (NoCs) are already a common choice of communication infrastructure for complex systems-on-chip (SoCs) containing a large number of processing resources and with critical communication requirements. A NoC provides several advantages, such as higher scalability, efficient energy management, higher bandwidth and lower average latency, when compared to bus-based systems. Experiments with applications running on NoCs with more than 10% of bandwidth usage show that most of a typical message latency refers to buffered packets waiting to enter the NoC, while the latency portion that depends on packets traversing the NoC is often negligible. This work proposes a Monitored NoC called MoNoC, which is based on a monitoring mechanism and on the exchange of high-priority control packets. Practical experiments show that our fast adaptation method enables transmitting packets with smaller latencies, by using non-congested NoC areas, which reduces the most significant part of message latency. Edson I. Moreno, Thais Webber, César A. M. Marcon, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
ISCAS | 3 |
| 2014 | MoNoC: A monitored network on chip with path adaptation mechanism
Edson I. Moreno, Thais Webber, César A. M. Marcon, Fernando Gehm Moraes, Ney Laert Vilar Calazans |
J. Syst. Archit. | 3 |
| 2013 | Phoenix NoC: A distributed fault tolerant architectureabstractThe advances in deep submicron technology have made the development of large Multiprocessor Systems-on-Chip (MPSoC) possible and Networks-on-Chip (NoCs) have been recognized to provide an efficient communication architecture for such systems. With the positive effects on the device's integration some drawbacks arise, such as the increase of fault susceptibility during the MPSoC manufacturing and operation. This work presents Phoenix, which is a direct mesh NoC that implements fault tolerant mechanisms in order to enable end-to-end communication when some links fail. Phoenix implements a distributed fault tolerant mechanism in software (i.e. in each processor) and in hardware (i.e. in each router). Experimental results show that Phoenix is scalable and allows the MPSoC operation even in the presence of several faulty links. César A. M. Marcon, Alexandre M. Amory, Thais Webber, Thomas Volpato, Letícia Maria Veiras Bolzani |
ICCD | 1 |
| 2013 | A flexible framework for modeling and simulation of multipurpose wireless networksabstractThe emergence of wireless networks has contributed to a growing number of studies and protocols regarding its performance and reliability requirements, among others. Several issues have to be considered when deploying such devices under harsh environmental conditions. These issues often force the designer to adopt decisions that are usually difficult to verify in real world settings. In order to mitigate such problems, an alternative resides in the use of simulation models for both homogeneous and heterogeneous devices. This paper describes an event-based Wireless Network Simulator (WiNeS) for devices operating in several topologies and configurations of networks. WiNeS is a Java-based framework specially built to support customized network options that offers hybrid simulation for virtual and physical nodes in the same environment. Some of WiNeS' features include the computation of maximum communication distances among devices in 2D and 3D spatial node distributions as well as pairing rules to evaluate the nodes connectivity. Vinicius Bohrer, Ramon Fernandes, César A. M. Marcon, Thais Webber, Letícia Maria Veiras Bolzani, Ricardo M. Czekster, Fabiano Hessel |
RSP | 3 |
| 2013 | An implementation of a distributed fault-tolerant mechanism for 2D mesh NoCsabstractAdvances in design integration have enabled the integration of large Multiprocessor Systems-on-Chip (MPSoC). Such systems are prone to the execution of complex applications if high degree of parallelism is employed on the communication infrastructure. Network-on-Chip (NoC) has emerged as a new communication paradigm for large MPSoCs with advantages such as the increase of reliability on components interactions. However, device's integration may convey few shortcomings during MPSoC manufacturing and operation, for instance, the vulnerability to faults. This paper describes Phoenix, which is a direct mesh NoC with fault detection scheme. The proposed architecture explores a fault-tolerant mechanism, which is implemented in a distributed manner as a fault monitor on processors and routers. Results demonstrate that Phoenix can be scalable in view of the stabilization time regarding to faults incidence, allowing MPSoC operation even with the occurrence of a large number of faults. César A. M. Marcon, Alexandre M. Amory, Felipe T. Bortolon, Thais Webber, Thomas Volpato, Jader Munareto |
RSP | 1 |
| 2012 | Buffer depth and traffic influence on 3D NoCs performanceabstract3D NoC-based architectures have emerged to reduce the network latency, the energy consumption and total area in comparison to 2D NoC topologies. However, they are characterized by various trade-offs with regard to the three dimensional structure and its performance specifications. In this paper, we present a 3D NoC mesh architecture called Lasio, whose latency and the throughput achieved, for both network and application, are evaluated considering two types of traffic patterns, varied buffer depth and a range of packet sizes. Cycle-accurate simulations demonstrated that there is a high impact of buffer depth and packet size on the NoC latency and on the application latency. Applying an appropriate buffer depth, for several sizes of packets, the application latency is reduced and throughput is increased. Yan Ghidini, Thais Webber, Edson I. Moreno, Fernando Grando, Rubem D. R. Fagundes, César A. M. Marcon |
RSP | 6 |
| 2011 | Evaluating energy consumption of homogeneous MPSoCs using spare tilesabstractThe yield of homogeneous network-on-chip based multi-processor chips can be improved with the addition of spare tiles. However, the impact of this reliability approach on the chip energy consumption is not documented. For instance, in a homogeneous MPSoC, application tasks can be placed onto any tile of a defect-free chip. On the other hand, a chip with defective tile needs a special task placement, where the faulty tile is avoided. This paper presents a task placement tool and the evaluation of energy consumption of homogeneous NoC-based MPSoCs with spare tiles. Results show NoC energy consumption overhead ranging from 1 to 10% when considering up to three faults randomly distributed over the tiles of a 3×4 mesh network. The results also indicate that faults on the central tiles typically have more impact on energy overhead. Alexandre M. Amory, Luciano Ost, César A. M. Marcon, Fernando Gehm Moraes, Marcelo Lubaszewski |
DATE | 3 |
| 2011 | CAFES: A framework for intrachip application modeling and communication architecture design
César A. M. Marcon, Ney Laert Vilar Calazans, Edson I. Moreno, Fernando Gehm Moraes, Fabiano Hessel, Altamiro Amadeu Susin |
J. Parallel Distributed Comput. | 1 |
| 2007 | Evaluation of Algorithms for Low Energy Mapping onto NoCsabstractSystems on chip (SoCs) congregate multiple modules and advanced interconnection schemes, such as networks on chip (NoCs). One relevant problem in SoC design is module mapping onto a NoC targeting low energy. To date, few works are available on design and evaluation of mapping algorithms. The main goal of this work is to propose some algorithms and evaluate its results and performance with regard to low energy NoC mappings. These include exhaustive and stochastic search methods and heuristic approaches, and some combinations. The use of combined approaches compared to pure stochastic algorithms provides average reductions above 98% in execution time, while keeping energy saving within at most 5% of the best results. In addition, one heuristic provided average reductions in execution time above 90% when compared to pure stochastic algorithms, and obtained better energy saving than combined approaches. César A. M. Marcon, Edson I. Moreno, Ney Laert Vilar Calazans, Fernando Gehm Moraes |
ISCAS | 1 |
| 2007 | A VHDL based approach for fast and accurate energy consumption estimationsabstractEfficient energy consumption became an important requirement and constraint to be considered in many systems implementations, mainly to the embedded ones. Accurate and efficient power estimation during the design phase is required, in order to meet the power specifications without a costly redesign. High abstraction levels descriptions enable fast energy consumption estimations, but hardly enable accurate estimations. It normally requires evaluations at low abstraction levels, such as electric ones. On the other hand, low abstraction levels require too much design effort and design time. In this sense, this work presents an approach for energy consumption estimation for systems written in synthesizable VHDLs. A VHDL cell library is the base of the methodology, which is characterized with some relevant energy consumption information according to foundry parameters. The use of this approach leads to high-quality energy consumption estimations and design time saving. César A. M. Marcon, Sergio Johann Filho, Fabiano Hessel |
VLSI-SoC | 1 |
| 2007 | A Flexible Design Flow for a Low Power RFID TagabstractThis paper describes the implementation of a passive RFID tag targeting low power implementation, which works on 915 MHz UHF frequency. The proposed architecture allows customizing the command sets implemented inside its digital block, according to the target application needs, saving area and reducing power consumption. A flexible design flow is proposed for the customization, verification and synthesis of the digital block, targeting low power requirements. José Carlos S. Palma, César A. M. Marcon, Fabiano Hessel, Eduardo Augusto Bezerra, Guilherme Rohde, Luciano Azevedo, Carlos Eduardo Reif, Carolina Metzler |
VLSI-SoC | 2 |
| 2006 | Scheduling refinement in abstract RTOS modelsabstractScheduling decision for real-time embedded software applications has a great impact on system performance and, therefore, is an important issue in RTOS design. Moreover, it is highly desirable to have the system designer able to evaluate and select the right scheduling policy at high abstraction levels, in order to allow faster exploration of the design space. In this paper, we address this problem by introducing an abstract RTOS model, as well as a new approach to refine an unscheduled high-level model to a high-level model with RTOS scheduling. This approach is based on SystemC language and enables the system designer to quickly evaluate different dynamic scheduling policies and make the optimal choice in early design stages. Furthermore, we present a case of study where our model is used to simulate and analyze a telecom system. Fabiano Hessel, Vitor M. da Rosa, Carlos Eduardo Reif, César A. M. Marcon, Tatiana Gadelha Serra dos Santos |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2005 | Time and energy efficient mapping of embedded applications onto NoCsabstractThis work analyzes, the mapping of applications onto generic regular Networks-on-Chip (NoCs). Cores must be placed considering communication requirements so as to minimize the overall application execution time and energy consumption. We expand previous mapping strategies by taking into consideration the dynamic behavior of the target application and thus potential contentions in the intercommunication of the cores. Experimental results for a suite of 22 benchmarks and various NoC sizes show that a 42% average reduction in the execution time of the mapped application can be obtained, together with a 21% average reduction in the total energy consumption for state-of-the-art technologies. César A. M. Marcon, André Borin Soares, Altamiro Amadeu Susin, Luigi Carro, Flávio Rech Wagner |
ASP-DAC | 1 |
| 2005 | Exploring NoC Mapping Strategies: An Energy and Timing Aware TechniqueabstractComplex applications implemented as systems on chip (SoC) demand extensive use of system level modeling and validation. Their implementation gathers a large number of complex IP cores and advanced interconnection schemes, such as hierarchical bus architectures or networks on chip (NoC). Modeling applications involves capturing its computation and communication characteristics. Previously proposed communication weighted models (CWM) consider only the application communication aspects. This work proposes a communication dependence and computation model (CDCM) that can simultaneously consider both aspects of an application. It presents a solution to the problem of mapping applications on regular NoC while considering execution time and energy consumption. The use of CDCM is shown to provide estimated average reductions of 40% in execution time, and 20% in energy consumption, for current technologies. César A. M. Marcon, Ney Laert Vilar Calazans, Fernando Gehm Moraes, Altamiro Amadeu Susin, Igor M. Reis, Fabiano Hessel |
DATE | 1 |
| 2005 | Modeling the Traffic Effect for the Application Cores Mapping Problem onto NoCs
César A. M. Marcon, José Carlos S. Palma, Ney Laert Vilar Calazans, Fernando Gehm Moraes, Altamiro Amadeu Susin, Ricardo Augusto da Luz Reis |
VLSI-SoC | 1 |
| 1993 | SHC-SLX: A levelized compiled, event driven interpreted VLSI simulator
Luigi Carro, César A. M. Marcon, Altamiro Amadeu Susin |
Microprocess. Microprogramming | 2 |