Willian Analdo Nunes

dblp:330/4438 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0001-5796-8416ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 61% Electronic design automation · 30% Energy-efficient computing · 9%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction set architecture
1.012026
Design Space Exploration of RISC-V Vector Extension Targeting Embedded Processors · IEEE Trans. Computers 2026
Electronic design automation › design space exploration
microarchitecture design space exploration
1.012026
Design Space Exploration of RISC-V Vector Extension Targeting Embedded Processors · IEEE Trans. Computers 2026
Processor architecture and microarchitecture › instruction set architecture › vector extension
RISC-V vector extension
1.012026
Design Space Exploration of RISC-V Vector Extension Targeting Embedded Processors · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

post-synthesis simulation · 1.0
YearPublicationVenuePosition
2026 Design Space Exploration of RISC-V Vector Extension Targeting Embedded Processors
abstract
The RISC-V Vector Extension (RVV) offers flexibility and scalability for the growing demand for data-parallel workloads in embedded systems. However, RVV configurability introduces a large design space, the implications of which for resource-constrained processors remain insufficiently quantified. Understanding how parameters such as vector length (VLEN), maximum element size (ELEN), lane count, LMUL, and SEW affect system-level metrics enables efficient acceleration while preserving software portability and reducing verification effort. This work explores the RVV design space for embedded processors using a Zve32x subset and evaluates the impact of architectural parameters on performance, area, power, and energy efficiency. Annotated post-synthesis simulations across representative benchmarks enable quantitative trade-off analysis over multiple RVV configurations. Results show that larger VLENs and aggressive LMUL settings can yield substantial speedups but often increase area and power, whereas moderate configurations provide improved energy efficiency while meeting performance targets. The paper derives practical design guidelines for selecting RVV parameters in power- and areaconstrained embedded systems. The RISC-V core used in this work is publicly available athttps://github.com/gaph-pucrs/RS5.
Willian Analdo Nunes, Antônio Vinicius Corrêa Dos Santos, Lucas Damo, Fernando Gehm Moraes
IEEE Trans. Computers1
2025 Accelerating Machine Learning with RISC-V Vector Extension and Auto-Vectorization Techniques
abstract
Convolutional neural networks (CNNs) have played a significant role in the recent evolution of machine learning (ML) due to their feature extraction and pattern recognition capabilities. CNNs involve numerous Multiply and Accumulate (MAC) operations, which are computationally expensive and often require hardware acceleration to achieve acceptable performance. The RISC-V Vector (RVV) Extension is a candidate for accelerating vector processing operations, such as those found in CNNs. Several studies have explored the application of the RVV extension in various areas of ML, where it is often implemented as a coprocessor due to its complexity. This work presents an RVV implementation as a tightly coupled accelerator for a small RISC-V processor, RS5, which implements a subset of the RVV extension to maintain a non-prohibitive area overhead. The paper explores a case study of a 1-D CNN combined with the recently introduced auto-vectorization feature in GCC 14.1. Performance results show an average speedup of 1.88x compared to a scalar core, achieved without modifying the original code.
Willian Analdo Nunes, Antônio Vinicius Corrêa Dos Santos, Fernando Gehm Moraes
ISCAS1
2025 Accelerating Machine Learning using RISC-V Vector Extension in a Manycore Platform
abstract
This work addresses the acceleration of convolutional neural network (CNN) inference in manycore architectures using coarse and fine-grain parallelism. Prior approaches focus on dedicated accelerators or modified NoCs, limiting flexibility. This work proposes integrating a RISC-V processor extended with the vector extensions (RVV) as general-purpose processing elements in a NoC-based manycore. The implementation applies depthwise convolution mapped across PEs and uses autovectorization provided by the compiler. Experiments on a $4 \times 4$ manycore running the first AlexNet layer achieved up to 5.70x speedup and reduced execution cycles by 82.45% compared to a scalar single-core baseline.
Willian Analdo Nunes, Antônio Vinicius Corrêa Dos Santos, César A. M. Marcon, Fernando Gehm Moraes
VLSI-SoC1