Vladimir Castro Alves

dblp:54/3127 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 61% Hardware accelerators and domain-specific architectures · 15% Memory systems · 15%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
computational storage
0.922020
Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications · ACM Trans. Storage 2020
Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020
Storage systems › computational storage
in-storage computing
0.922020
Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications · ACM Trans. Storage 2020
Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412020
Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020
Memory systems › processing-in-memory
near-data processing
0.412020
Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications · ACM Trans. Storage 2020
Distributed systems › distributed machine learning
distributed training
0.112020
Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112020
Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices · DAC 2020

Methods — techniques the papers use, named apart from their topics

distributed near-data processing · 0.9cost analysis · 0.9distributed training · 0.4batch size scheduling · 0.4
YearPublicationVenuePosition
2022 Leveraging Computational Storage for Power-Efficient Distributed Data Analytics
abstract
This article presents a family of computational storage drives (CSDs) and demonstrates their performance and power improvements due to in-storage processing (ISP) when running big data analytics applications. CSDs are an emerging class of solid state drives that are capable of running user code while minimizing data transfer time and energy. Applications that can benefit from in situ processing include distributed training, distributed inferencing, and databases. To achieve the full advantage of the proposed ISP architecture, we propose software solutions for workload balancing before and at runtime for training and inferencing applications. Other applications such as sharding-based databases can readily take advantage of our ISP structure without additional tooling. Experimental results on different capacity and form factors of CSDs show up to 3.1× speedup in processing while reducing the energy consumption and data transfer by up to 67% and 68%, respectively, compared to regular enterprise solid state drives.
Ali Heydarigorji, Siavash Rezaei, Mahdi Torabzadehkashi, Hossein Bobarshad, Vladimir Castro Alves, Pai H. Chou
ACM Trans. Embed. Comput. Syst.5
2020 Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices
abstract
Computational storage devices enable in-storage processing of data in place. These devices contain 64-bit application processors and hardware accelerators that can help improving performance and saving power by reducing or eliminating data movement between host computers and storage units. This paper proposes a framework, named Stannis, for distributed in-storage training of deep neural networks on clusters of computational storage devices. This in-storage processing style of training ensures that private data never leaves the storage while fully controlling the public sharing of data. The Stannis framework distributes the workload based on the processing power of each worker by determining the proper batch size for each node. Stannis also ensures the availability of input data for all nodes to avoid rank stall while maximizing the utilization and overall processing speed. Experimental results show up to 2.7x speedup and 69% reduction in energy consumption with no significant loss in accuracy.
Ali Heydarigorji, Mahdi Torabzadehkashi, Siavash Rezaei, Hossein Bobarshad, Vladimir Castro Alves, Pai H. Chou
DAC5
2020 HyperTune: Dynamic Hyperparameter Tuning for Efficient Distribution of DNN Training Over Heterogeneous Systems
abstract
Distributed training is a novel approach to accelerating training of Deep Neural Networks (DNN), but common training libraries fall short of addressing the distributed nature of heterogeneous processors or interruption by other workloads on the shared processing nodes. This paper describes distributed training of DNN on computational storage devices (CSD), which are NAND flash-based, high-capacity data storage with internal processing engines. A CSD-based distributed architecture incorporates the advantages of federated learning in terms of performance scalability, resiliency, and data privacy by eliminating the unnecessary data movement between the storage device and the host processor. The paper also describes Stannis, a DNN training framework that improves on the shortcomings of existing distributed training frameworks by dynamically tuning the training hyperparameters in heterogeneous systems to maintain the maximum overall processing speed in term of processed images per second and energy efficiency. Experimental results on image classification training benchmarks show up to 3.1x improvement in performance and 2.45x reduction in energy consumption when using Stannis plus CSD compare to the generic systems.
Ali Heydarigorji, Siavash Rezaei, Mahdi Torabzadehkashi, Hossein Bobarshad, Vladimir Castro Alves, Pai H. Chou
ICCAD5
2020 Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications
abstract
The growing volume of data produced continuously in the Cloud and at the Edge poses significant challenges for large-scale AI applications to extract and learn useful information from the data in a timely and efficient way. The goal of this article is to explore the use of computational storage to address such challenges by distributed near-data processing. We describe Newport, a high-performance and energy-efficient computational storage developed for realizing the full potential of in-storage processing. To the best of our knowledge, Newport is the first commodity SSD that can be configured to run a server-like operating system, greatly minimizing the effort for creating and maintaining applications running inside the storage. We analyze the benefits of using Newport by running complex AI applications such as image similarity search and object tracking on a large visual dataset. The results demonstrate that data-intensive AI workloads can be efficiently parallelized and offloaded, even to a small set of Newport drives with significant performance gains and energy savings. In addition, we introduce a comprehensive taxonomy of existing computational storage solutions together with a realistic cost analysis for high-volume production, giving a good big picture of the economic feasibility of the computational storage technology.
Jaeyoung Do, Victor da Cruz Ferreira, Hossein Bobarshad, Mahdi Torabzadehkashi, Siavash Rezaei, Ali Heydarigorji, Diego Fonseca Pereira de Souza, Brunno F. Goldstein, Leandro Santiago de Araújo, Min Soo Kim 0009, Priscila M. V. Lima, Felipe M. G. França, Vladimir Castro Alves
ACM Trans. Storage13
2019 A Feasible FPGA Weightless Neural Accelerator
abstract
AI applications have recently driven the computer architecture industry towards novel and more efficient dedicated hardware accelerators and tools. Weightless Neural Networks (WNNs) is a class of Artificial Neural Networks (ANNs) often applied to pattern recognition problems. It uses a set of Random Access Memories (RAMs) as the main mechanism for training and classifying information regarding a given input pattern. Due to its memory-based architecture, it can be easily mapped onto hardware and also greatly accelerated by means of a dedicated Register Transfer-Level (RTL) architecture designed to enable multiple memory accesses in parallel. On the other hand, a straightforward WNN hardware implementation requires too much memory resources for both ASIC and FPGA variants. This work aims at designing and evaluating a Weightless Neural accelerator designed in High-Level Synthesis (HLS). Our WNN accelerator implements Hash Tables, instead of regular RAMs, to substantially reduce its memory requirements, so that it can be implemented in a fairly small-sized Xilinx FPGA. Performance, circuit-area and power consumption results show that our accelerator can efficiently learn and classify the MNIST dataset in about 8 times faster than the system's embedded ARM processor.
Victor da Cruz Ferreira, Alexandre Solon Nery, Leandro A. J. Marzulo, Leandro Santiago de Araújo, Diego Fonseca Pereira de Souza, Brunno F. Goldstein, Felipe M. G. França, Vladimir Castro Alves
ISCAS8
2019 Catalina: In-Storage Processing Acceleration for Scalable Big Data Analytics
abstract
Cloud applications are increasingly playing a crucial role in big data analytics. New use cases such as autonomous cars and edge computing call for novel approaches mixing heterogeneous computing and machine learning. These applications typically process petabyte-scale datasets, therefore, requiring low-power and scalable storage providing low-latency and high-throughput data access. While data centers have been focusing on migrating from legacy HDDs and SATA SSDs by deploying high-throughput and low-latency NVMe SSDs, the data bottlenecks appear as capacity scales. One approach to tackle this problem is to enable processing to happen within the storage device -in-storage processing (ISP)- eliminating the need to move the data. In this paper, we investigated the deployment of storage units with embedded low-power application processors along with FPGA-based reconfigurable hardware accelerators to address both performance and energy efficiency. To this purpose, we developed a high-capacity solid-state drive (SSD) named Catalina equipped with a quad-core ARM A53 processor running a Linux operating system along with a highly efficient FPGA accelerator for running applications in-place. We evaluated our proposed approach on a case study application for a similarity search library called Faiss.
Mahdi Torabzadehkashi, Siavash Rezaei, Ali Heydarigorji, Hossein Bobarshad, Vladimir Castro Alves, Nader Bagherzadeh
PDP5
2002 Filters Designed for Testability Wrapped on the Mixed-Signal Test Bus
abstract
This work presents a design for test method for continuous time active filters of any order, using the IEEE 1149.4 as its backbone structure. The method relies on the synthesis of filter transfer functions using partial fraction extraction. Transfer functions are built from 1/sup st/ order blocks connected via the available standard infrastructure. Under this approach, structural test can be carried out using simple test vectors, which are disclosed according to a fault simulation process.
José Vicente Calvano, Vladimir Castro Alves, Antonio Carneiro de Mesquita Filho, Marcelo Lubaszewski
VTS2
2001 Fault Models and Test Generation for OpAmp Circuits - The FFM
José Vicente Calvano, Antonio Carneiro de Mesquita Filho, Vladimir Castro Alves, Marcelo Lubaszewski
J. Electron. Test.3
2000 Testing a PWM circuit using functional fault models and compact test vectors for operational amplifiers
abstract
The use of analog VLSI technology on ordinary but complex electronic products has in the test one of its last frontiers. The design for testability paradigm should allow the test plan implementation early in the design cycle. However, in a successful test strategy, fault simulation should be carried out in order to evaluate appropriate test patterns, fault grade, etc. This way adequate fault models must be established. This paper shows the suitability and the straightforward consequences on testing of complex analog circuits when using OpAmp functional fault macromodels. Due to the lack of fault models, suitable for operational amplifiers fault simulation, we propose methodology for functional fault modeling and a method for test pattern generation. A fault dictionary for OpAmps is built and a procedure for compact test vector construction is proposed. The method is used to detect OpAmp faults in a pulse width modulator. The obtained results show that the proposed method is able to verify high level OpAmp requirements, such as open loop gain, slew-rate and CMMR, with good compromise between fault modeling and the analog circuit simulation complexity.
José Vicente Calvano, Vladimir Castro Alves, Marcelo Lubaszewski
Asian Test Symposium2
2000 Fault Detection Methodology and BIST Method for 2nd Order Butterworth, Chebyshev and Bessel Filter Approximations
abstract
This work proposes a new BIST scheme for 2nd order Butterworth, Chebyshev and Bessel filter approximations, using the transient analysis of simple input test vectors. A functional approach for fault modeling in 2nd order filters is presented and the transient response method is used for fault detection. The approach considers the filter as a 2nd order dynamic system where /spl omega/c and Qp deviations are faults to be detected. The peak time is the observed parameter that is evaluated in order to verify the filter correctness. The obtained results are very promising since all of /spl omega/c and Qp deviations as well as 100% of passive components are detected for this BIST scheme.
José Vicente Calvano, Vladimir Castro Alves, Marcelo Lubaszewski
VTS2
1999 An FPGA-Based Fan Beam Image Reconstruction Module
abstract
Filtered Back-Projection (FBP) is a well-known algorithm for reconstruction of tomographic images from projections. Some of FBP's highlights are: (i) it allows agile software implementations, and; (ii) it produces images of good quality, i.e., relatively free of artifacts. Our goal is to reconstruct images from fan beam projections collected by detectors set in a linear array.
Luiz Maltar, Felipe M. G. França, Vladimir Castro Alves, Claudio Luis de Amorim
FCCM3
1998 A BIST Scheme for Asynchronous Logic
abstract
This work introduces a methodology to ease the implementation of BIST in asynchronous circuits. Scheduling by edge reversal (SER), a simple but powerful distributed synchronizer is used to implement a sequencer that allows testing the circuit at full speed. The methodology, which allows the detection of topological faults, is proved correct. Low hardware overhead and the absence of deadlocks are the main characteristics of the proposed methodology.
Vladimir Castro Alves, Felipe M. G. França, Edson do Prado Granja
Asian Test Symposium1
1998 Implementation of RNS Addition and RNS Multiplication into FPGAs
abstract
We investigate whether arithmetic operations based on Residue Number Systems (RNS) are cost-effective solutions to implement DSP applications into reconfigurable hardware. We simulated several RNS addition and multiplication implementations by varying the RNS parameters. For RNS addition, our results show that it can be implemented into a 3-stage 80.6-92.5 MHz pipeline using about 22 to 33 FPGAs' logic cells. For RNS multiplication, the attainable speed range was between 78.1 and 87.7 MHz, for operand lengths varying between 5 and 8 bits. Overall, a hybrid solution that combines logical elements and blocks of RAM is the best option, producing better average performance across the whole range of operand lengths.
Luiz Maltar, Felipe M. G. França, Vladimir Castro Alves, Claudio Luis de Amorim
FCCM3
1996 A Pragmatic, Systematic And Flexible Synthesis For Testability Methodology
abstract
Starting from the analysis of the most widely adopted methodologies and the most successful industrial tools in the fields of HLS and DFT, this paper proposes a general framework for a pragmatic, systematic and flexible SFT methodology. The prerequisites for such a methodology, together with the state of the art are first assessed, then an overview of the approach is presented, followed by step-by-step details through a case-study of High-Level Synthesis For BIST. Examples of first obtained results are also provided.
Vladimir Castro Alves, A. Ribeiro Antunes, Meryem Marzouki
Asian Test Symposium1
1995 Testing complex couplings in multiport memories
abstract
In this paper, the effects of simultaneous write access on the fault modeling of multiport RAMs are investigated. New fault models representing more accurately the actual faults in such memories are then defined. Subsequently, a general algorithm that ensures the detection of all faults belonging to the new fault model is proposed. Unfortunately, the obtained algorithms are of O(n/sup 2/) complexity which is not practical for real purposes. In order to reduce the complexity of the former test algorithm a topological approach has been developed. Finally, a BIST implementation of one of the proposed topological algorithms is presented.>
Michael Nicolaidis, Vladimir Castro Alves, Hakim Bederr
IEEE Trans. Very Large Scale Integr. Syst.2
1994 Trade-offs in scan path and BIST implementations for RAMs
Michael Nicolaidis, O. Kebichi, Vladimir Castro Alves
J. Electron. Test.3
1991 Built-In Self-Test for Multi-Port RAMs
abstract
The authors present a novel approach to the test of multi-port RAMs. A novel fault model that takes into account complex couplings resulting from simultaneous access of memory cells is used in order to ensure a very high fault coverage. A novel algorithm for the test of dual-port memories is detailed. This algorithm achieves O(n) complexity thanks to the use of some topological restrictions. The authors also present a novel built-in self-test (BIST) scheme, based on programmable schematic generators, that allows great flexibility for ASIC (application-specific integrated circuit) design.>
Vladimir Castro Alves, Michael Nicolaidis, P. Lestrat, Bernard Courtois
ICCAD1