VLDB 2026 Research / reviewers in the wild / expert
João Paulo C. de Lima
dblp:156/5239 · also João Paulo Cardoso de Lima
· DBLP profile ↗
18ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0001-9295-3519ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 8 first-author · 12 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Low-overhead Bitwise Shifting in DRAMabstractProcessing-in-Memory (PIM) architectures enable computation directly within DRAM and help combat the memory wall problem. While commodity DRAM has been demonstrated for data movement and bulk-bitwise PIM, data must be reorganized away from row-wise ordering and byte addressability used by typical processors into a column-ordering where bytes and words are spread across rows within a single column since there is no mechanism to move data between bitlines which is required for many computational operations including addition and multiplication. In this paper, we propose a DRAM subarray design that enables in-DRAM bit-shifting for open-bitline architectures. By adding migration cells at the top and bottom of each subarray, bidirectional bit-shifting may be conducted within any given row. In this paper we present initial findings and evaluation of potential timing and energy analysis using a combination of approaches from VLSI layout, spice simulation and demonstrate integration into NVMain-PIM a simulator extended with DRAM PIM primitives. William C. Tegge, João Paulo C. de Lima, Benjamin F. Morris III, Alex K. Jones |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Count2Multiply: Reliable In-Memory High-Radix CountingabstractComputing-in-memory (CIM) has been demonstrated across various memory technologies, from memristive crossbars for analog dot-products to large-scale digital bitwise operations in commodity DRAM and other non-volatile memory technologies. However, current CIM solutions face challenges related to latency and reliability. CIM fidelity lags considerably behind standard memory access fidelity. Furthermore, bulkbitwise CIM, though highly parallel, requires long latency for multiplication and addition due to bit-serial execution. This paper presents Count2Multiply, a digital CIM framework that performs multiplication, addition, and other operations using high-radix, massively parallel counting enabled by bulk-bitwise inmemory operations. Designed to meet fault tolerance requirements, Count2Multiply integrates traditional row-wise error correction codes, such as Hamming and BCH, to address the high error rates in existing CIM designs. We demonstrate Count2Multiply in commodity DRAMs, achieving on average 1.5× speedup, 4.6× higher energy efficiency (GOPS/Watt), and 4.4× better area efficiency (GOPS/mm2) over NVIDIA A100; and 9.7×, 8.1×, and 12.9× improvements, respectively, over SIMDRAM. João Paulo C. de Lima, Benjamin F. Morris III, Asif Ali Khan, Jerónimo Castrillón, Alex K. Jones |
HPCA | 1 |
| 2025 | Hardware-Aware Compilation and Simulation for In-Memory ComputingabstractThis brief presents an overview of recent tools and research efforts aimed at enhancing the programmability and reliability of In-Memory Computing (IMC)-based systems. We discuss hardware-aware training techniques that improve model resilience to analog device imperfections, and explore mapping strategies that balance accuracy and performance for heterogeneous IMC-based accelerators. Additionally, we examine a compiler framework that abstracts hardware complexities and enables seamless integration of these accelerators into existing deployment pipelines. By combining these approaches with advanced simulation tools, we propose an end-to-end workflow that facilitates the practical deployment and optimization of IMC technologies across diverse memory types and architectural designs. Asif Ali Khan, Hadjer Benmeziane, Hamid Farzaneh, João Paulo C. de Lima, William Andrew Simon, Yiyu Shi 0001, Zheyu Yan, Abu Sebastian, Xiaobo Sharon Hu, Jerónimo Castrillón, Corey Lammie |
CASES | 4 |
| 2025 | All-in-Memory Stochastic Computing using ReRAMabstractAs the demand for efficient, low-power computing in embedded and edge devices grows, traditional computing methods are becoming less effective for handling complex tasks. Stochastic computing (SC) offers a promising alternative by approximating complex arithmetic operations, such as addition and multiplication, using simple bitwise operations, like majority or AND, on random bit-streams. While SC operations are inherently fault-tolerant, their accuracy largely depends on the length and quality of the stochastic bit-streams (SBS). These bit-streams are typically generated by CMOS-based stochastic bit-stream generators that consume over 80% of the SC system’s power and area. Current SC solutions focus on optimizing the logic gates but often neglect the high cost of moving the bit-streams between memory and processor. This work leverages the physics of emerging ReRAM devices to implement the entire SC flow in place: ❶ generating low-cost true random numbers and SBSs, ❷ conducting SC operations, and ❸ converting SBSs back to binary. Considering the low reliability of ReRAM cells, we demonstrate how SC’s robustness to errors copes with ReRAM’s variability. Our evaluation shows significant improvements in throughput (1.39 ×, 2.16 ×) and energy consumption (1.15 ×, 2.8 ×) over state-of-the-art (CMOS- and ReRAM-based) solutions, respectively, with an average image quality drop of 5% across multiple SBS lengths and image processing tasks. João Paulo C. de Lima, Mehran Shoushtari Moghadam, Sercan Aygün, Jerónimo Castrillón, M. Hassan Najafi, Asif Ali Khan |
DAC | 1 |
| 2024 | C4CAM: A Compiler for CAM-based In-memory AcceleratorsabstractMachine learning and data analytics applications increasingly suffer from the high latency and energy consumption of conventional von Neumann architectures. Recently, several in-memory and near-memory systems have been proposed to overcome this von Neumann bottleneck. Platforms based on content-addressable memories (CAMs) are particularly interesting due to their efficient support for the search-based operations that form the foundation for many applications, including K-nearest neighbors (KNN), high-dimensional computing (HDC), recommender systems, and one-shot learning among others. Today, these platforms are designed by hand and can only be programmed with low-level code, accessible only to hardware experts. In this paper, we introduce C4CAM, the first compiler framework to quickly explore CAM configurations and seamlessly generate code from high-level Torch-Script code. C4CAM employs a hierarchy of abstractions that progressively lowers programs, allowing code transformations at the most suitable abstraction level. Depending on the type and technology, CAM arrays exhibit varying latencies and power profiles. Our framework allows analyzing the impact of such differences in terms of system-level performance and energy consumption, and thus supports designers in selecting appropriate designs for a given application. Hamid Farzaneh, João Paulo C. de Lima, Mengyuan Li 0001, Asif Ali Khan, Xiaobo Sharon Hu, Jerónimo Castrillón |
ASPLOS (3) | 2 |
| 2024 | SHERLOCK: Scheduling Efficient and Reliable Bulk Bitwise Operations in NVMsabstractBulk bitwise operations are commonplace in application domains such as databases, web search, cryptography, and image processing. The ever-growing volume of data and processing demands of these domains often result in high energy consumption and latency in conventional system architectures, mainly due to data movement between the processing and memory subsystems. Non-volatile memories (NVMs), such as RRAM, PCM and STT-MRAM, facilitate conducting bulk-bitwise logic operations in-memory (CIM). Efficient mapping of complex applications to these CIM-capable NVMs is non-trivial and can even lead to slowdowns. This paper presents Sherlock, a novel mapping and scheduling method for efficient execution of bulk bitwise operations in NVMs. Sherlock collaboratively optimizes for performance and energy consumption and outperforms the state-of-the-art by 10× and 4.6×, respectively. Hamid Farzaneh, João Paulo C. de Lima, Ali Nezhadi, Asif Ali Khan, Mahta Mayahinia, Mehdi Baradaran Tahoori, Jerónimo Castrillón |
DAC | 2 |
| 2024 | Full-Stack Optimization for CAM-Only DNN InferenceabstractThe accuracy of neural networks has greatly improved across various domains over the past years. Their ever-increasing complexity, however, leads to prohibitively high energy demands and latency in von-Neumann systems. Several computing-in-memory (CIM) systems have recently been proposed to overcome this, but trade-offs involving accuracy, hardware reliability, and scalability for large models remain a challenge. Additionally, for some CIM designs, the activation movement still requires considerable time and energy. This paper explores the combination of algorithmic optimizations for ternary weight neural networks and associative processors (APs) implemented using racetrack memory (RTM). We propose a novel compilation flow to optimize convolutions on APs by reducing their arithmetic intensity. By leveraging the benefits of RTM-based APs, this approach substantially reduces data transfers within the memory while addressing accuracy, energy efficiency, and reliability concerns. Concretely, our solution improves the energy efficiency of ResNet-18 inference on ImageNet by 7.5× compared to crossbar in-memory accelerators while retaining software accuracy. João Paulo C. de Lima, Asif Ali Khan, Luigi Carro, Jerónimo Castrillón |
DATE | 1 |
| 2024 | Smoothing Disruption Across the Stack: Tales of Memory, Heterogeneity, & CompilersabstractMultiple research vectors represent possible paths to improved energy and performance metrics at the application-level. There are active efforts with respect to emerging logic devices, new memory technologies, novel interconnects, and heterogeneous integration architectures. Of great interest is quantifying the potential impact of a given solution to prioritize research vectors accordingly. In this paper, we discuss two efforts - one focused on emerging memory technology, and another focused on heterogeneous integration technology - that speak to best practices for, and needed contributions from the design automation (DA) community to explore this vast design space. Furthermore, we highlight new research efforts that aim to develop the novel compiler abstractions and frameworks that are ultimately needed to derive maximum value from new memory and/or heterogeneous and monolithic integration architecture, and that can also play an important role with respect to design space exploration efforts. Michael T. Niemier, Zephan M. Enciso, M. Sharifi, Xiaobo Sharon Hu, Ian O'Connor, A. Graening, Jerónimo Castrillón, João Paulo C. de Lima, Asif Ali Khan, Hamid Farzaneh, N. Afroze, Julien Ryckaert |
DATE | 10 |
| 2023 | Special Session - Non-Volatile Memories: Challenges and Opportunities for Embedded System Architectures with Focus on Machine Learning ApplicationsabstractThis paper explores the challenges and opportunities of integrating non-volatile memories (NVMs) into embedded systems for machine learning. NVMs offer advantages such as increased memory density, lower power consumption, non-volatility, and compute-in-memory capabilities. The paper focuses on integrating NVMs into embedded systems, particularly in intermittent computing, where systems operate during periods of available energy. NVM technologies bring persistence closer to the CPU core, enabling efficient designs for energy-constrained scenarios. Next, computation in resistive NVMs is explored, highlighting its potential for accelerating machine learning algorithms. However, challenges related to reliability and device non-idealities need to be addressed. The paper also discusses memory-centric machine learning, leveraging NVMs to overcome the memory wall challenge. By optimizing memory layouts and utilizing probabilistic decision tree execution and neural network sparsity, NVM-based systems can improve cache behavior and reduce unnecessary computations. In conclusion, the paper emphasizes the need for further research and optimization for the widespread adoption of NVMs in embedded systems presenting relevant challenges, especially for machine learning applications. Jörg Henkel, Lokesh Siddhu, Lars Bauer, Jürgen Teich, Stefan Wildermann, Mehdi Baradaran Tahoori, Mahta Mayahinia, Jerónimo Castrillón, Asif Ali Khan, Hamid Farzaneh, João Paulo C. de Lima, Jian-Jia Chen, Christian Hakert, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng |
CASES | 11 |
| 2023 | Data and Computation Reuse in CNNs Using Memristor TCAMsabstractExploiting computational and data reuse in CNNs is crucial for the successful design of resource-constrained platforms. In image recognition applications, high levels of input locality and redundancy present in CNNs have become the golden goose for skipping costly arithmetic operations. One promising technique for this consists in storing function responses of some input patterns into offline lookup tables and replacing online computation with search operations, which are highly efficient when implemented by emerging non-volatile memory technologies. In this work, we rethink both algorithm and architecture for exploiting locality and reuse opportunities by replacing entire convolutions with searches on Content-addressable Memories. By previously calculating convolution results and building compact lookup tables with our novel clustering algorithm, one can evaluate activations at constant time complexity, also requiring a single read operation of the current input tensor. Then, we devise a reconfigurable array of processing elements based on memristive Ternary Content-addressable Memories to efficiently implement the algorithmic solution and meet the flexibility requirements of several CNN architectures. Results show that our design reduces the number of multiplications and memory accesses proportionally to the number of convolutional layer channels. The average performance is 1,172 and 82 FPS for AlexNet and VGG-16 models, thus outperforming state-of-the-art works by 13×. Rafael Fao de Moura, João Paulo C. de Lima, Luigi Carro |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Quantization-Aware In-situ Training for Reliable and Accurate Edge AIabstractIn-memory analog computation based on memristor crossbars has become the most promising approach for DNN inference. Because compute and memory requirements are larger during training, memristive crossbars are also an alternative to train DNN models within a feasible energy budget for edge devices, especially in the light of trends towards security, privacy, latency, and energy reduction, by avoiding data transfer over the Internet. To enable online training and inference on the same device, however, there are still challenges related to different minimum bitwidth needed in each phase, and memristor non-idealities to be addressed. We provide an in-situ training framework that allows the network to adapt to hardware imperfections, while practically eliminating errors from weight quantization. We validate our methodology with image classifiers, namely MNIST and CIFAR10, by training NN models with 8-bit weights and quantizing to 2 bits. The training algorithm recovers up to 12 % of the accuracy lost to quantization errors even under high variability, reduces training energy by up to 6 ×, and allows for energy-efficient inferences using a single cell per synapse, hence enhancing robustness and accuracy for a smooth training-to-inference transition. João Paulo C. de Lima, Luigi Carro |
DATE | 1 |
| 2022 | STAP: An Architecture and Design Tool for Automata Processing on Memristor TCAMsabstractAccelerating finite-state automata benefits several emerging application domains that are built on pattern matching. In-memory architectures, such as the Automata Processor (AP), are efficient to speed them up, at least for outperforming traditional von-Neumann architectures. In spite of the AP’s massive parallelism, current APs suffer from poor memory density, inefficient routing architectures, and limited capabilities. Although these limitations can be lessened by emerging memory technologies, its architecture is still the major source of huge communication demands and lack of scalability. To address these issues, we present STAP , a Scalable TCAM-based architecture for Automata Processing . STAP adopts a reconfigurable array of processing elements, which are based on memristive Ternary CAMs (TCAMs), to efficiently implement Non-deterministic finite automata (NFAs) through proper encoding and mapping methods. The CAD tool for STAP integrates the design flow of automata applications, a specific mapping algorithm, and place and route tools for connecting processing elements by RRAM-based programmable interconnects. Results showed 1.47× higher throughput when processing 16-bit input symbols, and improvements of 3.9× and 25× on state and routing densities over the state-of-the-art AP, while preserving 10 4 programming cycles. João Paulo C. de Lima, Marcelo Brandalero, Michael Hübner 0001, Luigi Carro |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2020 | Endurance-Aware RRAM-Based Reconfigurable Architecture using TCAM ArraysabstractField-Programmable Gate Arrays (FPGAs) have enabled the acceleration of important applications in the networking, cloud, and artificial intelligence domains, while providing a flexible fabric that can be reprogrammed on demand. Still, the high static power dissipation of FPGAs driven by Static Random Access Memories (SRAMs) leads them to energy consumption levels that may be unacceptable for several application domains. Reconfigurable fabrics with emerging Resistive RAM (RRAM) technologies have been considered as one of the most promising solutions to address these energy issues of current FPGAs. However, the low endurance and the high variability of these emerging devices present a threat to the demands for reconfiguration cycles of current applications, pushing for novel architectures and design strategies techniques for improving the device's lifetime. To address these challenges, we propose a novel reconfigurable architecture targeting classes of applications that require high flexibility in the field. More specifically, we introduce a reconfigurable architecture based on Ternary Content-Addressable Memories (TCAMs) that meets a double mission: to accelerate and tolerate endurance and variation issues supported by a CAD tool that foresees the reuse of data configuration, allowing for an increased endurance in the field. We present the potential of the proposed architecture and its synthesis flow for processing Regular Expression Matching (REM), widely used in network intrusion detection systems. The results show that the performance can achieve up to 32Gbps throughput at 0.89W, while improving the device's lifetime by two orders of magnitude. João Paulo C. de Lima, Marcelo Brandalero, Luigi Carro |
FPL | 1 |
| 2020 | Leveraging reuse and endurance by efficient mapping and placement for NVM-based FPGAsabstractAlthough NVM technologies can bring higher density, near-zero leakage power, and CMOS compatibility, the reduced endurance of such devices is still a major limitation for a broad range of applications, including NVM-based FPGAs. In this work, we propose to drastically reduce LUT writing and word flips occurrence by taking reconfiguration reuse as a target of optimization in the CAD flow. Specifically, we investigate circuit reuse in LUT- and CAM-based FPGAs, and propose a new endurance-aware mapping and placement algorithm. Simulation results show that much greater reuse can be achieved with PLA-style, and hence higher endurance is expected. João Paulo C. de Lima, Rafael Fao de Moura, Luigi Carro |
IOLTS | 1 |
| 2019 | A Compiler for Automatic Selection of Suitable Processing-in-Memory InstructionsabstractAlthough not a new technique, due to the advent of 3D-stacked technologies, the integration of large memories and logic circuitry able to compute large amount of data has revived the Processing-in-Memory (PIM) techniques. PIM is a technique to increase performance while reducing energy consumption when dealing with large amounts of data. Despite several designs of PIM are available in the literature, their effective implementation still burdens the programmer. Also, various PIM instances are required to take advantage of the internal 3D-stacked memories, which further increases the challenges faced by the programmers. In this way, this work presents the Processing-In-Memory cOmpiler (PRIMO). Our compiler is able to efficiently exploit large vector units on a PIM architecture, directly from the original code. PRIMO is able to automatically select suitable PIM operations, allowing its automatic offloading. Moreover, PRIMO concerns about several PIM instances, selecting the most suitable instance while reduces internal communication between different PIM units. The compilation results of different benchmarks depict how PRIMO is able to exploit large vectors, while achieving a near-optimal performance when compared to the ideal execution for the case study PIM. PRIMO allows a speedup of 38× for specific kernels, while on average achieves 11.8 × for a set of benchmarks from PolyBench Suite. Hameeza Ahmed, Paulo C. Santos 0001, João Paulo C. de Lima, Rafael Fao de Moura, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
DATE | 3 |
| 2018 | Design space exploration for PIM architectures in 3D-stacked memoriesabstractScaling existing architectures to large-scale data-intensive applications is limited by energy and performance losses caused by off-chip memory communication and data movements in the cache hierarchy. Processing-in-Memory (PIM) has been recently revisited to address the issues of memory and power wall, mainly due to the maturity of 3D-stacking manufacturing technology and the increasing demand for bandwidth and parallel access in emerging data-centric applications. Recent studies have shown a wide variety of processing mechanisms to be placed in the logic layer of 3D-stacked memories, not to mention the already available 3D-stacked DRAMs, such as Micron's Hybrid Memory Cube (HMC). Nevertheless, a few studies compare PIM accelerators to each other and have made efforts to indicate the trade-offs between power, area, and performance. In this paper, we review different state-of-the-art 3D-stacked in-memory accelerators, and we analyze them considering important constraints regarding area and power due to critical embedded nature of PIM. Aiming to point in the direction of massive parallel PIM designs, we take the simplest design found in this survey, and we explore the architectural design space to meet the constraints imposed by HMC. Our results show that the most straightforward approach can provide the highest performance while consuming the lowest amount of area and power, which makes it the most suitable design found in this survey for an energy-efficient in-memory accelerator, whether it goes in High-Performance Computing or Embedded Systems. For instance, the outstanding point in the design space indicates that a performance density of 320 GBps/mm2 and a performance efficiency of 0.6 GBps/mW can be achieved in the best scenario, that is, when a massive parallel application reaches the peak bandwidth. João Paulo C. de Lima, Paulo C. Santos 0001, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
CF | 1 |
| 2018 | Processing in 3D memories to speed up operations on complex data structuresabstractPointer chasing has been, for years, the kernel operation employed by diverse data structures, from graphs to hash tables and dictionaries. However, due to the bewildering growth in the volume of data that current applications have to deal with, performing pointer chasing operations have become a major source of performance and energy bottleneck, due to its sparse memory access behavior. In this work, we aim to tackle this problem by taking advantage of the already available parallelism present in today's 3D-stacked memories. We present a simple mechanism that can accelerate pointer chasing operations by making use of a state-of-the-art PIM design that executes in-memory vector operations. The key idea behind our design is to run speculative loads, in parallel, based on a given memory address in a reconfigurable window of addresses. Our design can perform pointer-chasing operations on b+tree 4.9 χ faster when compared to modern baseline systems. Besides that, since our device avoids data movement, we can also reduce energy consumption by 85% when compared to the baseline. Paulo C. Santos 0001, Geraldo F. Oliveira, João Paulo C. de Lima, Marco A. Z. Alves, Luigi Carro, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2014 | Application of remote experiments in basic education through mobile devicesabstractThis paper presents some findings related to the experience of developing and deploying an software application aimed at using remote experimentation with mobile devices and the use of Virtual Learning Environments (VLE) as tools to support teaching and learning. Currently the technological resources have been misused in education, and the potential use of mobile devices in education is virtually untapped. The architecture implemented enables the users to control real experiments and to monitor results via video streaming, which provides a sense of involvement similar to hands-on laboratories. The authors describe the deployment of this mobile learning tool in a second year Brazilian public high school. The mobile application uses HTML5, CSS3 and jQuery Mobile in order to maintain compatibility among the platforms widely used. The VLE Moodle has been used in order to provide homework, assignments and teaching material. The experiments have been automated using open hardware and software sources, which facilitate replication in different areas. João Paulo C. de Lima, Willian Rochadel, A. M. Silva, José Pedro Simão, Juarez Bento da Silva, João Bosco da Mota Alves |
EDUCON | 1 |