VLDB 2026 Research / reviewers in the wild / expert
Leonid Yavits
dblp:131/6615
· DBLP profile ↗
27ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0001-5248-3997ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 8 first-author · 13 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PatBiNN: A 65 nm Processing-in-CAM Based BNN Implementation for Pathogen Genome ClassificationabstractBinary Neural Networks (BNNs) are a cost-effective and highly efficient alternative to traditional neural networks. Genome classification is a frequent component of genome analysis pipelines, with a variety of applications spanning pandemic preparedness, AMR resistance control, drinking water and food safety. PatBiNN is a BNN based pathogen genome classifier optimized for edge and field use. It employs a binary multilayer perceptron (MLP) implemented using in-Hamming distance tolerant (similarity search) content addressable memory processing. PatBiNN was designed and manufactured in a commercial 65nm process. It achieves F1score of 88%, ROC AUC of 0.986, throughput of 0.8M inferences/s, power consumption of 4.8mW and energy efficiency of 237TOPs/s/W with silicon area of 0.87mm2. Yuval Harary, Almog Sharoni, Esteban Garzón, Leonid Yavits |
DATE | 4 |
| 2026 | BinDRAM: Binary neural network on unmodified commodity DRAM
Dan Yaron, Benjamin Wolfzon, Zuher Jahshan, Alexander Fish, Leonid Yavits |
Future Gener. Comput. Syst. | 5 |
| 2026 | GenMClass: Design and comparative analysis of genome classifier-on-chip platformabstractWe propose GenMClass, a genome classification system-on-chip (SoC) implementing two different classification approaches and comprising two separate classification engines: a DNN accelerator GenDNN, that classifies DNA reads converted to images using a classification neural network, and a similarity search-capable Error Tolerant Content Addressable Memory ETCAM, that classifies genomes by k-mer matching. Classification operations are controlled by an embedded RISCV processor. GenMClass classification platform was designed and manufactured in a commercial 65 nm process. We conduct a comparative analysis of ETCAM and GenDNN classification efficiency as well as their performance, silicon area and power consumption using silicon measurements. The size of GenMClass SoC is 3.4 mm 2 and its total power consumption (assuming both GenDNN and ETCAM perform classification at the same time) is 144 mW. This allows using GenMClass as a portable classifier for pathogen surveillance during pandemics, food safety and environmental monitoring, agriculture pathogen and antimicrobial resistance control, in the field or at points of care. Daria Bromot, Yehuda Kra, Zuher Jahshan, Esteban Garzón, Adam Teman, Leonid Yavits |
J. Syst. Archit. | 6 |
| 2026 | CADM: Content addressable commodity off-the-shelf DRAM-based genome classifierabstractProcessing using memory (PuM) leverages analog properties of memory infrastructure to implement logic and arithmetic operations. Commodity Off-The-Shelf (COTS) DRAM is particularly attractive for PuM because it requires no device modification, thereby preserving the ubiquity, availability, and cost advantages of modern DRAM while enabling massive column-level parallelism. We propose CADM (Content-Addressable DRAM), that enables exact and approximate (similarity) search in- and using- unmodified COTS DRAM. CADM targets genome classification, which is one of the most important applications in bioinformatics. Specifically, rapid and accurate detection of bacterial pathogens is critical for effective clinical decision-making, particularly in life-threatening conditions such as sepsis, where early identification of the causative agent significantly improves patient outcomes. We implement CADM in commercial DDR4 and show that it can achieve up to 185 × higher throughput and 73 × energy savings compared to CPU-run state-of-the-art classifier Kraken2. Using approximate search, CADM can achieve 9 × higher F 1 score when matching relatively short ( < 32 DNA bases) ambiguous and erroneous k -mers. Esteban Garzón, Alexander Fish, Leonid Yavits |
J. Syst. Archit. | 3 |
| 2026 | SPARCAM: Sparse matrix multiplication accelerator using multi-port dynamic CAMabstractSparse General matrix multiplication (SpGEMM) is a fundamental kernel in many scientific and engineering fields, including Artificial Intelligence (AI). However, its intrinsic computation complexity presents substantial challenges, making efficient hardware implementation particularly difficult. This paper proposes SPARCAM, a novel SpGEMM accelerator, developed and optimized for very energy-efficient AI edge applications. SPARCAM is designed using low-power dense Gain Cell embedded DRAM (GC-eDRAM) technology, a processing near memory paradigm, and a modified outer product matrix multiplication algorithm. Despite its quite limited peak theoretical performance, SPARCAM achieves very high energy efficiency due to its low-power architecture and almost 100% utilization of its computing resources. Designed in a commercial 28 nm FDSOI technology, SPARCAM achieves 13 . 9 × speedup over a high-performance embedded CPU when processing large-scale sparse matrices. When multiplying limited-size sparse matrices, SPARCAM obtains 193 × speedup over high-performance GPU. SPARCAM reaches about 4.3 orders-of-magnitude, on average, higher energy benefits, and 1892 × , 181 × , 2 × , and 3471 × , higher energy efficiency (over CPU) compared with state-of-the-art SpGEMM accelerators SpArch, OuterSPACE, MatRaptor, and high-performance GPU, respectively. Esteban Garzón, Benjamin Zambrano, David Sheinenzon, Marco Lanuzza, Adam Teman, Leonid Yavits |
J. Syst. Archit. | 6 |
| 2026 | GainP: A Gain Cell Embedded DRAM-based associative in-memory processorabstractAssociative processors (APs) are massively-parallel in-memory SIMD accelerators. While fairly well-known, APs have been revisited in recent years due to the proliferation of data-centric computing, and specifically, processing using memory. APs are based on Content Addressable Memory and utilize its unique ability to simultaneously search the entire memory content for a query pattern to implement massively parallel computations in memory. Several memory infrastructures have been considered for associative processing, including static CMOS, resistive, magnetoresistive, ferroelectric and even NAND flash memories. While all of these have certain merits (speed and low energy consumption for static CMOS, density for resistive and ferroelectric memories), they also face challenges (low density for static CMOS and magnetoresistive, limited write endurance and high write energy for resistive and ferroelectric memories), which limit the scalability and usefulness of APs. This work introduces GainP, an AP based on silicon-proven Gain Cell embedded DRAM (GCeDRAM). The latter combines relatively high density (compared to static CMOS memory) with low energy, high speed, practically unlimited endurance and low production costs (compared to emerging memory technologies). Using sparse-by-sparse matrix multiplication, we show that GainP outperforms high-performance CPU and GPU by 825 × and 41 × . We also show that GainP outperforms state-of-the-art processing-in-memory sparse matrix multiplication accelerators GAS, OuterSPACE and MatRaptor by 128 × , 125 × and 16 × , respectively, and provides average energy benefits of 96 × , 95 × and 15 × , respectively. Yaniv Levi, Odem Harel, Adam Teman, Leonid Yavits |
J. Syst. Archit. | 4 |
| 2025 | Towards Low-Power High-Performance Content-Addressable Memory: A Robust Precharge-Free ApproachabstractLow-Power high-performance content-addressable memories (CAMs) are important components in modern computing systems. In this work, we present a robust CAM that overcomes the power and performance limitations of conventional precharge-based CAMs. The proposed static transmission gate-based (STAT-TG) CAM design achieves low-power operation comparable to NAND CAMs while maintaining search speeds rivaling those of NOR CAMs. The STAT-TG CAM was designed using a 65nm CMOS technology and comprehensively evaluated under extensive Monte Carlo simulations. Compared to conventional CAMs, the STAT-TG CAM is 14% faster than NAND CAM, while consuming only 25% of the energy per operation relative to NOR CAM. This makes STAT-TG CAM a promising solution for high-performance yet energy-efficient applications. Ramiro Taco, Esteban Garzón, Adam Teman, Leonid Yavits, Marco Lanuzza |
ISCAS | 4 |
| 2024 | ViTAL: Vision TrAnsformer based Low coverage SARS-CoV-2 lineage assignmentabstractMOTIVATION: Rapid spread of viral diseases such as Coronavirus disease 2019 (COVID-19) highlights an urgent need for efficient surveillance of virus mutation and transmission dynamics, which requires fast, inexpensive and accurate viral lineage assignment. The first two goals might be achieved through low-coverage whole-genome sequencing (LC-WGS) which enables rapid genome sequencing at scale and at reduced costs. Unfortunately, LC-WGS significantly diminishes the genomic details, rendering accurate lineage assignment very challenging. RESULTS: We present ViTAL, a novel deep learning algorithm specifically designed to perform lineage assignment of low coverage-sequenced genomes. ViTAL utilizes a combination of MinHash for genomic feature extraction and Vision Transformer for fine-grain genome classification and lineage assignment. We show that ViTAL outperforms state-of-the-art tools across diverse coverage levels, reaching up to 87.7% lineage assignment accuracy at 1× coverage where state-of-the-art tools such as UShER and Kraken2 achieve the accuracy of 5.4% and 27.4% respectively. ViTAL achieves comparable accuracy results with up to 8× lower coverage than state-of-the-art tools. We explore ViTAL's ability to identify the lineages of novel genomes, i.e. genomes the Vision Transformer was not trained on. We show how ViTAL can be applied to preliminary phylogenetic placement of novel variants. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available in https://github.com/zuherJahshan/vital and can be accessed with 10.5281/zenodo.10688110. Zuher Jahshan, Leonid Yavits |
Bioinform. | 2 |
| 2024 | DIPER: Detection and Identification of Pathogens Using Edit Distance-Tolerant Resistive CAMabstractWe propose a novel resistive edit distance-tolerant content addressable memory for computational genomics applications, particularly for detection and identification of pathogens of pandemic importance. Unlike state-of-the-art approximate search solutions that tolerate small number of replacements between the query pattern and the stored data, DIPER tolerates insertions and deletions, ubiquitous in genomics. DIPER achieves up to 1.7× higherF1score for high-quality DNA reads and up to 6.2× higherF1score for DNA reads with 15% error rate, compared to state-of-the-art DNA classification tool Kraken2. Simulated at 500MHz, DIPER provides 910× average speedup over Kraken2. Itay Merlin, Esteban Garzón, Alexander Fish, Leonid Yavits |
IEEE Trans. Computers | 4 |
| 2024 | Designing Precharge-Free Energy-Efficient Content-Addressable MemoriesabstractContent-addressable memory (CAM) is a specialized type of memory that facilitates massively parallel comparison of a search pattern against its entire content. State-of-the-art (SOTA) CAM solutions are either fast but power-hungry (NOR CAM) or slow while consuming less power (nand CAM). These limitations stem from the dynamic precharge operation, leading to excessive power consumption in NOR CAMs and charge-sharing issues in NAND CAMs. In this work, we propose a precharge-free CAM (PCAM) class for energy-efficient applications. By avoiding precharge operation, PCAM consumes less energy than a NAND CAM, while achieving search speed comparable to a NOR CAM. PCAM was designed using a 65-nm CMOS technology and comprehensively evaluated under extensive Monte Carlo (MC) simulations while taking into account layout parasitics. When benchmarked against conventional NAND CAM, PCAM demonstrates improved search run time (reduced by more than 30%) and 15% less search energy. Moreover, PCAM can cut energy consumption by more than 75% when compared to conventional NOR CAM. We further extend our analysis to the application level, functionally evaluating the CAM designs as a fully associative cache using a CPU simulator running various benchmark workloads. This analysis confirms that PCAMs represent an optimal energy-performance design choice for associative memories and their broad spectrum of applications. Ramiro Taco, Esteban Garzón, Robert Hanhan, Adam Teman, Leonid Yavits, Marco Lanuzza |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | DASH-CAM: Dynamic Approximate SearcH Content Addressable Memory for genome classificationabstractWe propose a novel dynamic storage-based approximate search content addressable memory (DASH-CAM) for computational genomics applications, particularly for identification and classification of viral pathogens of epidemic significance. DASH-CAM provides 5.5 × better density compared to state-of-the-art SRAM-based approximate search CAM. This allows using DASH-CAM as a portable classifier that can be applied to pathogen surveillance in low-quality field settings during pandemics, as well as to pathogen diagnostics at points of care. DASH-CAM approximate search capabilities allow a high level of flexibility when dealing with a variety of industrial sequencers with different error profiles. DASH-CAM achieves up to 30% and 20% higher F1 score when classifying DNA reads with 10% error rate, compared to state-of-the-art DNA classification tools MetaCache-GPU and Kraken2 respectively. Simulated at 1GHz, DASH-CAM provides 1, 178 × and 1, 040 × average speedup over MetaCache-GPU and Kraken2 respectively. Zuher Jahshan, Itay Merlin, Esteban Garzón, Leonid Yavits |
MICRO | 4 |
| 2023 | ClaPIM: Scalable Sequence Classification Using Processing-in-MemoryabstractDeoxyribonucleic acid (DNA) sequence classification is a fundamental task in computational biology with vast implications for applications such as disease prevention and drug design. Therefore, fast high-quality sequence classifiers are significantly important. This article introduces ClaPIM, a scalable DNA sequence classification architecture based on the emerging concept of hybrid in-crossbar and near-crossbar memristive processing-in-memory (PIM). We enable efficient and high-quality classification by uniting the filter and search stages within a single algorithm. Specifically, we propose a custom filtering technique that drastically narrows the search space and a search approach that facilitates approximate string matching through a distance function. ClaPIM is the first PIM architecture for scalable approximate string matching that benefits from the high density of memristive crossbar arrays and the massive computational parallelism of PIM. Compared with Kraken2, a state-of-the-art software classifier, ClaPIM provides significantly higher classification quality (up to$20 \times $improvement in F1 score) and also demonstrates a$1.8 \times $throughput improvement. Compared with edit distance tolerant approximate matching (EDAM), a recently proposed static random-access memory (SRAM)-based accelerator that is restricted to small datasets, we observe both a$30.4 \times $improvement in normalized throughput per area and a 7% increase in classification precision. Marcel Khalifa, Barak Hoffer, Orian Leitersdorf, Robert Hanhan, Ben Perach, Leonid Yavits, Shahar Kvatinsky |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | EDAM: edit distance tolerant approximate matching content addressable memoryabstractWe propose a novel edit distance-tolerant content addressable memory (EDAM) for energy-efficient approximate search applications. Unlike state-of-the-art approximate search solutions that tolerate certain Hamming distance between the query pattern and the stored data, EDAM tolerates edit distance, which makes it especially efficient in applications such as text processing and genome analysis. EDAM was designed using a commercial 65 nm 1.2 V CMOS technology and evaluated through extensive Monte Carlo simulations, while considering different process corners. Simulation results show that EDAM can achieve robust approximate search operation with a wide range of edit distance threshold levels. EDAM is functionally evaluated as a pathogen DNA detection and classification accelerator. EDAM achieves up to 1.7× higher F1 score for high-quality DNA reads and up to 19.55× higher F1 score for DNA reads with 15% error rate, compared to state-of-the-art DNA classification tool Kraken2. Simulated at 667 MHz, EDAM provides 1, 214× average speedup over Kraken2. This makes EDAM suitable for hardware acceleration of genomic surveillance of outbreaks, such as the ongoing Covid-19 pandemic. Robert Hanhan, Esteban Garzón, Zuher Jahshan, Adam Teman, Marco Lanuzza, Leonid Yavits |
ISCA | 6 |
| 2022 | GIRAF: General Purpose In-Storage Resistive Associative FrameworkabstractGIRAF is a General purpose In-storage Resistive Associative Framework based on resistive content addressable memory (RCAM), which functions simultaneously as a storage and a massively parallel associative processor. GIRAF alleviates the bandwidth wall by connecting every memory bit to processing transistors and keeping computing inside the storage arrays, thus implementing deep in-data, rather than near-data, processing. We show that GIRAF outperformed a reference computer architecture with a bandwidth-limited external storage access on a variety of data-intensive workloads. The performance of GIRAF Dot Product and Sparse Matrix-Vector multiplication exceeds the attainable performance of a reference architecture by 1200 × and 130 ×, respectively. Leonid Yavits, Roman Kaplan, Ran Ginosar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | WoLFRaM: Enhancing Wear-Leveling and Fault Tolerance in Resistive Memories using Programmable Address DecodersabstractResistive memories have limited lifetime caused by limited write endurance and highly non-uniform write access patterns. Two main techniques to mitigate endurance-related memory failures are 1) wear-leveling, to evenly distribute the writes across the entire memory, and 2) fault tolerance, to correct memory cell failures. However, one of the main open challenges in extending the lifetime of existing resistive memories is to make both techniques work together seamlessly and efficiently. To address this challenge, we propose WoLFRaM, a new mechanism that combines both wear-leveling and fault tolerance techniques at low cost by using a programmable resistive address decoder (PRAD). The key idea of WoLFRaM is to use PRAD for implementing 1) a new efficient wear-leveling mechanism that remaps write accesses to random physical locations on the fly, and 2) a new effiCient fault tolerance mechanism that recovers from faults by remapping failed memory blocks to available physical locations. Our evaluations show that, for a Phase Change Memory (PCM) based system with cell endurance of 108writes, WoLFRaM increases the memory lifetime by 68% compared to a baseline that implements the best state-of-the-art wear-leveling and fault correction mechanisms. WoLFRaM's average / worst-case performance and energy overheads are 0.51% /3.8% and 0.47% /2.1% respectively. Leonid Yavits, Lois Orosa 0001, Suyash Mahar, João Dinis Ferreira, Mattan Erez, Ran Ginosar, Onur Mutlu |
ICCD | 1 |
| 2020 | Robust Dual Mode Pass Logic (DMPL) for Energy Efficiency and High PerformanceabstractIn the past, Pass Transistor Logic (PTL) was widely used due to benefits in terms of speed and power consumption coming from the reduced number of transistors. However, issues such as threshold drop across the single-channel pass transistors and high sensitivity to process variations have prevented the use of PTL in advanced nanometer technologies. In this paper, we propose a novel logic family named Dual Mode Pass Logic (DMPL), which allows for high speed and low power consumption while maintaining robustness down to the sub-threshold voltage region. The DMPL effectively combines PTL to reduce energy and power consumption along with the flexibility of Dual Mode Logic (DML) to switch to a speed improved operating mode according to the system requirement. Simulation analysis performed on basic NOR/NAND gates implemented in 16 nm Finfet technology demonstrates that DMPL can reduce energy and power by 33% and 42% as compared to logically equivalent static CMOS design. Moreover, running frequency of a DMPL circuit can exceed that of its static CMOS counterpart by 84% when speed is mandatory. Additionally, DMPL gates demonstrate similar robustness as static CMOS implementations under process and temperature variations at lower supply voltages. Inbal Stanger, Netanel Shavit, Ramiro Taco, Leonid Yavits, Marco Lanuzza, Alexander Fish |
ISCAS | 4 |
| 2020 | Exploiting Single-Well Design for Energy-Efficient Ultra-Wide Voltage Range Dual Mode Logic-Based Digital Circuits in 28nm FD-SOI TechnologyabstractIn this paper we evaluate the implementation options of energy-efficient dual mode logic (DML) circuits in 28nm fully depleted silicon-on-insulator (FD-SOI) technology. The combination of the flexibility of Dual Mode Logic (DML) and the unique characteristics of the FD-SOI technology has enormous potential to design energy-efficient adaptive digital circuits operating on an ultra-wide voltage range. As a main result, we demonstrate that single well option offered by the FD-SOI greatly extends the low-granularity energy-delay (E-D) optimization capability of DML-based designs. By exploiting the above implementation strategy, a 16-bit DML carry skip adder reduces its energy consumption by 41% and increases its speed of about 26% when changing its operation mode (from static to dynamic) at 0.4V as compared to its equivalent standard CMOS design. Ramiro Taco, Leonid Yavits, Netanel Shavit, Inbal Stanger, Marco Lanuzza, Alexander Fish |
ISCAS | 2 |
| 2020 | Dual Mode Logic Address DecoderabstractAddress decoders are integral components of random access memories. In higher-performance computing, the timing of address decoders is often critical, especially in applications such as translation lookaside buffer (TLB) and first level data cache. On the other hand, memory power budget and energy consumption are equally critically important for battery-powered devices. Dual Mode Logic (DML) has been shown to combine the support for both requirements in a single circuit. We present a novel DML based address decoder design and compare it with conventional static CMOS and np-CMOS address decoders. Simulations show that DML based address decoder in dynamic mode achieves 31% lower delay compared to conventional static CMOS implementation. In static mode, DML based address decoder reduces the energy consumption by 29% and reaches 10% lower energy-delay product compared to static CMOS address decoder. This is the first time DML is evaluated in 16nm FinFet process. Leonid Yavits, Ramiro Taco, Netanel Shavit, Inbal Stanger, Alexander Fish |
ISCAS | 1 |
| 2020 | BioSEAL: In-Memory Biological Sequence Alignment Accelerator for Large-Scale Genomic DataabstractGenome sequences contain hundreds of millions of DNA base pairs. Finding the degree of similarity between two genomes requires executing a compute-intensive dynamic programming algorithm, such as Smith-Waterman. Traditional von Neumann architectures have limited parallelism and cannot provide an efficient solution for large-scale genomic data. Approximate heuristic methods (e.g. BLAST) are commonly used. However, they are suboptimal and still compute-intensive. Roman Kaplan, Leonid Yavits, Ran Ginosar |
SYSTOR | 2 |
| 2019 | POSTER: BioSEAL: In-Memory Biological Sequence Alignment Accelerator for Large-Scale Genomic DataabstractGenome sequences contain hundreds of millions of DNA base pairs. Finding the degree of similarity between two genomes requires executing a compute-intensive dynamic programming algorithm, such as Smith-Waterman. Traditional von Neumann architectures have limited parallelism and cannot provide an efficient solution for large-scale genomic data. Approximate heuristic methods (e.g. BLAST) are commonly used. However, they are suboptimal and still compute-intensive. In this work, we present BioSEAL, a biological sequence alignment accelerator. BioSEAL is a massively parallel non-von Neumann processing-in-memory architecture for large-scale DNA and protein sequence alignment. BioSEAL is based on resistive content addressable memory, capable of energy-efficient and high-performance associative processing. We present an associative processing algorithm for entire database sequence alignment on BioSEAL and compare its performance and power consumption with state-of-art solutions. We show that BioSEAL can achieve up to 57x speedup and 156x better energy efficiency, compared with existing solutions for genome sequence alignment and protein sequence database search. Roman Kaplan, Leonid Yavits, Ran Ginosar |
PACT | 2 |
| 2019 | POSTER: GIRAF: General Purpose In-Storage Resistive Associative FrameworkabstractGIRAF is an in-storage architecture and algorithm framework based on Resistive Content Addressable Memory (RCAM). GIRAF functions simultaneously as a storage and a massively parallel associative processor. GIRAF alleviates the bandwidth wall by connecting every memory bit to processing transistors and keeping computing inside the storage arrays, thus implementing in-data, rather than near-data, processing. We show that GIRAF outperforms a reference computer architecture with a bandwidth-limited external storage access on a variety of data intensive workloads. The performance of GIRAF Euclidean distance, dot product and histogram implementation, exceeds the attainable performance of a reference architecture by up to four orders of magnitude, depending on the dataset size. The performance of GIRAF SpMV exceeds the attainable performance of such reference architecture by more than two orders of magnitude. Leonid Yavits, Roman Kaplan, Ran Ginosar |
PACT | 1 |
| 2016 | Resistive GP-SIMD Processing-In-MemoryabstractGP-SIMD, a novel hybrid general-purpose SIMD architecture, addresses the challenge of data synchronization by in-memory computing, through combining data storage and massive parallel processing. In this article, we explore a resistive implementation of the GP-SIMD architecture. In resistive GP-SIMD, a novel resistive row and column addressable 4F 2 crossbar is utilized, replacing the modified CMOS 190F 2 SRAM storage previously proposed for GP-SIMD architecture. The use of the resistive crossbar allows scaling the GP-SIMD from few millions to few hundred millions of processing units on a single silicon die. The performance, power consumption and power efficiency of a resistive GP-SIMD are compared with the CMOS version. We find that PiM architectures and, specifically, GP-SIMD benefit more than other many-core architectures from using resistive memory. A framework for in-place arithmetic operation on a single multivalued resistive cell is explored, demonstrating a potential to become a building block for next-generation PiM architectures. Amir Morad, Leonid Yavits, Shahar Kvatinsky, Ran Ginosar |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | The Effect of Temperature on Amdahl Law in 3D Multicore EraabstractThis work studies the influence of temperature on performance and scalability of 3D Chip Multiprocessors (CMP) from Amdahl's law perspective. We find that 3D CMP may reach its thermal limit before reaching its maximum power. We show that a high level of parallelism may lead to high peak temperatures even in small scale 3D CMPs, thus limiting 3D CMP scalability and calling for different, in-memory computing architectures. Leonid Yavits, Amir Morad, Ran Ginosar |
IEEE Trans. Computers | 1 |
| 2015 | Computer Architecture with Associative Processor Replacing Last-Level Cache and SIMD AcceleratorabstractThis study presents a computer architecture, where a last-level cache and a SIMD accelerator are replaced by an associative processor. Associative processor combines data storage and data processing, and functions as a massively parallel SIMD processor and a memory at the same time. An analytic performance model of this computer architecture is introduced. Comparative analysis supported by cycle-accurate simulation and emulation shows that this architecture may outperform a conventional computer architecture comprising a SIMD coprocessor and a shared last-level cache while consuming less power. Leonid Yavits, Amir Morad, Ran Ginosar |
IEEE Trans. Computers | 1 |
| 2015 | Sparse Matrix Multiplication On An Associative ProcessorabstractSparse matrix multiplication is an important component of linear algebra computations. Implementing sparse matrix multiplication on an associative processor (AP) enables high level of parallelism, where a row of one matrix is multiplied in parallel with the entire second matrix, and where the execution time of vector dot product does not depend on the vector size. Four sparse matrix multiplication algorithms are explored in this paper, combining AP and baseline CPU processing to various levels. They are evaluated by simulation on a large set of sparse matrices. The computational complexity of sparse matrix multiplication on AP is shown to be an O(nnz) where nnz is the number of nonzero elements. The AP is found to be especially efficient in binary sparse matrix multiplication. AP outperforms conventional solutions in power efficiency. Leonid Yavits, Amir Morad, Ran Ginosar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | The effect of communication and synchronization on Amdahl's law in multicore systems
Leonid Yavits, Amir Morad, Ran Ginosar |
Parallel Comput. | 1 |
| 2014 | GP-SIMD Processing-in-MemoryabstractGP-SIMD, a novel hybrid general-purpose SIMD computer architecture, resolves the issue of data synchronization by in-memory computing through combining data storage and massively parallel processing. GP-SIMD employs a two-dimensional access memory with modified SRAM storage cells and a bit-serial processing unit per each memory row. An analytic performance model of the GP-SIMD architecture is presented, comparing it to associative processor and to conventional SIMD architectures. Cycle-accurate simulation of four workloads supports the analytical comparison. Assuming a moderate die area, GP-SIMD architecture outperforms both the associative processor and conventional SIMD coprocessor architectures by almost an order of magnitude while consuming less power. Amir Morad, Leonid Yavits, Ran Ginosar |
ACM Trans. Archit. Code Optim. | 2 |