EDBT 2026 Demo / reviewers in the wild / expert
Yehuda Kra
dblp:70/5547
· DBLP profile ↗
9ranked-venue papers
6as first author
6since 2021 · last 2026
0009-0002-3377-6529ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenMClass: Design and comparative analysis of genome classifier-on-chip platformabstractWe propose GenMClass, a genome classification system-on-chip (SoC) implementing two different classification approaches and comprising two separate classification engines: a DNN accelerator GenDNN, that classifies DNA reads converted to images using a classification neural network, and a similarity search-capable Error Tolerant Content Addressable Memory ETCAM, that classifies genomes by k-mer matching. Classification operations are controlled by an embedded RISCV processor. GenMClass classification platform was designed and manufactured in a commercial 65 nm process. We conduct a comparative analysis of ETCAM and GenDNN classification efficiency as well as their performance, silicon area and power consumption using silicon measurements. The size of GenMClass SoC is 3.4 mm 2 and its total power consumption (assuming both GenDNN and ETCAM perform classification at the same time) is 144 mW. This allows using GenMClass as a portable classifier for pathogen surveillance during pandemics, food safety and environmental monitoring, agriculture pathogen and antimicrobial resistance control, in the field or at points of care. Daria Bromot, Yehuda Kra, Zuher Jahshan, Esteban Garzón, Adam Teman, Leonid Yavits |
J. Syst. Archit. | 2 |
| 2024 | Selfie5: An Autonomous, Self-Contained Verification Approach for High-Throughput Random Testing of Programmable ProcessorsabstractRandom testing plays a crucial role in processor designs, complementing other verification methodologies. This paper introduces Selfie5, an autonomous, self-contained verification approach that utilizes the device under verification (DUV) itself to generate, execute, and verify random sequences. This approach eliminates the overhead associated with testing environment interfaces, resulting in a substantial increase in throughput, a critical aspect for achieving comprehensive coverage. The utility can be deployed to FPGA prototypes, emulation platforms and fabricated ASICs and run at-speed to execute billions of tested scenarios per hour, while ensuring the reproducibility of captured failures in an observable simulation environment. This paper describes the Selfie5 approach, algorithms and utility, while also providing detailed insights into successful deployment of the utility for a RISC-V implementation. When deployed on a 16 nm test SoC featuring a RISC-V processor, Selfie5 delivered a testing throughput of 13.8 billion tested instructions per hour, which is$69\times$higher than other published works. Yehuda Kra, Naama Kra, Adam Teman |
DATE | 1 |
| 2024 | HAMSA-DI: A Low-Power Dual-Issue RISC-V Core Targeting Energy-Efficient Embedded SystemsabstractThe RISC-V architecture has recently emerged as a popular open source option for the design of general purpose cores with a wide spectrum of operating specifications. In this paper, we present HAMSA-DI, a small footprint, energy-efficient, embedded RISC-V core, featuring a dynamically scheduled, in-order, dual-issue processing pipeline, supporting the popular Xpulp extensions. The proposed cost-effective dual-issue implementation provides a significant performance boost and improved energy-efficiency over baseline low-power cores under common benchmarks. These include a CoreMark score of 3.48 CM/MHz (+22%) and an Embench score of 1.3 (+13%) with certain benchmarks displaying as much as 22% less energy than the baseline CV32E40P core. The proposed design was fabricated as part of a 16nm test chip, running at 1GHz with an 0.8V supply voltage. Silicon measurements demonstrate that the proposed core can improve performance by as much as 8$\times $for programs operating with full dual-issue utilization with energy-efficiency improving by as much as 6.5$\times $, as compared to compiled code on a single-issue core. Yehuda Kra, Yonatan Shoshan, Yehuda Rudin, Adam Teman |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | A RISC-V-based Research Platform for Rapid Design CycleabstractThis work proposes a novel platform for bringing a project from the concept to the tapeout stage in a short amount of time. An open-source and extendable RISC-V architecture is exploited to build a small area footprint core. This leads the research platform to be flexible in terms of design integration, while also allowing fast design cycles of research chips. Esteban Garzón, Roman Golman, Odem Harel, Tzachi Noy, Yehuda Kra, Asaf Pollock, Slava Yuzhaninov, Yonatan Shoshan, Yehuda Rudin, Yoav Weizman, Marco Lanuzza, Adam Teman |
ISCAS | 5 |
| 2022 | Silicon-Proven Clockless Wave-Propagated Pipelining for High-Throughput, Energy-Efficient ProcessingabstractThe vast majority of digital systems are designed using pipelined sequential logic, thanks to a well-known and robust implementation flow with the ability to increase throughput simply by introducing intermediate sampling stages. However, adding these registers results in significant area and power overheads. Clockless Wave-Propagated Pipelining (CWPP) is a design approach that reaches high throughputs without the need for intermediate sampling registers. As opposed to traditional sequential design, which increases frequency by minimizing the longest delay through a combinational path, the performance of a CWPP scheme is set according to the difference between the longest and shortest paths through the logic, as captured by the following constraint [1]: Yehuda Kra, Adam Teman |
ISCAS | 1 |
| 2021 | WP 2.0: Signoff-Quality Implementation and Validation of Energy-Efficient Clock-Less Wave Propagated PipeliningabstractThe design of computational datapaths with the clockless wave-propagated pipelining (CWPP) approach is an area and energy-efficient alternative to traditional pipelined logic. Removal of the internal registers saves both area and the toggling power of these complex gates, while also simplifying the clock tree. However, this approach is rarely used in modern scaled technologies, due to the complexity of implementation and the lack of a robust, scalable, and automated design methodology that meets rigid industry standards. In this paper, we present WP 2.0, an extension of the original WP algorithm and automation utility, which demonstrated how to apply CWPP to any generic combinatorial circuit using a CMOS standard cell library. WP 2.0 advances this concept to provide full-flow implementation capabilities, providing a post-layout CWPP-ready design that meets signoff-quality industry timing requirements. The WP 2.0 utility interfaces with commercial design automation software for balancing a post-synthesis netlist to achieve a high CWPP launch rate (frequency). We demonstrate the calculation of an fused dot-product accumulation unit, implemented with a 65nm standard cell library, providing a worst-case launch rate that is comparable to a design implemented with a 3-stage clocked pipeline with a 12 % area reduction and between 37 % -54 % power savings. Furhtermore, the CWPP design is equipped with unique post-silicon field configuration capabilities for optimizing operation and overcoming variation. Yehuda Kra, Tzachi Noy, Adam Teman |
DATE | 1 |
| 2020 | WavePro: Clock-less Wave-Propagated Pipeline Compiler for Low-Power and High-Throughput ComputationabstractClock-less Wave-Propagated Pipelining is a long-known approach to achieve high-throughput without the over-head of costly sampling registers. However, due to many design challenges, which have only increased with technology scaling, this approach has never been widely accepted and has generally been limited to small and very specific demonstrations. This paper addresses this barrier by presenting WavePro, a generic and scalable algorithm, capable of skew balancing any combinatorial logic netlist for the application of wave-pipelining. The algorithm was implemented in the WavePro Compiler automation utility, which interfaces with industry delays extraction and standard timing analysis tools to produce a sign-off quality result. The utility is demonstrated upon a dot-product accelerator in a 65 nm CMOS technology, using a vendor-provided standard cell library and commercial timing analysis tools. By reducing the worst-case output skew by over 70%, the test case example was able to achieve equivalent throughput of an 8-staged sequentially pipelined implementation with power savings of almost 3×. Yehuda Kra, Tzachi Noy, Adam Teman |
DATE | 1 |
| 2020 | Physically Aware Affinity-Driven Multiplier ImplementationabstractOptimized hardware for the execution of large dot-product (DP) calculations is central to many of today's integrated circuits. These arithmetic blocks are often implemented with the parallel fused DP (FDP) approach, and to achieve high performance, are realized with a tree-based compression algorithm, using on commercially available synthesis macros. However, these macros are based on performance optimization of the gate-level netlist, and fail to take into account the consequences of the applied heuristics on the physical-implementation (layout) of these large circuits. In this article, we propose a physical-aware approach to FDP implementation based on the affinity between the logic gates that make up the gate-level structure. The proposed clustered DP (CDP) algorithm, enables the place and route tools to cluster gates with high-affinity, leading to higher placement utilization and lower routing congestion. DP calculations with up to 78 multipliers were implemented with a 65-nm CMOS standard cell library, providing power reduction of up to 63%, up to 60% lower area, and performance improvements as high as 2.5×, as compared to similar implementations based on commercial macros based on post-layout results. Or Maltabashi, Yehuda Kra, Adam Teman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1993 | A Cross-Debugging Method for Hardware/Software Co-design EnvironmentsabstractSof2ware debugging in hardware simulation environments is often inconvenient and extremely slow.The BackC method described in this paper allows one to reproduce accurately the simulated code flow under a conventional pure software debugging environment, while greatly decreasing the test turnaround time, thereby dmstically reducing debugging time.The described method has been successfully implemented and used during a chip hardware and embedded soflware co-design.It can be applied to any conventional simulation environment used for co-design verification. Yehuda Kra |
DAC | 1 |