Adam Teman

dblp:92/4765 · DBLP profile ↗
← Back
46ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-8233-4711ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46 · 5 first-author · 17 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 GenMClass: Design and comparative analysis of genome classifier-on-chip platform
abstract
We propose GenMClass, a genome classification system-on-chip (SoC) implementing two different classification approaches and comprising two separate classification engines: a DNN accelerator GenDNN, that classifies DNA reads converted to images using a classification neural network, and a similarity search-capable Error Tolerant Content Addressable Memory ETCAM, that classifies genomes by k-mer matching. Classification operations are controlled by an embedded RISCV processor. GenMClass classification platform was designed and manufactured in a commercial 65 nm process. We conduct a comparative analysis of ETCAM and GenDNN classification efficiency as well as their performance, silicon area and power consumption using silicon measurements. The size of GenMClass SoC is 3.4 mm 2 and its total power consumption (assuming both GenDNN and ETCAM perform classification at the same time) is 144 mW. This allows using GenMClass as a portable classifier for pathogen surveillance during pandemics, food safety and environmental monitoring, agriculture pathogen and antimicrobial resistance control, in the field or at points of care.
Daria Bromot, Yehuda Kra, Zuher Jahshan, Esteban Garzón, Adam Teman, Leonid Yavits
J. Syst. Archit.5
2026 SPARCAM: Sparse matrix multiplication accelerator using multi-port dynamic CAM
abstract
Sparse General matrix multiplication (SpGEMM) is a fundamental kernel in many scientific and engineering fields, including Artificial Intelligence (AI). However, its intrinsic computation complexity presents substantial challenges, making efficient hardware implementation particularly difficult. This paper proposes SPARCAM, a novel SpGEMM accelerator, developed and optimized for very energy-efficient AI edge applications. SPARCAM is designed using low-power dense Gain Cell embedded DRAM (GC-eDRAM) technology, a processing near memory paradigm, and a modified outer product matrix multiplication algorithm. Despite its quite limited peak theoretical performance, SPARCAM achieves very high energy efficiency due to its low-power architecture and almost 100% utilization of its computing resources. Designed in a commercial 28 nm FDSOI technology, SPARCAM achieves 13 . 9 × speedup over a high-performance embedded CPU when processing large-scale sparse matrices. When multiplying limited-size sparse matrices, SPARCAM obtains 193 × speedup over high-performance GPU. SPARCAM reaches about 4.3 orders-of-magnitude, on average, higher energy benefits, and 1892 × , 181 × , 2 × , and 3471 × , higher energy efficiency (over CPU) compared with state-of-the-art SpGEMM accelerators SpArch, OuterSPACE, MatRaptor, and high-performance GPU, respectively.
Esteban Garzón, Benjamin Zambrano, David Sheinenzon, Marco Lanuzza, Adam Teman, Leonid Yavits
J. Syst. Archit.5
2026 GainP: A Gain Cell Embedded DRAM-based associative in-memory processor
abstract
Associative processors (APs) are massively-parallel in-memory SIMD accelerators. While fairly well-known, APs have been revisited in recent years due to the proliferation of data-centric computing, and specifically, processing using memory. APs are based on Content Addressable Memory and utilize its unique ability to simultaneously search the entire memory content for a query pattern to implement massively parallel computations in memory. Several memory infrastructures have been considered for associative processing, including static CMOS, resistive, magnetoresistive, ferroelectric and even NAND flash memories. While all of these have certain merits (speed and low energy consumption for static CMOS, density for resistive and ferroelectric memories), they also face challenges (low density for static CMOS and magnetoresistive, limited write endurance and high write energy for resistive and ferroelectric memories), which limit the scalability and usefulness of APs. This work introduces GainP, an AP based on silicon-proven Gain Cell embedded DRAM (GCeDRAM). The latter combines relatively high density (compared to static CMOS memory) with low energy, high speed, practically unlimited endurance and low production costs (compared to emerging memory technologies). Using sparse-by-sparse matrix multiplication, we show that GainP outperforms high-performance CPU and GPU by 825 × and 41 × . We also show that GainP outperforms state-of-the-art processing-in-memory sparse matrix multiplication accelerators GAS, OuterSPACE and MatRaptor by 128 × , 125 × and 16 × , respectively, and provides average energy benefits of 96 × , 95 × and 15 × , respectively.
Yaniv Levi, Odem Harel, Adam Teman, Leonid Yavits
J. Syst. Archit.3
2026 A Cryogenic 2T GC-eDRAM With an Inverter-Based Readout Scheme and an 80 ms Retention Time
abstract
This article presents a novel sensing scheme for a 2T gain-cell embedded dynamic random access memory (GC-eDRAM) optimized for cryogenic temperatures. The proposed design integrates a highly accurate sense amplifier (SA) with a$\pm 3\sigma $sensitivity of 11.16 mV and an internally generated, adjustable reference voltage (VREF) with 9.5 mV step resolution. Additionally, an on-chip ring-oscillator-based design-for-test (RO-DFT) mechanism is introduced to determine the optimal VREF postfabrication, improving readout reliability while adapting to process variations. Simulation results, based on a 65nm CMOS process calibrated for 77 K operation, demonstrate a significant extension of the data retention time (DRT) to 80 ms, reducing refresh cycle frequency and lowering power consumption—critical for cryogenic applications such as infrared imaging. The proposed architecture provides a power-efficient, high-density embedded memory solution for low-temperature computing environments.
Elisheva Berkowitz, Yosi Greenblatt, Claudio G. Jakobson, Adam Teman, Joseph Shor
IEEE Trans. Very Large Scale Integr. Syst.4
2025 LPRE: Logarithmic Posit-enabled Reconfigurable edge-AI Engine
abstract
Edge-AI applications face huge challenges in resource-constrained environments, particularly in enhancing computational efficiency within bandwidth limitations. This work proposes the Logarithmic-Posit-enabled Reconfigurable edgeAI Engine (LPRE) that enhances hardware efficiency without compromising accuracy. The proposed architecture utilizes time-multiplexed dynamically configurable single-layer hardware to balance resource reuse and bandwidth for multi-layer perceptron and CNN models. Evaluations on LeNet-5 using MNIST demonstrate that LPRE achieves up to 4× throughput enhancement at 8-bit precision with negligible accuracy loss (compared to FP32 baseline), while requiring up to 80% and 50% fewer resources than fixed-point arithmetic and state-of-the-art works, respectively. The design is viable for various edge-AI applications, such as real-time number plate recognition, offering scalable, energy-efficient IoT solutions.
Omkar Kokane, Mukul Lokhande, Gopal Raut, Adam Teman, Santosh Kumar Vishvakarma
ISCAS4
2025 Towards Low-Power High-Performance Content-Addressable Memory: A Robust Precharge-Free Approach
abstract
Low-Power high-performance content-addressable memories (CAMs) are important components in modern computing systems. In this work, we present a robust CAM that overcomes the power and performance limitations of conventional precharge-based CAMs. The proposed static transmission gate-based (STAT-TG) CAM design achieves low-power operation comparable to NAND CAMs while maintaining search speeds rivaling those of NOR CAMs. The STAT-TG CAM was designed using a 65nm CMOS technology and comprehensively evaluated under extensive Monte Carlo simulations. Compared to conventional CAMs, the STAT-TG CAM is 14% faster than NAND CAM, while consuming only 25% of the energy per operation relative to NOR CAM. This makes STAT-TG CAM a promising solution for high-performance yet energy-efficient applications.
Ramiro Taco, Esteban Garzón, Adam Teman, Leonid Yavits, Marco Lanuzza
ISCAS3
2024 Selfie5: An Autonomous, Self-Contained Verification Approach for High-Throughput Random Testing of Programmable Processors
abstract
Random testing plays a crucial role in processor designs, complementing other verification methodologies. This paper introduces Selfie5, an autonomous, self-contained verification approach that utilizes the device under verification (DUV) itself to generate, execute, and verify random sequences. This approach eliminates the overhead associated with testing environment interfaces, resulting in a substantial increase in throughput, a critical aspect for achieving comprehensive coverage. The utility can be deployed to FPGA prototypes, emulation platforms and fabricated ASICs and run at-speed to execute billions of tested scenarios per hour, while ensuring the reproducibility of captured failures in an observable simulation environment. This paper describes the Selfie5 approach, algorithms and utility, while also providing detailed insights into successful deployment of the utility for a RISC-V implementation. When deployed on a 16 nm test SoC featuring a RISC-V processor, Selfie5 delivered a testing throughput of 13.8 billion tested instructions per hour, which is$69\times$higher than other published works.
Yehuda Kra, Naama Kra, Adam Teman
DATE3
2024 HAMSA-DI: A Low-Power Dual-Issue RISC-V Core Targeting Energy-Efficient Embedded Systems
abstract
The RISC-V architecture has recently emerged as a popular open source option for the design of general purpose cores with a wide spectrum of operating specifications. In this paper, we present HAMSA-DI, a small footprint, energy-efficient, embedded RISC-V core, featuring a dynamically scheduled, in-order, dual-issue processing pipeline, supporting the popular Xpulp extensions. The proposed cost-effective dual-issue implementation provides a significant performance boost and improved energy-efficiency over baseline low-power cores under common benchmarks. These include a CoreMark score of 3.48 CM/MHz (+22%) and an Embench score of 1.3 (+13%) with certain benchmarks displaying as much as 22% less energy than the baseline CV32E40P core. The proposed design was fabricated as part of a 16nm test chip, running at 1GHz with an 0.8V supply voltage. Silicon measurements demonstrate that the proposed core can improve performance by as much as 8$\times $for programs operating with full dual-issue utilization with energy-efficiency improving by as much as 6.5$\times $, as compared to compiled code on a single-issue core.
Yehuda Kra, Yonatan Shoshan, Yehuda Rudin, Adam Teman
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Revisiting Dynamic Logic - A True Candidate for Energy-Efficient Cryogenic Operation in Nanoscaled Technologies
abstract
Dynamic logic is a high-speed technology that was previously used in mature technologies, but lost popularity due to the increased leakage and process variations in advanced technologies. However, the recent popularity of circuits running in the cryogenic region provides a new opportunity for dynamic operation, thanks to the reduced leakages at such low temperatures. This paper revisits dynamic logic as a true candidate for high-performance and energy-efficient circuits for cryogenic operation in nanoscaled technologies. The paper first overviews and analyzes transistor operation at cryogenic temperatures and how it influences digital circuit design targeted to this regime. With these effects in mind, the use of dynamic logic families, including the classical dynamic (NORA) logic and the recently introduced Dual Mode Logic (DML) and Dual Mode Pass Logic (DMPL) families, are examined under cryogenic operation, showcasing improved performance and power efficiency. Measurements conducted on a 16 nm FinFET test chip validate their operation at low temperatures down to 4K, with supply voltages ranging 0.4–0.8-V. Furthermore, the considered dual mode logic families exhibit performance enhancements of up to 26% in dynamic mode and power efficiency increases up to 53% in static mode, compared to CMOS.
Inbal Stanger, Noam Roknian, Netanel Shavit, Yonatan Shoshan, Yoav Weizman, Adam Teman, Edoardo Charbon, Alexander Fish
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 Designing Precharge-Free Energy-Efficient Content-Addressable Memories
abstract
Content-addressable memory (CAM) is a specialized type of memory that facilitates massively parallel comparison of a search pattern against its entire content. State-of-the-art (SOTA) CAM solutions are either fast but power-hungry (NOR CAM) or slow while consuming less power (nand CAM). These limitations stem from the dynamic precharge operation, leading to excessive power consumption in NOR CAMs and charge-sharing issues in NAND CAMs. In this work, we propose a precharge-free CAM (PCAM) class for energy-efficient applications. By avoiding precharge operation, PCAM consumes less energy than a NAND CAM, while achieving search speed comparable to a NOR CAM. PCAM was designed using a 65-nm CMOS technology and comprehensively evaluated under extensive Monte Carlo (MC) simulations while taking into account layout parasitics. When benchmarked against conventional NAND CAM, PCAM demonstrates improved search run time (reduced by more than 30%) and 15% less search energy. Moreover, PCAM can cut energy consumption by more than 75% when compared to conventional NOR CAM. We further extend our analysis to the application level, functionally evaluating the CAM designs as a fully associative cache using a CPU simulator running various benchmark workloads. This analysis confirms that PCAMs represent an optimal energy-performance design choice for associative memories and their broad spectrum of applications.
Ramiro Taco, Esteban Garzón, Robert Hanhan, Adam Teman, Leonid Yavits, Marco Lanuzza
IEEE Trans. Very Large Scale Integr. Syst.4
2022 EDAM: edit distance tolerant approximate matching content addressable memory
abstract
We propose a novel edit distance-tolerant content addressable memory (EDAM) for energy-efficient approximate search applications. Unlike state-of-the-art approximate search solutions that tolerate certain Hamming distance between the query pattern and the stored data, EDAM tolerates edit distance, which makes it especially efficient in applications such as text processing and genome analysis. EDAM was designed using a commercial 65 nm 1.2 V CMOS technology and evaluated through extensive Monte Carlo simulations, while considering different process corners. Simulation results show that EDAM can achieve robust approximate search operation with a wide range of edit distance threshold levels. EDAM is functionally evaluated as a pathogen DNA detection and classification accelerator. EDAM achieves up to 1.7× higher F1 score for high-quality DNA reads and up to 19.55× higher F1 score for DNA reads with 15% error rate, compared to state-of-the-art DNA classification tool Kraken2. Simulated at 667 MHz, EDAM provides 1, 214× average speedup over Kraken2. This makes EDAM suitable for hardware acceleration of genomic surveillance of outbreaks, such as the ongoing Covid-19 pandemic.
Robert Hanhan, Esteban Garzón, Zuher Jahshan, Adam Teman, Marco Lanuzza, Leonid Yavits
ISCA4
2022 A RISC-V-based Research Platform for Rapid Design Cycle
abstract
This work proposes a novel platform for bringing a project from the concept to the tapeout stage in a short amount of time. An open-source and extendable RISC-V architecture is exploited to build a small area footprint core. This leads the research platform to be flexible in terms of design integration, while also allowing fast design cycles of research chips.
Esteban Garzón, Roman Golman, Odem Harel, Tzachi Noy, Yehuda Kra, Asaf Pollock, Slava Yuzhaninov, Yonatan Shoshan, Yehuda Rudin, Yoav Weizman, Marco Lanuzza, Adam Teman
ISCAS12
2022 Silicon-Proven Clockless Wave-Propagated Pipelining for High-Throughput, Energy-Efficient Processing
abstract
The vast majority of digital systems are designed using pipelined sequential logic, thanks to a well-known and robust implementation flow with the ability to increase throughput simply by introducing intermediate sampling stages. However, adding these registers results in significant area and power overheads. Clockless Wave-Propagated Pipelining (CWPP) is a design approach that reaches high throughputs without the need for intermediate sampling registers. As opposed to traditional sequential design, which increases frequency by minimizing the longest delay through a combinational path, the performance of a CWPP scheme is set according to the difference between the longest and shortest paths through the logic, as captured by the following constraint [1]:
Yehuda Kra, Adam Teman
ISCAS2
2022 Evaluation of Dual Mode Logic Under Cryogenic Temperatures
abstract
Dual Mode Logic (DML) enables the dynamical operation of digital circuits optimized for energy-delay efficiency. Here, for the first time, DML is examined under cryogenic conditions, and its characteristics are evaluated for future applications. As a proof-of-concept, a DML testchip designed in 65nm technology was measured under cryogenic temperatures down to 4K. Measurements at supply voltages from 0.8V to 1.2V and temperatures ranging from 300K (room temperature) to 4K, confirm the effectiveness of DML under extreme temperatures.
Inbal Stanger, Noam Roknian, Yonatan Shoshan, Zafrir Levy, Yoav Weizman, Edoardo Charbon, Adam Teman, Alexander Fish
ISCAS7
2021 WP 2.0: Signoff-Quality Implementation and Validation of Energy-Efficient Clock-Less Wave Propagated Pipelining
abstract
The design of computational datapaths with the clockless wave-propagated pipelining (CWPP) approach is an area and energy-efficient alternative to traditional pipelined logic. Removal of the internal registers saves both area and the toggling power of these complex gates, while also simplifying the clock tree. However, this approach is rarely used in modern scaled technologies, due to the complexity of implementation and the lack of a robust, scalable, and automated design methodology that meets rigid industry standards. In this paper, we present WP 2.0, an extension of the original WP algorithm and automation utility, which demonstrated how to apply CWPP to any generic combinatorial circuit using a CMOS standard cell library. WP 2.0 advances this concept to provide full-flow implementation capabilities, providing a post-layout CWPP-ready design that meets signoff-quality industry timing requirements. The WP 2.0 utility interfaces with commercial design automation software for balancing a post-synthesis netlist to achieve a high CWPP launch rate (frequency). We demonstrate the calculation of an fused dot-product accumulation unit, implemented with a 65nm standard cell library, providing a worst-case launch rate that is comparable to a design implemented with a 3-stage clocked pipeline with a 12 % area reduction and between 37 % -54 % power savings. Furhtermore, the CWPP design is equipped with unique post-silicon field configuration capabilities for optimizing operation and overcoming variation.
Yehuda Kra, Tzachi Noy, Adam Teman
DATE3
2021 4T Gain-Cell Providing Unlimited Availability through Hidden Refresh with 1W1R Functionality
abstract
Modern SoCs area and power budgets are often dominated by embedded memories on board of the chip. Gain-cell embedded DRAM is a dense, low power memory solution, supporting low supply voltages; however, it suffers from limited data retention time (DRT) and requires periodic refresh operations, limiting its use only to applications that can tolerate temporary memory blockages. This work presents a novel gain cell design, with robust dual read mechanism, exploiting GC-eDRAM characteristics for double write throughput, supporting low cost hidden refresh mechanism and 100% array availability, providing continuous 1W1R functionality. A 16 kbit memory macro was implemented in 65nm bulk technology offering up- to 20% reduction in bitcell area compared to standard SRAM solution, and up to 3 x area reduction compared to 1R1W memory solutions.
Einat Levy, Aharon Sfez, Roman Golman, Odem Harel, Adam Teman
ISCAS5
2021 Gain-Cell Embedded DRAM Under Cryogenic Operation - A First Study
abstract
Operating circuits under cryogenic conditions is effective for a large spectrum of applications. However, the refrigeration requirement for the cooling of cryogenic systems introduces serious issues in terms of power dissipation. Gain-cell embedded dynamic random access memory (GC-eDRAM) is a low-area, logic-compatible embedded memory alternative to static random access memory (SRAM), which has the potential to provide ultralow-power operation under cryogenic conditions due to the lower leakages at these temperatures. In this article, we present the first comparative design exploration of GC-eDRAM under cryogenic conditions performed with transistor models characterized based on actual silicon measurements under temperatures as low as 77 K. Our study shows that the two-transistor (2T)-based GC-eDRAM configurations turn out to be the best solutions for very low-temperature operation. In particular, the 2T mixed GC-eDRAM configurations allow read sensing margin improvements (up to 99%) within the 2T-based configurations while at the same time excel in terms of data retention time (+44%) and power consumption (-27%) when compared to more complex GC-eDRAM topologies. Moreover, even better improvements in terms of area (-73%), leakage power (-97%), retention power (-76%), and energy (-66%) are observed when compared to conventional 6T-SRAM.
Esteban Garzón, Yosi Greenblatt, Odem Harel, Marco Lanuzza, Adam Teman
IEEE Trans. Very Large Scale Integr. Syst.5
2020 WavePro: Clock-less Wave-Propagated Pipeline Compiler for Low-Power and High-Throughput Computation
abstract
Clock-less Wave-Propagated Pipelining is a long-known approach to achieve high-throughput without the over-head of costly sampling registers. However, due to many design challenges, which have only increased with technology scaling, this approach has never been widely accepted and has generally been limited to small and very specific demonstrations. This paper addresses this barrier by presenting WavePro, a generic and scalable algorithm, capable of skew balancing any combinatorial logic netlist for the application of wave-pipelining. The algorithm was implemented in the WavePro Compiler automation utility, which interfaces with industry delays extraction and standard timing analysis tools to produce a sign-off quality result. The utility is demonstrated upon a dot-product accelerator in a 65 nm CMOS technology, using a vendor-provided standard cell library and commercial timing analysis tools. By reducing the worst-case output skew by over 70%, the test case example was able to achieve equivalent throughput of an 8-staged sequentially pipelined implementation with power savings of almost 3×.
Yehuda Kra, Tzachi Noy, Adam Teman
DATE3
2020 Gain-Cell Embedded DRAMs: Modeling and Design Space
abstract
Among the different types of DRAMs, gain-cell embedded DRAM (GC-eDRAM) is a compact, low-power and CMOS-compatible alternative to conventional SRAM. GC-eDRAM achieves high memory density as it relies on a storage cell that can be implemented with as few as two transistors and that can be fabricated without additional process steps. However, since the performance of GC-eDRAMs relies on many interdependent variables, the optimization of the performance of these memories for the integration into their hosting system, as well as the design investigation of future GC-eDRAMs, prove to be highly complex tasks. In this context, modeling tools of memories are key enablers for the exploration of this large design space in a short amount of time. In this paper, we present GEMTOO, the first modeling tool that estimates timing, memory availability, bandwidth, and area of GC-eDRAMs. The tool considers parameters related to technology, circuits, and memory architecture and it enables the evaluation of architectural transformations as well as of advanced transistor-level effects, such as the increase of the access delay due to deterioration of the stored data. The timing is estimated with a maximum deviation of 15% from post-layout simulations in a 28nm FD-SOI technology for different memory sizes and architectures. Moreover, the measured random cycle frequency of a GC-eDRAM fabricated in 28nm CMOS bulk process is estimated with a 9% deviation when considering 6-sigma random process variations of the bitcells. The proposed GEMTOO modeling tool is used to show the intricacies in design optimization of GC-eDRAMs and, based on the results, optimal design practices are derived.
Andrea Bonetti, Roman Golman, Robert Giterman, Adam Teman, Andreas Peter Burg
ISCAS4
2020 GC-eDRAM with Body-Bias Compensated Readout and Error Detection in 28nm FD-SOI
abstract
Gain-cell embedded DRAM (GC-eDRAM) is an attractive alternative to conventional SRAM due to its high-density, low-leakage, and inherent two-ported functionality. However, its dynamic storage mechanism requires power-hungry refresh cycles to maintain data. This problem is aggravated due to the impact of Process-Voltage-Temperature (PVT) variations at deeply-scaled technology nodes and low voltages. In this paper, we present a GC-eDRAM with body-bias compensated readout, which is dynamically configured to extend the data retention time (DRT) of the memory under varying operating conditions. The proposed GC-eDRAM exploits the body-biasing capabilities of FD-SOI technology to adjust the switching threshold of the sense inverter under PVT variations. An additional, unbiased, sense inverter is added to provide a dual-sampling mechanism to the readout path, enabling error detection to further reduce design guard bands. An 8 kb GC-eDRAM with integrated body-bias compensated readout and error detection was implemented in 28 nm FD-SOI technology. Silicon measurements of the manufactured array demonstrate up-to 75% DRT improvement and up-to 86% energy savings under PVT and frequency variations compared to a conventional guard banded memory design.
Robert Giterman, Andrea Bonetti, Andreas Peter Burg, Adam Teman
ISCAS4
2020 Improved Read Access in GC-eDRAM Memory by Dual-Negative Word-Line Technique
abstract
Embedded memories occupy an increasingly dominant portion of the area and power budgets of modern SoCs and are also a limiting factor in VDDscaling. GC-eDRAM is a dense, low power option for embedded memory implementation, supporting low supply voltages; however, it suffers from limited data retention time (DRT) and requires an additional boosted voltage supply for successful write operations. This work presents a novel technique that uses the same negative voltage applied to the write port in many GC-eDRAMs topologies to expedite the read operation and/or further increase the DRT by using it during read operations. An 8 kbit memory macro was implemented in a 28nm FD-SOI technology, demonstrating over 20× read latency reduction, an order-of-magnitude longer DRT, and up-to 4 order-of-magnitude lower retention power consumption over a conventional 2T GC-eDRAM.
Roman Golman, Robert Giterman, Odem Harel, Adam Teman
ISCAS4
2020 Physically Aware Affinity-Driven Multiplier Implementation
abstract
Optimized hardware for the execution of large dot-product (DP) calculations is central to many of today's integrated circuits. These arithmetic blocks are often implemented with the parallel fused DP (FDP) approach, and to achieve high performance, are realized with a tree-based compression algorithm, using on commercially available synthesis macros. However, these macros are based on performance optimization of the gate-level netlist, and fail to take into account the consequences of the applied heuristics on the physical-implementation (layout) of these large circuits. In this article, we propose a physical-aware approach to FDP implementation based on the affinity between the logic gates that make up the gate-level structure. The proposed clustered DP (CDP) algorithm, enables the place and route tools to cluster gates with high-affinity, leading to higher placement utilization and lower routing congestion. DP calculations with up to 78 multipliers were implemented with a 65-nm CMOS standard cell library, providing power reduction of up to 63%, up to 60% lower area, and performance improvements as high as 2.5×, as compared to similar implementations based on commercial macros based on post-layout results.
Or Maltabashi, Yehuda Kra, Adam Teman
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Gain-Cell Embedded DRAMs: Modeling and Design Space
abstract
Among the different types of dynamic random-access memories (DRAMs), gain-cell embedded DRAM (GC-eDRAM) is a compact, low-power, and CMOS-compatible alternative to conventional static random-access memory (SRAM). GC-eDRAM achieves high memory density, as it relies on a storage cell that can be implemented with as few as two transistors and that can be fabricated without additional process steps. However, since the performance of GC-eDRAMs relies on many interdependent variables, the optimization of the performance of these memories for the integration into their hosting system, as well as the design investigation of future GC-eDRAMs, proves to be highly complex tasks. In this context, modeling tools of memories are key enablers for the exploration of this large design space in a short amount of time. In this article, we present GC-eDRAM modeling tool (GEMTOO), the first modeling tool that estimates timing, memory availability, bandwidth, and area of GC-eDRAMs. The tool considers parameters related to technology, circuits, and memory architecture, and it enables the evaluation of architectural transformations as well as advanced transistor-level effects, such as the increase in the access delay due to the deterioration of the stored data. The timing is estimated with a maximum deviation of 15% from postlayout simulations in a 28-nm FD-SOI technology for different memory sizes and architectures. Moreover, the measured random cycle frequency of a GC-eDRAM fabricated in a 28-nm CMOS bulk process is estimated with a 9% deviation when considering 6-sigma random process variations of the bitcells. The proposed GEMTOO modeling tool is used to show the intricacies in design optimization of GC-eDRAMs, and based on the results, optimal design practices are derived.
Andrea Bonetti, Roman Golman, Robert Giterman, Adam Teman, Andreas Peter Burg
IEEE Trans. Very Large Scale Integr. Syst.4
2018 Live Demonstration: An 800 Mhz Gain-Cell Embedded DRAM in 28 nm CMOS Bulk Process for Approximate Computing Applications
abstract
Gain-cell embedded DRAM (GC-eDRAM) is an attractive alternative to traditional SRAM, due to its high-density, low-leakage, and inherent 2-ported operation, yet, its dynamic nature leads to limited retention time that requires periodic, power-hungry refresh cycles. However, the emerging approximate computing paradigm utilizes the inherent error resilience of some applications to tolerate data errors. Such error tolerance can be exploited by reducing the refresh rate in GC-eDRAM to achieve a substantial decrease in power consumption, at the cost of an increase in cell failure probability. In this demonstration, we present the first fabricated and fully functional GC-eDRAM in a 28 nm bulk CMOS technology. The array, which is based on a novel mixed-VT 4T bitcell, can be used in both traditional and for approximate computing applications, featuring a small silicon footprint and supporting high-performance operation. Silicon measurements demonstrate successful operation at 800 Mhz under a 900 mV supply, while retaining almost 30% lower area than a single-ported 6T SRAM in the same technology.
Robert Giterman, Roman Golman, Amir Shalom, Or Maltabashi, Alexander Fish, Adam Teman
ISCAS6
2018 A 5-Transistor Ternary Gain-Cell eDRAM with Parallel Sensing
abstract
Embedded memories dominate area, power, and cost of modern VLSI system-on-chips. While static random access memory (SRAM) is the dominant technology for implementing these memories, Gain-cell embedded DRAM (GC-eDRAM) has been suggested as a possible alternative in recent years. This technology has been shown to provide low-power, logic compatible storage in a reduced silicon footprint, as compared to conventional SRAM. In this paper we suggest a novel GC-eDRAM topology that is capable of storing three voltage levels within a single cell, further improving upon the area and energy-per-bit of the storage solution. The proposed ternary gain-cell is designed in a standard CMOS 65 nm technology node using a low overhead1/2VDDwrite driver for ternary writes and a parallel sensing scheme composed of skewed sense inverters for ternary readout. The proposed approach provides over 3× reduction in static power with a 48% reduction of area-per-bit in comparison with a conventional SRAM cell in the same technology.
Or Maltabashi, Hanan Marinberg, Robert Giterman, Adam Teman
ISCAS4
2018 A 588-Gb/s LDPC Decoder Based on Finite-Alphabet Message Passing
abstract
An ultrahigh throughput low-density parity-check (LDPC) decoder with an unrolled full-parallel architecture is proposed, which achieves the highest decoding throughput compared to previously reported LDPC decoders in the literature. The decoder benefits from a serial message-transfer approach between the decoding stages to alleviate the well-known routing congestion problem in parallel LDPC decoders. Furthermore, a finite-alphabet message passing algorithm is employed to replace the VN update rule of the standard min-sum (MS) decoder with lookup tables, which are designed in a way that maximizes the mutual information between decoding messages. The proposed algorithm results in an architecture with reduced bit-width messages, leading to a significantly higher decoding throughput and to a lower area compared to an MS decoder when serial message transfer is used. The architecture is placed and routed for the standard MS reference decoder and for the proposed finite-alphabet decoder using a custom pseudo-hierarchical backend design strategy to further alleviate routing congestions and to handle the large design. Postlayout results show that the finite-alphabet decoder with the serial message-transfer architecture achieves a throughput as large as 588 Gb/s with an area of 16.2 mm2and dissipates an average power of 22.7 pJ per decoded bit in a 28-nm fully depleted silicon on isulator library. Compared to the reference MS decoder, this corresponds to 3.1 times smaller area and 2 times better energy efficiency.
Reza Ghanaatian, Alexios Balatsoukas-Stimming, Thomas Christoph Müller, Michael Meidlinger, Gerald Matz, Adam Teman, Andreas Peter Burg
IEEE Trans. Very Large Scale Integr. Syst.6
2017 Automated Integration of Dual-Edge Clocking for Low-Power Operation in Nanometer Nodes
abstract
Clocking power, including both clock distribution and registers, has long been one of the primary factors in the total power consumption of many digital systems. One straightforward approach to reduce this power consumption is to apply dual-edge-triggered (DET) clocking, as sequential elements operate at half the clock frequency while maintaining the same throughput as with conventional single-edge-triggered (SET) clocking. However, the DET approach is rarely taken in modern integrated circuits, primarily due to the perceived complexity of integrating such a clocking scheme. In this article, we first identify the most promising conditions for achieving low-power operation with DET clocking and then introduce a fully automated design flow for applying DET to a conventional SET design. The proposed design flow is demonstrated on three benchmark circuits in a 40nm CMOS technology, providing as much as a 50% reduction in clock distribution and register power consumption.
Andrea Bonetti, Nicholas Preyss, Adam Teman, Andreas Peter Burg
ACM Trans. Design Autom. Electr. Syst.3
2017 Area and Energy-Efficient Complementary Dual-Modular Redundancy Dynamic Memory for Space Applications
abstract
The limited size and power budgets of space-bound systems often contradict the requirements for reliable circuit operation within high-radiation environments. In this paper, we propose the smallest solution for soft-error tolerant embedded memory yet to be presented. The proposed complementary dual-modular redundancy (CDMR) memory is based on a four-transistor dynamic memory core that internally stores complementary data values to provide an inherent per-bit error detection capability. By adding simple, low-overhead parity, an error-correction capability is added to the memory architecture for robust soft-error protection. The proposed memory was implemented in a 65-nm CMOS technology, displaying as much as a 3.5×1 smaller silicon footprint than other radiation-hardened bitcells. In addition, the CDMR memory consumes between 48% and 87% less standby power than other considered solutions across the entire operating region.
Robert Giterman, Lior Atias, Adam Teman
IEEE Trans. Very Large Scale Integr. Syst.3
2017 A 0.65-V, 500-MHz Integrated Dynamic and Static RAM for Error Tolerant Applications
abstract
The diminishing returns provided by voltage scaling have led to a recent paradigm shift toward so-called “approximate computing,” where computation accuracy is traded off for cost in error-tolerant applications. In this paper, a novel approach to achieving the power-performance-area versus data integrity tradeoff is proposed by integrating robust static memory cells and error-prone dynamic cells within a single array. In addition, the resulting integrated dynamic and static random access memory (iD-SRAM) provides the ability to trade off power consumption and accuracy on-the-fly according to the current conditions and operating mode. A 4-kB iD-SRAM array was implemented in a low-power, 65-nm CMOS technology, providing as much as an 80% power reduction and a 20% area reduction as compared with standard approaches, when applied to a video decoder at 500 MHz.
Amit Kazimirsky, Adam Teman, Noa Edri, Alexander Fish
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A low-power correlator for wakeup receivers with algorithm pruning through early termination
abstract
A low-complexity, low-power digital correlator for wakeup receivers is presented. With the proposed algorithm, unnecessary computational cycles are dynamically pruned from the correlation using an early threshold check. For the algorithm, we provide a rigorous mathematical analysis for the associated complexity/performance trade-offs. Furthermore, a low overhead hardware architecture with early-termination capability is developed and implemented in a 0.18μm CMOS technology. The post layout power analysis shows that the presented architecture can reduce power by up to 32% when compared to the conventional architecture with negligible degradation in detection probability and without degradation in false-alarm probability.
Reza Ghanaatian, Paul N. Whatmough, Jeremy Constantin, Adam Teman, Andreas Peter Burg
ISCAS4
2016 A process compensated gain cell embedded-DRAM for ultra-low-power variation-aware design
abstract
Gain cell embedded DRAM (GC-eDRAM) is a high-density alternative to SRAM for ultra-low-power systems. However, due to its dynamic nature, GC-eDRAM requires power-hungry refresh cycles to ensure data retention. Traditional design approaches dictate configuration of the refresh rate according to the worst bitcell, when biased at low-probability, worst-case conditions. However, due to the process variations and local mismatch that can significantly deteriorate the data retention time of a GC-eDRAM bitcell, this design approach often leads to a large power overhead. In this paper, we present a novel GC-eDRAM architecture, incorporating several techniques for variation-aware operation. The primary feature of this architecture is an improved replica scheme for process compensated access tracking that enables calibration for process variations and adaptive refresh according to the array access statistics. The array is shown to ensure data integrity, providing as much as a 7x reduction in retention power over worst-case refresh-rate design for 20% write activity.
Robert Giterman, Adam Teman, Pascal Andreas Meinerzhagen, Alexander Fish, Andreas Peter Burg
ISCAS2
2016 Synthesis of Dual Mode Logic
Lior Moyal, Itamar Levi, Adam Teman, Alexander Fish
Integr.3
2016 Power, Area, and Performance Optimization of Standard Cell Memory Arrays Through Controlled Placement
abstract
Embedded memory remains a major bottleneck in current integrated circuit design in terms of silicon area, power dissipation, and performance; however, static random access memories (SRAMs) are almost exclusively supplied by a small number of vendors through memory generators, targeted at rather generic design specifications. As an alternative, standard cell memories (SCMs) can be defined, synthesized, and placed and routed as an integral part of a given digital system, providing complete design flexibility, good energy efficiency, low-voltage operation, and even area efficiency for small memory blocks. Yet implementing an SCM block with a standard digital flow often fails to exploit the distinct and regular structure of such an array, leaving room for optimization. In this article, we present a design methodology for optimizing the physical implementation of SCM macros as part of the standard design flow. This methodology introduces controlled placement, leading to a structured, noncongested layout with close to 100% placement utilization, resulting in a smaller silicon footprint, reduced wire length, and lower power consumption compared to SCMs without controlled placement. This methodology is demonstrated on SCM macros of various sizes and aspect ratios in a state-of-the-art 28nm fully depleted silicon-on-insulator technology, and compared with equivalent macros designed with the noncontrolled, standard flow, as well as with foundry-supplied SRAM macros. The controlled SCMs provide an average 25% reduction in area as compared to noncontrolled implementations while achieving a smaller size than SRAM macros of up to 1Kbyte. Power and performance comparisons of controlled SCM blocks of a commonly found 256 × 32 (1 Kbyte) memory with foundry-provided SRAMs show greater than 65% and 10% reduction in read and write power, respectively, while providing faster access than their SRAM counterparts, despite being of an aspect ratio that is typically unfavorable for SCMs. In addition, the SCM blocks function correctly with a supply voltage as low as 0.3V, well below the lower limit of even the SRAM macros optimized for low-voltage operation. The controlled placement methodology is applied within a full-chip physical implementation flow of an OpenRISC-based test chip, providing more than 50% power reduction compared to equivalently sized compiled SRAMs under a benchmark application.
Adam Teman, Davide Rossi 0001, Pascal Andreas Meinerzhagen, Luca Benini, Andreas Peter Burg
ACM Trans. Design Autom. Electr. Syst.1
2016 A Low-Voltage Radiation-Hardened 13T SRAM Bitcell for Ultralow Power Space Applications
abstract
Continuous transistor scaling, coupled with the growing demand for low-voltage, low-power applications, increases the susceptibility of VLSI circuits to soft-errors, especially when exposed to extreme environmental conditions, such as those encountered by space applications. The most vulnerable of these circuits are memory arrays that cover large areas of the silicon die and often store critical data. Radiation hardening of embedded memory blocks is commonly achieved by implementing extremely large bitcells or redundant arrays and maintaining a relatively high operating voltage; however, in addition to the resulting area overhead, this often limits the minimum operating voltage of the entire system leading to significant power consumption. In this paper, we propose the first radiation-hardened static random access memory (SRAM) bitcell targeted at low-voltage functionality, while maintaining high soft-error robustness. The proposed 13T employs a novel dual-driven separated-feedback mechanism to tolerate upsets with charge deposits as high as 500 fC at a scaled 500-mV supply voltage. A 32×32 bit memory macro was designed and fabricated in a standard 0.18-μm CMOS process, showing full read and write functionality down to the subthreshold voltage of 300 mV. This is achieved with a cell layout that is only 2× larger than a reference 6T SRAM cell drawn with standard design rules.
Lior Atias, Adam Teman, Robert Giterman, Pascal Andreas Meinerzhagen, Alexander Fish
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Single-Supply 3T Gain-Cell for Low-Voltage Low-Power Applications
abstract
Logic compatible gain cell (GC)-embedded DRAM (eDRAM) arrays are considered an alternative to SRAM due to their small size, nonratioed operation, low static leakage, and two-port functionality. However, traditional GC-eDRAM implementations require boosted control signals in order to write full voltage levels to the cell to reduce the refresh rate and shorten access times. These boosted levels require either an extra power supply or on-chip charge pumps, as well as nontrivial level shifting and toleration of high voltage levels. In this brief, we present a novel, logic compatible, 3T GC-eDRAM bitcell that operates with a single-supply voltage and provides superior write capability to the conventional GC structures. The proposed circuit is demonstrated with a 2-kb memory macro that was designed and fabricated in a mature 0.18-μm CMOS process, targeted at low-power, energy-efficient applications. The test array is powered with a single supply of 900 mV, showing a 0.8-ms worst case retention time, a 1.3-ns write-access time, and a 2.4-pW/bit retention power. The proposed topology provides a bitcell area reduction of 43%, as compared with a redrawn 6-transistor SRAM in the same technology, and an overall macro area reduction of 67% including peripherals.
Robert Giterman, Adam Teman, Pascal Andreas Meinerzhagen, Lior Atias, Andreas Peter Burg, Alexander Fish
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Controlled placement of standard cell memory arrays for high density and low power in 28nm FD-SOI
abstract
Standard cell memories (SCMs) are becoming a popular alternative to SRAM IPs due to their design flexibility, ease of implementation, and robust operation at low supply voltages. Exclusively composed of standard cells, these memory arrays are implemented as part of the standard digital design flow. However, the synthesis and place and route (P&R) algorithms employed by this flow do not exploit the distinct and regular structure of an SCM array, leaving room for optimization. In this paper, we present a controlled placement design methodology for optimizing the physical implementation of SCM macros, leading to a structured, non-congested layout with close to 100% placement utilization and reduced wirelength as compared to unstructured layouts. Three sample SCM macro sizes were implemented according to the proposed methodology in a state-of-the-art 28nm FD-SOI technology, and compared with equivalent macros designed with the non-controlled, standard flow, achieving as much as a 22% reduction in area, a 57% reduction in switching power, and a 42% reduction in leakage power. In addition, these macros provide as much as an 88% reduction in switching power, as compared to equivalently sized, foundry provided SRAM IPs, while enabling robust functionality well below the minimum operating voltage of these IPs.
Adam Teman, Davide Rossi 0001, Pascal Andreas Meinerzhagen, Luca Benini, Andreas Peter Burg
ASP-DAC1
2015 Mitigating the impact of faults in unreliable memories for error-resilient applications
abstract
Inherently error-resilient applications in areas such as signal processing, machine learning and data analytics provide opportunities for relaxing reliability requirements, and thereby reducing the overhead incurred by conventional error correction schemes. In this paper, we exploit the tolerable imprecision of such applications by designing an energy-efficient fault-mitigation scheme for unreliable data memories to meet target yield. The proposed approach uses a bit-shuffling mechanism to isolate faults into bit locations with lower significance. This skews the bit-error distribution towards the low order bits, substantially limiting the output error magnitude. By controlling the granularity of the shuffling, the proposed technique enables trading-off quality for power, area, and timing overhead. Compared to error-correction codes, this can reduce the overhead by as much as 83% in read power, 77% in read access time, and 89% in area, when applied to various data mining applications in 28nm process technology.
Shrikanth Ganapathy, Georgios Karakonstantis, Adam Teman, Andreas Peter Burg
DAC3
2015 Energy versus data integrity trade-offs in embedded high-density logic compatible dynamic memories
Adam Teman, Georgios Karakonstantis, Robert Giterman, Pascal Andreas Meinerzhagen, Andreas Peter Burg
DATE1
2015 An overlap-contention free true-single-phase clock dual-edge-triggered flip-flop
abstract
Dual-edge-triggered (DET) synchronous operation is a very attractive option for low-power, high-performance designs. Compared to conventional single-edge synchronous systems, DET operation is capable of providing the same throughput at half the clock frequency. This can lead to significant power savings on the clock network that is often one of the major contributors to total system power. However, in order to implement DET operation, special registers need to be introduced that sample data on both clock-edges. These registers are more complex than their single-edge counterparts, and often suffer from a certain amount of clock-overlap between the main clock and the internally generated inverted clock. This overlap can cause contention inside the cell and lead to logic failures, especially when operating at scaled power supplies and under process variations that characterize nanometer technologies. This paper presents a novel, static DET flip-flop (DET-FF) with a true-single-phase clock that completely avoids clock overlap hazards by eliminating the need for an inverted clock edge for functionality. The proposed DET FF was implemented in a standard 40nm CMOS technology, showing full functionality at low-voltage operating points, where conventional DET-FFs fail. Under a near-threshold, 500mV supply voltage, the proposed cell also provides a 35% lower CK-to-Q delay and the lowest power-delay-product compared to all considered DET-FF implementations.
Andrea Bonetti, Adam Teman, Andreas Peter Burg
ISCAS2
2015 A Fast Modular Method for True Variation-Aware Separatrix Tracing in Nanoscaled SRAMs
abstract
As memory density continues to grow in modern systems, accurate analysis of static RAM (SRAM) stability is increasingly important to ensure high yields. Traditional static noise margin metrics fail to capture the dynamic characteristics of SRAM behavior, leading to expensive over design and disastrous under design. One of the central components of more accurate dynamic stability analysis is the separatrix; however, its straightforward extraction is extremely time-consuming, and efficient methods are either nonaccurate or extremely difficult to implement. In this paper, we propose a novel algorithm for fast separatrix tracing of any given SRAM topology, designed with industry standard transistor models in nanoscaled technologies. The proposed algorithm is applied to both standard 6T SRAM bitcells, as well as previously proposed alternative subthreshold bitcells, providing up to three orders-of-magnitude speedup, as compared with brute force methods. In addition, for the first time, statistical Monte Carlo separatrix distributions are plotted.
Adam Teman, Roman Visotsky
IEEE Trans. Very Large Scale Integr. Syst.1
2014 4T Gain-Cell with internal-feedback for ultra-low retention power at scaled CMOS nodes
abstract
Gain-Cell embedded DRAM (GC-eDRAM) has recently been recognized as a possible alternative to traditional SRAM. While GC-eDRAM inherently provides high-density, low-leakage, low-voltage, and 2-ported operation, its limited retention time requires periodic, power-hungry refresh cycles. This drawback is further enhanced at scaled technologies, where increased subthreshold leakage currents and decreased in-cell storage capacitances result in faster data deterioration. In this paper, we present a novel 4T GC-eDRAM bitcell that utilizes an internal feedback mechanism to significantly increase the data retention time in scaled CMOS technologies. A 2 kb memory macro was implemented in a low-power 65nm CMOS technology, displaying an over 3× improvement in retention time over the best previous publication at this node. The resulting array displays a nearly 5× reduction in retention power (despite the refresh power component) with a 40% reduction in bitcell area, as compared to a standard 6T SRAM.
Robert Giterman, Adam Teman, Pascal Andreas Meinerzhagen, Andreas Peter Burg, Alexander Fish
ISCAS2
2012 A GIDL free tunneling gate driver for a low power non-volatile memory array
abstract
A recently presented single-poly non-volatile C-Flash memory bitcell provides an ultra-low power low cost option for embedded RFID design. This cell requires the application of a 10V potential difference between the cell's control lines for program and erase operations. Providing the required voltages includes several challenges in the design of the voltage driver, such as the elimination of Gate Induced Drain Leakage (GIDL) currents. In this paper, we present a voltage driver architecture that utilizes novel techniques to overcome the power consumption problems during high voltage propagation. This driver was implemented in the TowerJazz 0.18μm CMOS technology, providing the required functionality with a low static-power figure of 34.6pW.
Hadar Dagan, Adam Teman, Alexander Fish, Evgeny Pikhay, Vladislav Dayan, Yakov Roizin
ISCAS2
2012 A low-cost low-power non-volatile memory for RFID applications
abstract
One of the main obstacles delaying a more widespread use of radio frequency identification (RFID) tags is cost. A critical element of any RFID system is a low power embedded non-volatile memory (NVM) that can be fabricated without additional masks to the core CMOS process. In this paper, we present a 256-bit re-writeable NVM array, implemented in the TowerJazz 0.18µm CMOS process using only standard logic process steps and masks. Based on the single-poly C-Flash bitcell, this array achieves an extremely low static power figure of 3.8µW during operation cycles.
Hadar Dagan, Adam Teman, Alexander Fish, Evgeny Pikhay, Vladislav Dayan, Yakov Roizin
ISCAS2
2012 State space modeling for sub-threshold SRAM stability analysis
abstract
Continuous technology scaling has made traditional Static Noise Margin metrics for stability analysis of SRAM bitcells insufficient. Today, Dynamic Noise Margin analyses and metrics are necessary for state-of-the-art bitcell design, especially under problematic low-voltage operation. In this paper, we overview the concept of state-space modeling for dynamic stability analysis, and then develop an analytical method for evaluating SRAM bitcell operation in the sub-threshold regime. An algorithm for state-space and phase-portrait plotting is proposed and shown to correctly predict subthreshold hold and write behavior of standard bitcells in a 40nm CMOS technology. Implementation of the presented technique in mathematical CAD tools provides orders of magnitude faster evaluation than using traditional brute force approaches.
Janna Mezhibovsky, Adam Teman, Alexander Fish
ISCAS2
2009 Ultra-low Power Subthreshold Flip-flop Design
abstract
In recent years, low power design has become one of the main focuses of digital VLSI circuits. As technology scales, leakage currents in contemporary CMOS logic have become one of the main power consumers. Contrary to conventional methods for power reduction, where efforts are taken to reduce subthreshold leakage, operation of digital circuits in the subthreshold region, utilizes this current, minimizing power consumption in low-frequency systems. This paper proposes two architectures for implementing flip-flop cells, designed to operate in the subthreshold region. Both cells integrate a gate-diffusion input (GDI) multiplexer in their designs to minimize area and capacitance. Timing parameters of the flip-flops are calculated and techniques for improving the timing characteristics are proposed. The proposed designs are simulated in a standard 90 nm process achieving a power dissipation of 8.4 nW in a typical corner at VDD = 300 mV with a delay of 51.7 nsec.
Sagi Fisher, Adam Teman, Dmitry Vaysman, Alexander Gertsman, Orly Yadid-Pecht, Alexander Fish
ISCAS2
2008 Autonomous CMOS image sensor for real time target detection and tracking
abstract
An autonomous image sensor for real time target detection and tracking is presented. The sensor is based on a CMOS APS array, equipped with in-pixel functionality and integrates analog and digital components to achieve autonomous operation with minimal power dissipation. The system employs a two-phased operation flow; during the initial acquisition stage, the digital controller detects and acquires the brightest targets in the field of view within a single frame and defines windows of interest (WOI) around the center of mass coordinates of each object. Subsequently, the system moves into the analog tracking mode during which all areas outside of the WOI are entirely shut down, thus saving power to a number of orders of magnitude. In addition to its low power dissipation, the sensor features real-time operation, low fixed pattern noise, linearity and the ability to track a predefined number of targets throughout the entire field of view. A 64x64 pixel sensor array has been designed in 0.18μm CMOS technology and is operated via a 1.8V supply. The imager architecture is discussed, the circuits’ descriptions are shown and simulation results are presented.
Adam Teman, Sagi Fisher, Liby Sudakov, Alexander Fish, Orly Yadid-Pecht
ISCAS1