EDBT 2026 Demo / reviewers in the wild / expert
Peter A. Beerel
dblp:29/6330 · also Peter Anthony Beerel
· DBLP profile ↗
106ranked-venue papers
11as first author
48since 2021 · last 2025
0000-0002-8283-0168ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 78 · 9 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 16 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Software engineering, systems software and programming languages · 12 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating Hallucinations in Vision-Language Models through Image-Guided Head SuppressionabstractDespite their remarkable progress in multimodal understanding tasks, large vision language models (LVLMs) often suffer from "hallucination", generating texts misaligned with the visual context.Existing methods aimed at reducing hallucinations through inference time intervention incur a significant increase in latency.To mitigate this, we present SPIN, a task-agnostic attention-guided head suppression strategy that can be seamlessly integrated during inference without incurring any significant compute or latency overhead.We investigate whether hallucination in LVLMs can be linked to specific model components.Our analysis suggests that hallucinations can be attributed to a dynamic subset of attention heads in each layer.Leveraging this insight, for each text query token, we selectively suppress attention heads that exhibit low attention to image tokens, keeping the top-k attention heads intact.Extensive evaluations on visual question answering and image description tasks demonstrate the efficacy of SPIN in reducing hallucination scores up to 2.7× while maintaining F1, and improving throughput by 1.8× compared to existing alternatives.Code is available here. Sreetama Sarkar, Yue Che, Alex Gavin, Peter A. Beerel, Souvik Kundu 0002 |
EMNLP | 4 |
| 2025 | An Efficient Distributed Machine Learning Inference Framework with Byzantine Fault Detection
Utkarsh Mohan, Peter A. Beerel |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Dynamic SpikFormer: Low-Latency & Energy-Efficient Spiking Neural Networks with Dynamic Time Steps for Vision TransformersabstractSpiking Neural Networks (SNNs) have emerged as a popular spatio-temporal computing paradigm for complex vision tasks. Recently proposed SNN training algorithms have significantly reduced the number of time steps (down to 1) for improved latency and energy efficiency, however, they target only convolutional neural networks (CNN). These algorithms, when applied to the recently spotlighted vision transformers (ViT), either require a large number of time steps or fail to converge. Based on the analysis of the histograms of the ANN and SNN activation maps, we hypothesize that each ViT block has a different sensitivity to the number of time steps. We propose a novel training framework that dynamically allocates the number of time steps to each ViT module depending on a trainable score assigned to each timestep. In particular, we generate a scalar binary time step mask that filters spikes emitted by each neuron in a leaky-integrate-and-fire (LIF) layer. The resulting SNNs have high activation sparsity and require only accumulate operations (AC), except for the input embedding layer, in contrast to expensive multiply-and-accumulates (MAC) needed in traditional ViTs. This yields significant improvements in energy efficiency. We evaluate our training framework and resulting SNNs on image recognition tasks including CIFAR10, CIFAR100, and ImageNet with different ViT architectures. We obtain a test accuracy of 95.97% with 4.97 time steps with direct encoding on CIFAR10. Gourav Datta, Zeyu Liu 0003, Anni Li, Peter A. Beerel |
ICASSP | 4 |
| 2025 | Optimizing SFQ Circuit Design: A Timing-Driven Framework for Performance-Constrained Area MinimizationabstractSingle Flux Quantum (SFQ) digital logic offers a promising path to energy-efficient, high-performance computing, but faces significant scalability challenges–particularly due to the area overhead associated with gate-level path balancing and the pipelining of splitter trees. Prior efforts to reduce this overhead often unnecessarily compromise throughput, diminishing SFQ’s key performance advantage. To mitigate this, we propose a timing-aware optimization framework that unifies and enhances both traditional full path balancing (FPB) and multi-phase clocking approaches. Inspired by timing-driven EDA strategies in CMOS, our approach explicitly models gate delays and jointly optimizes clock phase assignments, pipelining, and splitter tree construction, minimizing area and latency overhead while achieving a given performance target. On benchmark circuits, our timing-aware 2-phase clocking achieves up to 30% area and 29% latency reductions over FPB at a 27ps clock period. Compared to prior multi-phase methods operating at 42ps, we achieve 22% lower area and 23% lower latency. At a tighter 21ps performance target, our timing-aware single-phase clocking improves area and latency by 28% and 7%, respectively. All optimizations are implemented using scalable, polynomial-time algorithms. Robert Aviles, Rassul Bairamkulov, Peter A. Beerel |
ICCAD | 4 |
| 2025 | Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon PhotonicsabstractVision Transformers (ViTs) have emerged as a powerful architecture for computer vision tasks due to their ability to model long-range dependencies and global contextual relationships. However, their substantial compute and memory demands hinder efficient deployment in scenarios with strict energy and bandwidth limitations. In this work, we propose Opto-ViT, the first near-sensor, region-aware ViT accelerator leveraging silicon photonics (SiPh) for real-time and energy-efficient vision processing. Opto-ViT features a hybrid electronic-photonic architecture, where the optical core handles compute-intensive matrix multiplications using Vertical-Cavity Surface-Emitting Lasers (VCSELs) and Microring Resonators (MRs), while nonlinear functions and normalization are executed electronically. To reduce redundant computation and patch processing, we introduce a lightweight Mask Generation Network (MGNet) that identifies regions of interest in the current frame and prunes irrelevant patches before ViT encoding. We further co-optimize the ViT backbone using quantization-aware training and matrix decomposition tailored for photonic constraints. Experiments across device fabrication, circuit and architecture co-design, to classification, detection, and video tasks demonstrate that Opto-ViT achieves 100.4 KFPS/W with up to 84% energy savings with less than 1.6% accuracy loss, while enabling scalable and efficient ViT deployment at the edge. Mehrdad Morsali, Chengwei Zhou, Deniz Najafi, Sreetama Sarkar, Pietro Mercati, Navid Khoshavi, Peter A. Beerel, Mahdi Nikdast, Gourav Datta, Shaahin Angizi |
ICCAD | 7 |
| 2025 | MaskVD: Region Masking for Efficient Video Object DetectionabstractVideo tasks are compute-heavy and thus pose a challenge when deploying in real-time applications, particularly for tasks that require state-of-the-art Vision Trans-formers (ViTs). Several research efforts have tried to address this challenge by leveraging the fact that large portions of the video undergo very little change across frames leading to redundant computations in frame-based video processing. In particular, some works leverage pixel or semantic differences across frames, however, this yields limited latency benefits with significantly increased memory overhead. This paper, in contrast, presents a strategy for masking regions in video frames that leverages the semantic information in images and the temporal correlation between frames to significantly reduce FLOPs and latency with little to no penalty in performance over base-line models. In particular, we demonstrate that by lever-aging extracted features from previous frames, ViT back-bones directly benefit from region masking, skipping up to 80% of input regions, improving FLOPs and latency by 3.14x and 1.5x. We improve memory and latency over the state-of-the-art (SOTA) by 2.3x and 1.14x, while maintaining similar detection performance. Additionally, our approach demonstrates promising results on convolutional neural networks (CNNs) and provides latency improvements over the SOTA up to 1.3x using specialized computational kernels. Code is available at: https://github.com/sreetamasarkar/MaskVD Sreetama Sarkar, Gourav Datta, Souvik Kundu 0002, Chirayata Bhattacharyya, Peter A. Beerel |
WACV | 6 |
| 2024 | Toward High-Accuracy, Programmable Extreme-Edge Intelligence for Neuromorphic Vision Sensors utilizing Magnetic Domain Wall Motion-based MTJabstractThe desire to empower resource-limited edge devices with computer vision (CV) must overcome the high energy consumption of collecting and processing vast sensory data. To address the challenge, this work proposes an energy-efficient non-von-Neumann in-pixel processing solution for neuromorphic vision sensors employing emerging (X) magnetic domain wall magnetic tunnel junction (MDWMTJ) for the first time, in conjunction with CMOS-based neuromorphic pixels. Our hybrid CMOS+X approach performs in-situ massively parallel asynchronous analog convolution, exhibiting low power consumption and high accuracy across various CV applications by leveraging the non-volatility and programmability of the MDWMTJ. Moreover, our developed device-circuit-algorithm co-design framework captures device constraints (low tunnel-magnetoresistance, low dynamic range) and circuit constraints (non-linearity, process variation, area consideration) based on monte-carlo simulations and device parameters utilizing GF22nm FD-SOI technology. Our experimental results suggest we can achieve an average of 45.3% reduction in backend-processor energy, maintaining similar front-end energy compared to the state-of-the-art and high accuracy of 79.17% and 95.99% on the DVS-CIFAR10 and IBM DVS128-Gesture datasets, respectively. Md. Abdullah-Al Kaiser, Gourav Datta, Peter A. Beerel, Akhilesh Jaiswal 0001 |
DAC | 3 |
| 2024 | Challenges and Unexplored Frontiers in Electronic Design Automation for Superconducting Digital LogicabstractPositioned as a highly promising post-CMOS computing technology, superconductor electronics (SCE) offer the potential for unparalleled performance and energy efficiency gains compared to end-of-roadmap CMOS circuits. However, achieving very large-scale integration poses numerous challenges. These challenges span from the modeling and analysis of superconducting devices and logic gates to the intricate design of complex SCE circuits and systems. Addressing power and clock distribution issues, minimizing adverse effects of flux trappings, and mitigating stray electromagnetic fields in sensitive SCE circuitry are key challenges that need attention. Verification and testing of SCE circuits also remain open problems. Moreover, scaling the minimum feature sizes of SCE circuits, currently set at 150nm, presents critical scaling and physical design challenges that must be overcome. This review aims to delve into these issues, providing detailed insights while exploring existing or potential solutions to overcome them. Sasan Razmkhah, Robert Aviles, Mingye Li, Sandeep Gupta 0001, Peter A. Beerel, Massoud Pedram |
DATE | 5 |
| 2024 | Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural Networks: from Algorithms to TechnologyabstractNeuromorphic computing and, in particular, spiking neural networks (SNNs) have become an attractive alternative to deep neural networks for a broad range of signal processing applications, processing static and/or temporal inputs from different sensory modalities, including audio and vision sensors. In this paper, we start with a description of recent advances in algorithmic and optimization innovations to efficiently train and scale low-latency, and energy-efficient spiking neural networks (SNNs) for complex machine learning applications. We then discuss the recent efforts in algorithm-architecture co-design that explores the inherent trade-offs between achieving high energy-efficiency and low latency while still providing high accuracy and trustworthiness. We then describe the underlying hardware that has been developed to leverage such algorithmic innovations in an efficient way. In particular, we describe a hybrid method to integrate significant portions of the model’s computation within both memory components as well as the sensor itself. Finally, we discuss the potential path forward for research in building deployable SNN systems identifying key challenges in the algorithm-hardware-application co-design space with an emphasis on trustworthiness. Souvik Kundu 0002, Rui-Jie Zhu 0003, Akhilesh Jaiswal 0001, Peter A. Beerel |
ICASSP | 4 |
| 2024 | Training Ultra-Low-Latency Spiking Neural Networks from ScratchabstractSpiking Neural networks (SNN) have emerged as an attractive spatio-temporal computing paradigm for a wide range of low-power vision tasks. However, state-of-the-art (SOTA) SNN models either incur multiple time steps which hinder their deployment in real-time use cases or increase the training complexity significantly. To mitigate this concern, we present a training framework (from scratch) for SNNs with ultra-low (down to 1) time steps that leverages the Hoyer regularizer. We calculate the threshold for each BANN layer as the Hoyer extremum of a clipped version of its activation map. The clipping value is determined through training using gradient descent with our Hoyer regularizer. We evaluate the efficacy of our training framework on large-scale vision tasks, including traditional and event-based image recognition and object detection. Our experiments demonstrate up to 34× increase in compute efficiency with a marginal accuracy/mAP drop compared to non-spiking networks. Finally, we implement our framework in the Lava-DL library, thereby enabling the deployment of our SNN models in the Loihi neuromorphic chip. Gourav Datta, Zeyu Liu 0003, Peter A. Beerel |
ICASSP | 3 |
| 2024 | Mitigate Replication and Copying in Diffusion Models with Generalized Caption and Dual Fusion EnhancementabstractWhile diffusion models demonstrate a remarkable capability for generating high-quality images, their tendency to ‘replicate’ training data raises privacy concerns. Although recent research suggests that this replication may stem from the insufficient generalization of training data captions and duplication of training images, effective mitigation strategies remain elusive. To address this gap, our paper first introduces a generality score that measures the caption generality and employ large language model (LLM) to generalize training captions. Subsequently, we leverage generalized captions and propose a novel dual fusion enhancement approach to mitigate the replication of diffusion models. Our empirical results demonstrate that our proposed methods can significantly reduce replication by 43.5% compared to the original diffusion model while maintaining the diversity and quality of generations. Code is available at https://github.com/HowardLi0816/dual-fusion-diffusion. Dake Chen, Peter A. Beerel |
ICASSP | 4 |
| 2024 | A Joint Optimization of Buffer and Splitter Insertion for Phase-Skipping Adiabatic Quantum - Flux - Parametron Circuits
Robert Aviles, Peter A. Beerel |
ICCD | 2 |
| 2024 | Can we get the best of both Binary Neural Networks and Spiking Neural Networks for Efficient Computer Vision?abstractBinary Neural networks (BNN) have emerged as an attractive computing paradigm for a wide range of low-power vision tasks. However, state-of-the-art (SOTA) BNNs do not yield any sparsity, and induce a significant number of non-binary operations. On the other hand, activation sparsity can be provided by spiking neural networks (SNN), that too have gained significant traction in recent times. Thanks to this sparsity, SNNs when implemented on neuromorphic hardware, have the potential to be significantly more power-efficient compared to traditional artifical neural networks (ANN). However, SNNs incur multiple time steps to achieve close to SOTA accuracy. Ironically, this increases latency and energy---costs that SNNs were proposed to reduce---and presents itself as a major hurdle in realizing SNNs’ theoretical gains in practice. This raises an intriguing question: *Can we obtain SNN-like sparsity and BNN-like accuracy and enjoy the energy-efficiency benefits of both?* To answer this question, in this paper, we present a training framework for sparse binary activation neural networks (BANN) using a novel variant of the Hoyer regularizer. We estimate the threshold of each BANN layer as the Hoyer extremum of a clipped version of its activation map, where the clipping value is trained using gradient descent with our Hoyer regularizer.
This approach shifts the activation values away from the threshold, thereby mitigating the effect of noise that can otherwise degrade the BANN accuracy. Our approach outperforms existing BNNs, SNNs, and adder neural networks (that also avoid energy-expensive multiplication operations similar to BNNs and SNNs) in terms of the accuracy-FLOPs trade-off for complex image recognition tasks. Downstream experiments on object detection further demonstrate the efficacy of our approach. Lastly, we demonstrate the portability of our approach to SNNs with multiple time steps. Codes are publicly available [here](https://github.com/godatta/Ultra-Low-Latency-SNN). Gourav Datta, Zeyu Liu 0003, Peter A. Beerel |
ICLR | 3 |
| 2024 | LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory UnitsabstractTransformer models have demonstrated high accuracy in numerous applications but have high complexity and lack sequential processing capability making them ill-suited for many streaming applications at the edge where devices are heavily resource-constrained. Thus motivated, many researchers have proposed reformulating the transformer models as RNN modules which modify the self-attention computation with explicit states. However, these approaches often incur significant performance degradation.
The ultimate goal is to develop a model that has the following properties: parallel training, streaming and low-cost inference, and state-of-the-art (SOTA) performance. In this paper, we propose a new direction to achieve this goal. We show how architectural modifications to a fully-sequential recurrent model can help push its performance toward Transformer models while retaining its sequential processing capability. Specifically, inspired by the recent success of Legendre Memory Units (LMU) in sequence learning tasks, we propose LMUFormer, which augments the LMU with convolutional patch embedding and convolutional channel mixer.
Moreover, we present a spiking version of this architecture, which introduces the benefit of states within the patch embedding and channel mixer modules while simultaneously reducing the computing complexity.
We evaluated our architectures on multiple sequence datasets. Of particular note is our performance on the Speech Commands V2 dataset (35 classes). In comparison to SOTA transformer-based models within the ANN domain, our LMUFormer demonstrates comparable performance while necessitating a remarkable $70\times$ reduction in parameters and a substantial $140\times$ decrement in FLOPs. Furthermore, when benchmarked against extant low-complexity SNN variants, our model establishes a new SOTA with an accuracy of 96.12\%.
Additionally, owing to our model's proficiency in real-time data processing, we are able to achieve a 32.03\% reduction in sequence length, all while incurring an inconsequential decline in performance. Zeyu Liu 0003, Gourav Datta, Anni Li, Peter A. Beerel |
ICLR | 4 |
| 2024 | FixPix: Fixing Bad Pixels using Deep Learning
Sreetama Sarkar, Xinan Ye, Gourav Datta, Peter A. Beerel |
ICPR (3) | 4 |
| 2024 | What Makes Vision Transformers Robust Towards Bit-Flip Attack?
Souvik Kundu 0002, Dake Chen, Peter A. Beerel |
ICPR (8) | 5 |
| 2024 | FireLoc: Low-latency Multi-modal Wildfire GeolocationabstractFirefighters still rely on coarse remote sensing and inaccurate eyewitness reports to localize spreading wildfires. Despite advances in sensing, UAVs, and computer vision, the community has yet to combine the right modalities to achieve effective wildfire geolocalization and spotting. We present FireLoc, a fast and accurate wildfire crowdsensing system that localizes and maps wildfires combining ground cameras and landscape data. Yue Hu 0019, Prashanth Sutrave, Peter A. Beerel, Barath Raghavan |
SenSys | 4 |
| 2024 | Guest Editorial: Special Issue on Learning, Optimization, and Implementation for Circuits and Systems Driven by Artificial IntelligenceabstractCircuits and systems, such as multidimensional and nonlinear ones, large-scale integration circuits, and power networks, play a significant role in the whole spectrum of science and technology, from basic scientific theories to various real-world applications. With the increasing demand from applications, it is vital to develop circuits and systems with high accuracy, stability, flexibility, and security through efficient learning, design optimization, and integrated implementation. The rapid advancement of artificial intelligence (AI) has fostered a symbiotic relationship between circuits and systems and AI in both theory and applications. On the one hand, research in circuits and systems on efficient learning, design optimization, and integrated implementation aided by AI has recently gained a promising development, where energy-efficient circuits and systems have a very broad range of applications. On the other hand, the utilization of AI in real-world applications has become indispensable for the optimization and implementation of circuits and systems with high efficiency and low-power computation. Overall, through advanced learning, optimization, and implementation driven by AI, efficient circuits and systems running in real-time with low power can be realized for wider applications. Yang Tang 0001, Peter A. Beerel, Jürgen Kurths, Guanrong Chen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | C2PI: An Efficient Crypto-Clear Two-Party Neural Network Private InferenceabstractRecently, private inference (PI) has addressed the rising concern over data and model privacy in machine learning inference as a service. However, existing PI frameworks suffer from high computational and communication costs due to the expensive multi-party computation (MPC) protocols. Existing literature has developed lighter MPC protocols to yield more efficient PI schemes. We, in contrast, propose to lighten them by introducing an empirically-defined privacy evaluation. To that end, we reformulate the threat model of PI and use inference data privacy attacks (IDPAs) to evaluate data privacy. We then present an enhanced IDPA, named distillation-based inverse-network attack (DINA), for improved privacy evaluation. Finally, we leverage the findings from DINA and propose C2PI, a two-party PI framework presenting an efficient partitioning of the neural network model and requiring only the initial few layers to be performed with MPC protocols. Based on our experimental evaluations, relaxing the formal data privacy guarantees C2PI can speed up existing PI frameworks, including Delphi [1] and Cheetah [2], up to 2.89× and 3.88× under LAN and WAN settings, respectively, and save up to 2.75× communication costs. Dake Chen, Souvik Kundu 0002, Haomei Liu, Ruiheng Peng, Peter A. Beerel |
DAC | 6 |
| 2023 | Island-based Random Dynamic Voltage Scaling vs ML-Enhanced Power Side-Channel AttacksabstractIn this paper, we describe and analyze an island-based random dynamic voltage scaling (iRDVS) approach to thwart power side-channel attacks. We first analyze the impact of the number of independent voltage islands on the resulting signal-to-noise ratio and trace misalignment. As part of our analysis of misalignment, we propose a novel unsupervised machine learning (ML) based attack that is effective on systems with three or fewer independent voltages. Our results show that iRDVS with four voltage islands, however, cannot be broken with 200k encryption traces, suggesting that iRDVS can be effective. We finish the talk by describing an iRDVS test chip in a 12nm FinFet process that incorporates three variants of an AES-256 accelerator, all originating from the same RTL. This included a synchronous core, an asynchronous core with no protection, and a core employing the iRDVS technique using asynchronous logic. Lab measurements from the chips indicated that both unprotected variants failed the test vector leakage assessment (TVLA) security metric test, while the iRDVS was proven secure in a variety of configurations. Dake Chen, Christine Goins, Maxwell Waugaman, Georgios D. Dimou, Peter A. Beerel |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | Technology-Circuit-Algorithm Tri-Design for Processing-in-Pixel-in-Memory (P2M)abstractThe massive amounts of data generated by camera sensors motivate data processing inside pixel arrays, i.e., at the extreme-edge. Several critical developments have fueled recent interest in the processing-in-pixel-in-memory paradigm for a wide range of visual machine intelligence tasks, including (1) advances in 3D integration technology to enable complex processing inside each pixel in a 3D integrated manner while maintaining pixel density, (2) analog processing circuit techniques for massively parallel low-energy in-pixel computations, and (3) algorithmic techniques to mitigate non-idealities associated with analog processing through hardware-aware training schemes. This article presents a comprehensive technology-circuit-algorithm landscape that connects technology capabilities, circuit design strategies, and algorithmic optimizations to power, performance, area, bandwidth reduction, and application-level accuracy metrics. We present our results using a comprehensive co-design framework incorporating hardware and algorithmic optimizations for various complex real-life visual intelligence tasks mapped onto our P2M paradigm. Md. Abdullah-Al Kaiser, Gourav Datta, Sreetama Sarkar, Souvik Kundu 0002, Zihan Yin, Manas Garg, Ajey P. Jacob, Peter A. Beerel, Akhilesh Jaiswal 0001 |
ACM Great Lakes Symposium on VLSI | 8 |
| 2023 | In-Sensor & Neuromorphic Computing Are all You Need for Energy Efficient Computer VisionabstractDue to the high activation sparsity and use of accumulates (AC) instead of expensive multiply-and-accumulates (MAC), neuromorphic spiking neural networks (SNNs) have emerged as a promising low-power alternative to traditional DNNs for several computer vision (CV) applications. However, most existing SNNs require multiple time steps for acceptable inference accuracy, hindering real-time deployment and increasing spiking activity and, consequently, energy consumption. Recent works proposed direct encoding that directly feeds the analog pixel values in the first layer of the SNN in order to significantly reduce the number of time steps. Although the overhead for the first layer MACs with direct encoding is negligible for deep SNNs and the CV processing is efficient using SNNs, the data transfer between the image sensors and the downstream processing costs significant bandwidth and may dominate the total energy. To mitigate this concern, we propose an in-sensor computing hardware-software co-design framework for SNNs targeting image recognition tasks. Our approach reduces the bandwidth between sensing and processing by 12−96× and the resulting total energy by 2.32× compared to traditional CV processing, with a 3.8% reduction in accuracy on ImageNet. Gourav Datta, Zeyu Liu 0003, Md. Abdullah-Al Kaiser, Souvik Kundu 0002, Joe Mathai, Zihan Yin, Ajey P. Jacob, Akhilesh Jaiswal 0001, Peter A. Beerel |
ICASSP | 9 |
| 2023 | Sparse Mixture Once-for-all Adversarial Training for Efficient in-situ Trade-off between Accuracy and Robustness of DNNsabstractExisting deep neural networks (DNNs) that achieve state-of-the-art (SOTA) performance on both clean and adversarially-perturbed images rely on either activation or weight conditioned convolution operations. However, such conditional learning costs additional multiply-accumulate (MAC) or addition operations, increasing inference memory and compute costs. To that end, we present a sparse mixture once for all adversarial training (SMART), that allows a model to train once and then in-situ trade-off between accuracy and robustness, that too at a reduced compute and parameter overhead. In particular, SMART develops two expert paths, for clean and adversarial images, respectively, that are then conditionally trained via respective dedicated sets of binary sparsity masks. Extensive evaluations on multiple image classification datasets across different models show SMART to have up to 2.72× fewer non-zero parameters costing proportional reduction in compute overhead, while yielding SOTA accuracy-robustness trade-off. Additionally, we present insightful observations in designing sparse masks to successfully condition on both clean and perturbed images. Souvik Kundu 0002, Sairam Sundaresan, Sharath Nittur Sridhar, Shunlin Lu, Peter A. Beerel |
ICASSP | 6 |
| 2023 | Quantpipe: Applying Adaptive Post-Training Quantization For Distributed Transformer Pipelines In Dynamic Edge EnvironmentsabstractPipeline parallelism has achieved great success in deploying large-scale transformer models in cloud environments, but has received less attention in edge environments. Unlike in cloud scenarios with high-speed and stable network inter-connects, dynamic bandwidth in edge systems can degrade distributed pipeline performance. We address this issue with QuantPipe, a communication-efficient distributed edge system that introduces post-training quantization (PTQ) to compress the communicated tensors. QuantPipe uses adaptive PTQ to change bitwidths in response to bandwidth dynamics, maintaining transformer pipeline performance while incurring limited inference accuracy loss. We further improve the accuracy with a directed-search analytical clipping for integer quantization method (DS-ACIQ), which bridges the gap between estimated and real data distributions. Experimental results show that QuantPipe adapts to dynamic bandwidth to maintain pipeline performance while achieving a practical model accuracy using a wide range of quantization bitwidths, e.g., improving accuracy under 2-bit quantization by 15.85% on ImageNet compared to naive quantization. Connor Imes, Souvik Kundu 0002, Peter A. Beerel, Stephen P. Crago, John Paul Walters |
ICASSP | 4 |
| 2023 | RNA-ViT: Reduced-Dimension Approximate Normalized Attention Vision Transformers for Latency Efficient Private InferenceabstractThe concern over data and model privacy in machine learning inference as a service (MLaaS) has led to the development of private inference (PI) techniques. However, existing PI frameworks, especially those designed for large models such as vision transformers (ViT), suffer from high computational and communication overheads caused by the expensive multi-party computation (MPC) protocols. The encrypted attention module that involves the softmax operation contributes significantly to this overhead. In this work, we present a family of models dubbed RNA-ViT, that leverage a novel attention module called reduced-dimension approximate normalized attention and a latency efficient GeLU-alternative layer. In particular, RNA-ViT uses two novel techniques to improve PI efficiency in ViTs: a reduced-dimension normalized attention (RNA) architecture and a high order polynomial (HOP) softmax approximation for latency efficient normalization. We also propose a novel metric, accuracy-to-latency ratio (A2L), to evaluate modules in terms of their accuracy and PI latency. Based on this metric, we perform an analysis to identify a nonlinearity module with improved PI efficiency. Our extensive experiments show that RNA-ViT can achieve average 3.53×, 3.54×, 1.66× lower PI latency with an average accuracy improvement of 0.93%, 2.04%, and 2.73% compared to the state-of-the-art scheme MPCViT [1], on CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively. Dake Chen, Souvik Kundu 0002, Peter A. Beerel |
ICCAD | 5 |
| 2023 | SAL-ViT: Towards Latency Efficient Private Inference on ViT using Selective Attention Search with a Learnable Softmax ApproximationabstractRecently, private inference (PI) has addressed the rising concern over data and model privacy in machine learning inference as a service. However, existing PI frameworks suffer from high computational and communication overheads due to the expensive multi-party computation (MPC) protocols, particularly for large models such as vision transformers (ViT). The majority of this overhead is due to the encrypted softmax operation in each self-attention layer. In this work, we present SAL-ViT with two novel techniques to boost PI efficiency on ViTs. Our first technique is a learnable PI-efficient approximation to softmax, namely, learnable 2Quad (L2Q), that introduces learnable scaling and shifting parameters to the prior 2Quad softmax approximation, enabling improvement in accuracy. Then, given our observation that external attention (EA) presents lower PI latency than widely-adopted self-attention (SA) at the cost of accuracy, we present a selective attention search (SAS) method to integrate the strength of EA and SA. Specifically, for a given lightweight EA ViT, we leverage a constrained optimization procedure to selectively search and replace EA modules with SA alternatives to maximize the accuracy. Our extensive experiments show that our SAL-ViT can averagely achieve 1.28×, 1.28×, 1.14× lower PI latency with 1.79%, 1.41%, and 2.08% higher accuracy compared to the existing alternatives, on CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively. Dake Chen, Souvik Kundu 0002, Peter A. Beerel |
ICCV | 5 |
| 2023 | Learning to Linearize Deep Neural Networks for Secure and Efficient Private Inference
Souvik Kundu 0002, Shunlin Lu, Jacqueline Tiffany Liu, Peter A. Beerel |
ICLR | 5 |
| 2023 | ViTA: A Vision Transformer Inference Accelerator for Edge ApplicationsabstractVision Transformer models, such as ViT, Swin Transformer, and Transformer-in-Transformer, have recently gained significant traction in computer vision tasks due to their ability to capture the global relation between features which leads to superior performance. However, they are compute-heavy and difficult to deploy in resource-constrained edge devices. Existing hardware accelerators, including those for the closely-related BERT transformer models, do not target highly resource-constrained environments. In this paper, we address this gap and propose ViTA - a configurable hardware accelerator for inference of vision transformer models, targeting resource-constrained edge computing devices and avoiding repeated off-chip memory accesses. We employ a head-level pipeline and inter-layer MLP optimizations, and can support several commonly used vision transformer models with changes solely in our control logic. We achieve nearly 90% hardware utilization efficiency on most vision transformer models, report a power of 0.88W when synthesised with a clock of 150 MHz, and get reasonable frame rates - all of which makes ViTA suitable for edge applications. Shashank Nag, Gourav Datta, Souvik Kundu 0002, Nitin Chandrachoodan, Peter A. Beerel |
ISCAS | 5 |
| 2023 | Bridging the Gap Between Spiking Neural Networks & LSTMs for Latency & Energy EfficiencyabstractSpiking Neural Networks (SNNs) have emerged as an attractive spatio-temporal computing paradigm for complex vision tasks. However, most existing works yield models that require many time steps and do not leverage the inherent temporal dynamics of spiking neural networks, even for sequential tasks. Motivated by this observation, we propose an optimized spiking long short-term memory networks (LSTM) training framework that involves a novel ANN-to-SNN conversion framework, followed by SNN fine-tuning via backpropagation through time (BPTT). In particular, we propose novel activation functions in the source LSTM architecture and convert a judiciously selected subset of them to leaky-integrate-and-fire (LIF) activations with optimal bias shifts. Moreover, we propose a pipelined parallel processing scheme that hides the SNN time steps, significantly improving system latency, especially for long sequences. The resulting SNNs have high activation sparsity and require only accumulate operations (AC), in contrast to expensive multiply-and-accumulates (MAC) needed for ANNs, except for the input layer when using direct encoding, yielding significant improvements in energy efficiency. We evaluate our framework on sequential learning tasks including temporal MNIST, Google Speech Commands (GSC), and UCI Smartphone datasets on different LSTM architectures. We obtain test accuracy of 94.75 % with only 2 time steps on the GSC dataset with$\sim 4.1\times$lower energy than an iso-architecture standard LSTM. Gourav Datta, Haoqin Deng, Robert Aviles, Zeyu Liu 0003, Peter A. Beerel |
ISLPED | 5 |
| 2023 | Self-Attentive Pooling for Efficient Deep LearningabstractEfficient custom pooling techniques that can aggressively trim the dimensions of a feature map for resource-constrained computer vision applications have recently gained significant traction. However, prior pooling works extract only the local context of the activation maps, limiting their effectiveness. In contrast, we propose a novel non-local self-attentive pooling method that can be used as a drop-in replacement to the standard pooling layers, such as max/average pooling or strided convolution. The proposed self-attention module uses patch embedding, multihead self-attention, and spatial-channel restoration, followed by sigmoid activation and exponential soft-max. This self-attention mechanism efficiently aggregates dependencies between non-local activation patches during downsampling. Extensive experiments on standard object classification and detection tasks with various convolutional neural network (CNN) architectures demonstrate the superiority of our proposed mechanism over the state-of-the-art (SOTA) pooling techniques. In particular, we surpass the test accuracy of existing pooling techniques on different variants of MobileNet-V2 on ImageNet by an average of ~1.2%. With the aggressive down-sampling of the activation maps in the initial layers (providing up to 22x reduction in memory consumption), our approach achieves 1.43% higher test accuracy compared to SOTA techniques with iso-memory footprints. This enables the deployment of our models in memory-constrained devices, such as micro-controllers (without losing significant accuracy), because the initial activation maps consume a significant amount of on-chip memory for high-resolution images required for complex vision tasks. Our pooling method also leverages channel pruning to further reduce memory footprints. Codes are available at https://github.com/CFun/Non-Local-Pooling. Gourav Datta, Souvik Kundu 0002, Peter A. Beerel |
WACV | 4 |
| 2023 | Enabling ISPless Low-Power Computer VisionabstractCurrent computer vision (CV) systems use an image signal processing (ISP) unit to convert the high resolution raw images captured by image sensors to visually pleasing RGB images. Typically, CV models are trained on these RGB images and have yielded state-of-the-art (SOTA) performance on a wide range of complex vision tasks, such as object detection. In addition, in order to deploy these models on resource-constrained low-power devices, recent works have proposed in-sensor and in-pixel computing approaches that try to partly/fully bypass the ISP and yield significant bandwidth reduction between the image sensor and the CV processing unit by downsampling the activation maps in the initial convolutional neural network (CNN) layers. However, direct inference on the raw images degrades the test accuracy due to the difference in covariance of the raw images captured by the image sensors compared to the ISP-processed images used for training. Moreover, it is difficult to train deep CV models on raw images, because most (if not all) large-scale open-source datasets consist of RGB images. To mitigate this concern, we propose to invert the ISP pipeline, which can convert the RGB images of any dataset to its raw counterparts, and enable model training on raw images. We release the raw version of the COCO dataset, a large-scale benchmark for generic high-level vision tasks. For ISP-less CV systems, training on these raw images result in a ∼7.1% increase in test accuracy on the visual wake works (VWW) dataset compared to relying on training with traditional ISP-processed RGB datasets. To further improve the accuracy of ISP-less CV models and to increase the energy and bandwidth benefits obtained by in-sensor/in-pixel computing, we propose an energy-efficient form of analog in-pixel demosaicing that may be coupled with in-pixel CNN computations. When evaluated on raw images captured by real sensors from the PASCALRAW dataset, our approach results in a 8.1% increase in mAP. Lastly, we demonstrate a further 20.5% increase in mAP by using a novel application of few-shot learning with thirty shots each for the novel PASCALRAW dataset, constituting 3 classes. Codes are available at https://github.com/godatta/ISP-less-CV. Gourav Datta, Zeyu Liu 0003, Zihan Yin, Linyu Sun, Akhilesh Jaiswal 0001, Peter A. Beerel |
WACV | 6 |
| 2023 | FLOAT: Fast Learnable Once-for-All Adversarial Training for Tunable Trade-off between Accuracy and RobustnessabstractExisting models that achieve state-of-the-art (SOTA) performance on both clean and adversarially-perturbed images rely on convolution operations conditioned with feature-wise linear modulation (FiLM) layers. These layers require additional parameters and are hyperparameter sensitive. They significantly increase training time, memory cost, and potential latency which can be costly for resource-limited or real-time applications. In this paper, we present a fast learnable once-for-all adversarial training (FLOAT) algorithm, which instead of the existing FiLM-based conditioning, presents a unique weight conditioned learning that requires no additional layer, thereby incurring no significant increase in parameter count, training time, or network latency compared to standard adversarial training. In particular, we add configurable scaled noise to the weight tensors that enables a trade-off between clean and adversarial performance. Extensive experiments show that FLOAT can yield SOTA performance improving both clean and perturbed image classification by up to ~6% and ~10%, respectively. Moreover, real hardware measurement shows that FLOAT can reduce the training time by up to 1.43× with fewer model parameters of up to 1.47× on iso-hyperparameter settings compared to the FiLM-based alternatives. Additionally, to further improve memory effi ciency we introduce FLOAT sparse (FLOATS), a form of non-iterative model pruning and provide detailed empirical analysis in yielding a three-way accuracy-robustness-complexity trade-off for these new class of pruned conditionally trained models. Souvik Kundu 0002, Sairam Sundaresan, Massoud Pedram, Peter A. Beerel |
WACV | 4 |
| 2023 | On the Security of Sequential Logic Locking Against Oracle-Guided AttacksabstractThe Boolean satisfiability (SAT) attack is an oracle-guided attack that can break most combinational logic locking schemes by efficiently pruning out all the wrong keys from the search space. Extending such an attack to sequential logic locking requires multiple time-consuming rounds of SAT solving, performed using an “unrolled” version of the sequential circuit, and model checking, used to determine the successful termination of the attack. This article addresses these challenges by formally characterizing the relation between the minimum unrolling depth required to prune out the wrong keys of an SAT-based attack and a notion of functional corruptibility (FC) for sequential circuits, which can be efficiently estimated from a locked circuit to indicate the progress of an SAT-based attack. Based on this analysis, we present an FC-guided SAT-based attack that can significantly reduce unnecessary SAT and model-checking tasks. We present two versions of the attack, namely,Fun-SATandFun-SAT+, based on whether the attacker has a priori knowledge of the key length.Fun-SATaims to find the correct key sequence, whileFun-SAT+aims to retrieve the correct initial state of the circuit. The numerical evaluation shows thatFun-SATcan be, on average,$90\boldsymbol {\times }$faster than previous attacks against state-of-the-art locking methods. On the other hand, when using an approximate termination condition,Fun-SAT+can find an initial state that leads to at most 0.1% FC in 76.9% instances that would otherwise time out after one day. Yinghua Hu, Kaixin Yang, Dake Chen, Peter A. Beerel, Pierluigi Nuzzo 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Can Deep Neural Networks be Converted to Ultra Low-Latency Spiking Neural Networks?abstractSpiking neural networks (SNNs), that operate via binary spikes distributed over time, have emerged as a promising energy efficient ML paradigm for resource-constrained devices. However, the current state-of-the-art (SOTA) SNNs require multiple time steps for acceptable inference accuracy, increasing spiking activity and, consequently, energy consumption. SOTA training strategies for SNNs involve conversion from a non-spiking deep neural network (DNN). In this paper, we determine that SOTA conversion strategies cannot yield ultra low latency because they incorrectly assume that the DNN and SNN pre-activation values are uniformly distributed. We propose a new training algorithm that accurately captures these distributions, minimizing the error between the DNN and converted SNN. The resulting SNNs have ultra low latency and high activation sparsity, yielding significant improvements in compute efficiency. In particular, we evaluate our framework on image recognition tasks from CIFAR-10 and CIFAR-100 datasets on several VGG and ResNet architectures. We obtain top-1 accuracy of 64.19% with only 2 time steps on the CIFAR-100 dataset with ∼159.2× lower compute energy compared to an iso-architecture standard DNN. Compared to other SOTA SNN models, our models perform inference 2.5-8× faster (i.e., with fewer time steps). Gourav Datta, Peter A. Beerel |
DATE | 2 |
| 2022 | BMPQ: Bit-Gradient Sensitivity-Driven Mixed-Precision Quantization of DNNs from ScratchabstractLarge DNNs with mixed-precision quantization can achieve ultra-high compression while retaining high classification performance. However, because of the challenges in finding an accurate metric that can guide the optimization process, these methods either sacrifice significant performance compared to the 32-bit floating-point (FP-32) baseline or rely on a compute-expensive, iterative training policy that requires the availability of a pre-trained baseline. To address this issue, this paper presents BMPQ, a training method that uses bit gradients to analyze layer sensitivities and yield mixed-precision quantized models. BMPQ requires a single training iteration but does not need a pre-trained baseline. It uses an integer linear program (ILP) to dynamically adjust the precision of layers during training, subject to a fixed hardware budget. To evaluate the efficacy of BMPQ, we conduct extensive experiments with VGG16 and ResNet18 on CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets. Compared to the baseline FP-32 models, BMPQ can yield models that have 15.4x fewer parameter bits with negligible drop in accuracy. Compared to the SOTA “during training”, mixed-precision training scheme, our models are 2.1 x, 2.2x, and 2.9x smaller, on CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively, with an improved accuracy of up to 14.54%. Souvik Kundu 0002, Shikai Wang, Qirui Sun, Peter A. Beerel, Massoud Pedram |
DATE | 4 |
| 2022 | TriLock: IC Protection with Tunable Corruptibility and Resilience to SAT and Removal AttacksabstractSequential logic locking has been studied over the last decade as a method to protect sequential circuits from reverse engineering. However, most of the existing sequential logic locking techniques are threatened by increasingly more sophisticated SAT-based attacks, efficiently using input queries to a SAT solver to rule out incorrect keys, as well as removal attacks based on structural analysis. In this paper, we propose TriLock, a sequential logic locking method that simultaneously addresses these vulnerabilities. TriLock can achieve high, tunable functional corruptibility while still guaranteeing exponential queries to the SAT solver in a SAT-based attack. Further, it adopts a state re-encoding method to obscure the boundary between the original state registers and those inserted by the locking method, thus making it more difficult to detect and remove the locking-related components. Yinghua Hu, Pierluigi Nuzzo 0002, Peter A. Beerel |
DATE | 4 |
| 2022 | PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge DevicesabstractDeep neural networks with large model sizes achieve state-of-the-art results for tasks in computer vision and natural language processing. However, such models are too compute- or memory-intensive for resource-constrained edge devices. Prior works on parallel and distributed execution primarily focus on training-rather than inference-using homogeneous accelerators in data centers. We propose PipeEdge, a distributed framework for edge systems that uses pipeline parallelism to both speed up inference and enable running larger, more accurate models that otherwise cannot fit on single edge devices. PipeEdge uses an optimal partition strategy that considers heterogeneity in compute, memory, and network bandwidth. Our empirical evaluation demonstrates that PipeEdge achieves 11.88× and 12.78× speedup using 16 edge devices for the ViT-Huge and BERT-Large models, respectively, with no accuracy loss. Similarly, PipeEdge improves throughput for ViT-Huge (which cannot fit in a single device) by 3.93× over a 4-device baseline using 16 edge devices. Finally, we show up to 4.16× throughput improvement over the state-of-the-art PipeDream when using a heterogeneous set of devices. Connor Imes, Xuanang Zhao, Souvik Kundu 0002, Peter A. Beerel, Stephen P. Crago, John Paul Walters |
DSD | 5 |
| 2022 | Radiation Hardening by Design Techniques for the Mutual Exclusion ElementabstractCircuits in advanced CMOS technology are increasingly more sensitive to transient pulses caused by radiation particles that strike vulnerable circuit components, specially turned off transistors, often generating multiple voltage upsets. Towards mitigating these issues, this paper presents a novel Radiation Hardened by Design (RHBD) mutual exclusion element (mutex) that incorporates multiple RHBD techniques with reduced area overhead. Moisés Herrera, Peter A. Beerel |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | P2M-DeTrack: Processing-in-Pixel-in-Memory for Energy-efficient and Real-Time Multi-Object Detection and TrackingabstractToday’s high resolution, high frame rate cameras in autonomous vehicles generate a large volume of data that needs to be transferred and processed by a downstream processor or machine learning (ML) accelerator to enable intelligent computing tasks, such as multi-object detection and tracking. The massive amount of data transfer incurs significant energy, latency, and bandwidth bottlenecks, which hinders real-time processing. To mitigate this problem, we propose an algorithm-hardware co-design framework called Processing-in-Pixel-in-Memory-based object Detection and Tracking (P2M-DeTrack). P2M-DeTrack is based on a custom faster R-CNN-based model that is distributed partly inside the pixel array (front-end) and partly in a separate FPGA/ASIC (back-end). The proposed front-end in-pixel processing down-samples the input feature maps significantly with judiciously optimized strided convolution and pooling. Compared to a conventional baseline design that transfers frames of RGB pixels to the back-end, the resulting P2M-DeTrack designs reduce the data bandwidth between sensor and back-end by up to 24×. The designs also reduce the sensor and total energy (obtained from in-house circuit simulations at Globalfoundries 22nm technology node) per frame by 5.7× and 1.14×, respectively. Lastly, they reduce the sensing and total frame latency by an estimated 1.7× and 3×, respectively. We evaluate our approach on the multi-object object detection (tracking) task of the large-scale BDD100K dataset and observe only a 0.5% reduction in the mean average precision (0.8% reduction in the identification F1 score) compared to the state-of-the-art. Gourav Datta, Souvik Kundu 0002, Zihan Yin, Joe Mathai, Zeyu Liu 0003, Mulin Tian, Shunlin Lu, Ravi Teja Lakkireddy, Andrew G. Schmidt, Wael Abd-Almageed, Ajey P. Jacob, Akhilesh Jaiswal 0001, Peter A. Beerel |
VLSI-SoC | 14 |
| 2022 | Converting Flip-Flop to Clock-Gated 3-Phase Latch-Based Designs Using Graph-Based RetimingabstractLatches have the advantages of timing-borrowing, smaller cell area, lower input capacitance, and lower power compared to flip-flops (FFs). This article presents a CAD flow that converts any arbitrarily complex single-clock-domain FF-based RTL design into an efficient 3-phase latch-based design. The flow includes a novel 3-phase aware retiming algorithm for power and area optimization. Post place-and-route results demonstrate that our new 3-phase designs achieve 23.5% and 23.9% average power reductions compared to more traditional FF and master–slave-based alternatives across a board range of benchmarks with no degradation in performance and on average less area. Huimei Cheng, Yichen Gu, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Toward Adversary-aware Non-iterative Model Pruning through Dynamic Network Rewiring of DNNsabstractWe present a dynamic network rewiring (DNR) method to generate pruned deep neural network (DNN) models that both are robust against adversarially generated images and maintain high accuracy on clean images. In particular, the disclosed DNR training method is based on a unified constrained optimization formulation using a novel hybrid loss function that merges sparse learning with robust adversarial training. This training strategy dynamically adjusts inter-layer connectivity based on per-layer normalized momentum computed from the hybrid loss function. To further improve the robustness of the pruned models, we propose DNR++, an extension of the DNR method where we introduce the idea of sparse parametric Gaussian noise tensor that is added to the weight tensors to yield robust regularization. In contrast to existing robust pruning frameworks that require multiple training iterations, the proposed DNR and DNR++ achieve an overall target pruning ratio with only a single training iteration and can be tuned to support both irregular and structured channel pruning. To demonstrate the efficacy of the proposed method under the no-increased-training-time “free” adversarial training scenario, we finally present FDNR++, a simple yet effective training modification that can yield robust yet compressed models requiring training time comparable to that of an unpruned non-adversarial training. To evaluate the merits of our disclosed training methods, experiments were performed with two widely accepted models, namely VGG16 and ResNet18, on CIFAR-10 and CIFAR-100 as well as with VGG16 on Tiny-ImageNet. Compared to the baseline uncompressed models, our methods provide over 20× compression on all the datasets without any significant drop of either clean or adversarial classification performance. Moreover, extensive experiments show that our methods consistently find compressed models with better clean and adversarial image classification performance than what is achievable through state-of-the-art alternatives. We provide insightful observations to help make various model, parameter density, and prune-type selection choices and have open-sourced our saved models and test codes to ensure reproducibility of our results. Souvik Kundu 0002, Bill Ye, Peter A. Beerel, Massoud Pedram |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | DNR: A Tunable Robust Pruning Framework Through Dynamic Network Rewiring of DNNsabstractThis paper presents a dynamic network rewiring (DNR) method to generate pruned deep neural network (DNN) models that are robust against adversarial attacks yet maintain high accuracy on clean images. In particular, the disclosed DNR method is based on a unified constrained optimization formulation using a hybrid loss function that merges ultra-high model compression with robust adversarial training. This training strategy dynamically adjusts inter-layer connectivity based on per-layer normalized momentum computed from the hybrid loss function. In contrast to existing robust pruning frameworks that require multiple training iterations, the proposed learning strategy achieves an overall target pruning ratio with only a single training iteration and can be tuned to support both irregular and structured channel pruning. To evaluate the merits of DNR, experiments were performed with two widely accepted models, namely VGG16 and ResNet-18, on CIFAR-10, CIFAR-100 as well as with VGG16 on Tiny-ImageNet. Compared to the baseline uncompressed models, DNR provides over 20x compression on all the datasets with no significant drop in either clean or adversarial classification accuracy. Moreover, our experiments show that DNR consistently finds compressed models with better clean and adversarial image classification performance than what is achievable through state-of-the-art alternatives. Our models and test codes are available at https://github.com/ksouvik52/DNR_ASP_DAC2021. Souvik Kundu 0002, Mahdi Nazemi, Peter A. Beerel, Massoud Pedram |
ASP-DAC | 3 |
| 2021 | HIRE-SNN: Harnessing the Inherent Robustness of Energy-Efficient Deep Spiking Neural Networks by Training with Crafted Input NoiseabstractLow-latency deep spiking neural networks (SNNs) have become a promising alternative to conventional artificial neural networks (ANNs) because of their potential for increased energy efficiency on event-driven neuromorphic hardware. Neural networks, including SNNs, however, are subject to various adversarial attacks and must be trained to remain resilient against such attacks for many applications. Nevertheless, due to prohibitively high training costs associated with SNNs, an analysis and optimization of deep SNNs under various adversarial attacks have been largely overlooked. In this paper, we first present a detailed analysis of the inherent robustness of low-latency SNNs against popular gradient-based attacks, namely fast gradient sign method (FGSM) and projected gradient descent (PGD). Motivated by this analysis, to harness the model’s robustness against these attacks we present an SNN training algorithm that uses crafted input noise and incurs no additional training time. To evaluate the merits of our algorithm, we conducted extensive experiments with variants of VGG and ResNet on both CIFAR-10 and CIFAR-100 dataset. Compared to standard trained direct-input SNNs, our trained models yield improved classification accuracy of up to 13.7% and 10.1% on FGSM and PGD attack generated images, respectively, with negligible loss in clean image accuracy. Our models also outperform inherently-robust SNNs trained on rate-coded inputs with improved or similar classification performance on attack-generated images while having up to 25× and ∼4.6× lower latency and computation energy, respectively. For reproducibility, we have open-sourced the code at github.com/ksouvik52/hiresnn2021. Souvik Kundu 0002, Massoud Pedram, Peter A. Beerel |
ICCV | 3 |
| 2021 | Training Energy-Efficient Deep Spiking Neural Networks with Single-Spike Hybrid Input EncodingabstractSpiking Neural Networks (SNNs) have emerged as an attractive alternative to traditional deep learning frameworks, since they provide higher computational efficiency in event driven neuromorphic hardware. However, the state-of-the-art (SOTA) SNNs suffer from high inference latency, resulting from inefficient input encoding and training techniques. The most widely used input coding schemes, such as Poisson based rate-coding, do not leverage the temporal learning capabilities of SNNs. This paper presents a training framework for low-latency energy-efficient SNNs that uses a hybrid encoding scheme at the input layer in which the analog pixel values of an image are directly applied during the first timestep and a novel variant of spike temporal coding is used during subsequent timesteps. In particular, neurons in every hidden layer are restricted to fire at most once per image which increases activation sparsity. To train these hybrid-encoded SNNs, we propose a variant of the gradient descent based spike timing dependent backpropagation (STDB) mechanism using a novel cross entropy loss function based on both the output neurons' spike time and membrane potential. The resulting SNNs have reduced latency and high activation sparsity, yielding significant improvements in computational efficiency. In particular, we evaluate our proposed training scheme on image classification tasks from CIFAR-10 and CIFAR-100 datasets on several VGG architectures. We achieve top-l accuracy of 66.46% with 5 timesteps on the CIFAR-100 dataset with ~125x less compute energy than an equivalent standard ANN. Additionally, our proposed SNN performs 5–300 x faster inference compared to other state-of-the-art rate or temporally coded SNN models. Gourav Datta, Souvik Kundu 0002, Peter A. Beerel |
IJCNN | 3 |
| 2021 | Analyzing the Confidentiality of Undistillable Teachers in Knowledge DistillationabstractKnowledge distillation (KD) has recently been identified as a method that can unintentionally leak private information regarding the details of a teacher model to an unauthorized student. Recent research in developing undistillable nasty teachers that can protect model confidentiality has gained significant attention. However, the level of protection these nasty models offer has been largely untested. In this paper, we show that transferring knowledge to a shallow sub-section of a student can largely reduce a teacher’s influence. By exploring the depth of the shallow subsection, we then present a distillation technique that enables a skeptical student model to learn even from a nasty teacher. To evaluate the efficacy of our skeptical students, we conducted experiments with several models with KD on both training data-available and data-free scenarios for various datasets. While distilling from nasty teachers, compared to the normal student models, skeptical students consistently provide superior classification performance of up to ∼59.5%. Moreover, similar to normal students, skeptical students maintain high classification accuracy when distilled from a normal teacher, showing their efficacy irrespective of the teacher being nasty or not. We believe the ability of skeptical students to largely diminish the KD-immunity of potentially nasty teachers will motivate the research community to create more robust mechanisms for model confidentiality. We have open-sourced the code at https://github.com/ksouvik52/Skeptical2021 Souvik Kundu 0002, Qirui Sun, Massoud Pedram, Peter A. Beerel |
NeurIPS | 5 |
| 2021 | Spike-Thrift: Towards Energy-Efficient Deep Spiking Neural Networks by Limiting Spiking Activity via Attention-Guided CompressionabstractThe increasing demand for on-chip edge intelligence has motivated the exploration of algorithmic techniques and specialized hardware to reduce the computation energy of current machine learning models. In particular, deep spiking neural networks (SNNs) have gained interest because their event-driven hardware implementations can consume very low energy. However, minimizing average spiking activity and thus energy consumption while preserving accuracy in deep SNNs remains a significant challenge and opportunity. This paper proposes a novel two-step SNN compression technique to reduce their spiking activity while maintaining accuracy that involves compressing specifically-designed artificial neural networks (ANNs) that are then converted into the target SNNs. Our approach uses an ultra-high ANN compression technique that is guided by the attention-maps of an uncompressed meta-model. We then evaluate the firing threshold of each ANN layer and start with the trained ANN weights to perform a sparse-learning-based supervised SNN training to minimize the number of time steps required while retaining compression. To evaluate the merits of the proposed approach, we performed experiments with variants of VGG and ResNet, on both CIFAR-10 and CIFAR-100, and VGG16 on Tiny-ImageNet. SNN models generated through the proposed technique yield state-of-the-art compression ratios of up to 33.4× with no significant drop in accuracy compared to baseline unpruned counterparts. As opposed to the existing SNN pruning methods we achieve up to 8.3× better compression with no drop in accuracy. Moreover, compressed SNN models generated by our methods can have up to 12.2× better compute energy-efficiency compared to ANNs that have a similar number of parameters. Souvik Kundu 0002, Gourav Datta, Massoud Pedram, Peter A. Beerel |
WACV | 4 |
| 2021 | Metastability in Superconducting Single Flux Quantum (SFQ) LogicabstractSuperconducting digital electronics, especially Single Flux Quantum (SFQ), has emerged as a promising beyond-CMOS technology with Josephson junctions (JJ) as the active device. It has the potential to meet the booming demands of lower power consumption and higher operation speeds in the electronics industry and future exascale supercomputing systems. Despite these promises, scaling SFQ circuits remains a serious challenge that motivates the support of multiple SFQ clock domains. Towards this end, this paper analyzes the impact of setup time violations and metastability in SFQ circuits comparing the derived analytical models to their CMOS counterparts. It also proposes new techniques to reduce the average latency in metastability-tolerant SFQ synchronizers, and evaluates their effects on the layout and critical margin of the design. It further extends the proposed model to estimate the Mean Time Between Failure (MTBF) of flip-flop-based synchronizers and shows that their MTBF with the current feature sizes is unaffected by noise, similar to CMOS. Finally, it curve fits this model to simulations using the state-of-the-art SFQ5ee process and shows that a two-flop SFQ synchronizer with a clock frequency of 25 GHz has an estimated MTBF of ~106years. Gourav Datta, Yunkun Lin, Bo Zhang 0098, Peter A. Beerel |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | A Variation-aware Hold Time Fixing Methodology for Single Flux Quantum Logic CircuitsabstractSingle flux quantum (SFQ) logic is a promising technology to replace complementary metal-oxide-semiconductor logic for future exa-scale supercomputing but requires the development of reliable EDA tools that are tailored to the unique characteristics of SFQ circuits, including the need for active splitters to support fanout and clocked logic gates. This article is the first work to present a physical design methodology for inserting hold buffers in SFQ circuits. Our approach is variation-aware, uses common path pessimism removal and incremental placement to minimize the overhead of timing fixes, and can trade off layout area and timing yield. Compared to a previously proposed approach using fixed hold time margins, Monte Carlo simulations show that, averaging across 10 ISCAS’85 benchmark circuits, our proposed method can reduce the number of inserted hold buffers by 8.4% with a 6.2% improvement in timing yield and by 21.9% with a 1.7% improvement in timing yield. Soheil Nazar Shahsavani, Massoud Pedram, Peter A. Beerel |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2020 | Deep-n-Cheap: An Automated Search Framework for Low Complexity Deep LearningabstractWe present Deep-n-Cheap – an open-source AutoML framework to search for deep learning models. This search includes both architecture and training hyperparameters, and supports convolutional neural networks and multilayer perceptrons. Our framework is targeted for deployment on both benchmark and custom datasets, and as a result, offers a greater degree of search space customizability as compared to a more limited search over only pre-existing models from literature. We also introduce the technique of ’search transfer’, which demonstrates the generalization capabilities of our models to multiple datasets. Deep-n-Cheap includes a user-customizable complexity penalty which trades off performance with training time or number of parameters. Specifically, our framework results in models offering performance comparable to state-of-the-art while taking 1-2 orders of magnitude less time to train than models from other AutoML and model search frameworks. Additionally, this work investigates and develops various insights regarding the search process. In particular, we show the superiority of a greedy strategy and justify our choice of Bayesian optimization as the primary search methodology over random / grid search. Sourya Dey, Saikrishna C. Kanala, Keith M. Chugg, Peter A. Beerel |
ACML | 4 |
| 2020 | Saving Power by Converting Flip-Flop to 3-Phase Latch-Based DesignsabstractLatches are smaller and lower power than flip-flops (FFs) and are typically used in a time-borrowing master-slave configuration. This paper presents an automatic flow for converting arbitrarily-complex single-clock-domain FF-based RTL designs to efficient 3-phase latch-based designs with reduced number of required latches, saving both register and clock-tree power. Post place-and-route results demonstrate that our 3-phase latch-based designs save an average of 15.5% and 18.5% power on a variety of ISCAS, CEP, and CPU benchmark circuits, compared to their more traditional FF and master-slave based alternatives. Huimei Cheng, Yichen Gu, Peter A. Beerel |
DATE | 4 |
| 2020 | Neural Network Training with Approximate Logarithmic ComputationsabstractThe high computational complexity associated with training deep neural networks limits online and real-time training on edge devices. This paper proposed an end-to-end training and inference scheme that eliminates multiplications by approximate operations in the log-domain which has the potential to significantly reduce implementation complexity. We implement the entire training procedure in the log-domain, with fixed-point data representations. This training procedure is inspired by hardware-friendly approximations of log-domain addition which are based on look-up tables and bit-shifts. We show that our 16-bit log-based training can achieve classification accuracy within approximately 1% of the equivalent floating-point baselines for a number of commonly used datasets. Arnab Sanyal, Peter A. Beerel, Keith M. Chugg |
ICASSP | 2 |
| 2020 | Modeling and Characterization of Metastability in Single Flux Quantum (SFQ) SynchronizersabstractDespite the promises of low-power and high-frequency of single-flux quantum (SFQ) technology, scaling these circuits remains a serious challenge that motivates the support of multiple SFQ clock domains. Towards this end, this paper analyzes the impact of setup time violations and metastability in SFQ circuits comparing the derived analytical models to their CMOS counterparts. It then extends this model to estimate the Mean Time Between Failure (MTBF) of flip-flop-based synchronizers and curve fits this model to simulations in the state-of-the-art SFQ5ee process. Interestingly, we find a two-flop SFQ synchronizer has an estimated MTBF of ~106years. Gourav Datta, Peter A. Beerel |
ISCAS | 2 |
| 2020 | Pre-Defined Sparsity for Low-Complexity Convolutional Neural NetworksabstractThe high energy cost of processing deep convolutional neural networks impedes their ubiquitous deployment in energy-constrained platforms such as embedded systems and IoT devices. This article introduces convolutional layers with pre-defined sparse 2D kernels that have support sets that repeat periodically within and across filters. Due to the efficient storage of our periodic sparse kernels, the parameter savings can translate into considerable improvements in energy efficiency due to reduced DRAM accesses, thus promising significant improvements in the trade-off between energy consumption and accuracy for both training and inference. To evaluate this approach, we performed experiments with two widely accepted datasets, CIFAR-10 and Tiny ImageNet in sparse variants of the ResNet18 and VGG16 architectures. Compared to baseline models, our proposed sparse variants require up to ~82% fewer model parameters with 5.6× fewer FLOPs with negligible loss in accuracy for ResNet18 on CIFAR-10. For VGG16 trained on Tiny ImageNet, our approach requires 5.8× fewer FLOPs and up to ~83.3% fewer model parameters with a drop in top-5 (top-1) accuracy of only 1.2% (~2.1%). We also compared the performance of our proposed architectures with that of ShuffleNet and MobileNetV2. Using similar hyperparameters and FLOPs, our ResNet18 variants yield an average accuracy improvement of ~2.8%. Souvik Kundu 0002, Mahdi Nazemi, Massoud Pedram, Keith M. Chugg, Peter A. Beerel |
IEEE Trans. Computers | 5 |
| 2020 | A Theoretical Foundation for Timing Synchronous Systems Using Asynchronous StructuresabstractTiming of synchronous systems is an everlasting stumbling block to the booming demands for lower power consumption and higher operation speeds in the electronics industry. This hardship is aggravated by the growing levels of variability in state-of-the-art silicon dimensions and in other beyond-CMOS technologies. Although some designers continue to strongly believe in the performance advantages of being fully synchronous, others have radically shifted toward extremely robust delay-insensitive domains. Targeting a different compromise of both performance and robustness, this article provides sufficient conditions for an asynchronous system to be able to generate the periodic signals necessary for the timing of a fully synchronous system and highlights a specific hierarchical clocking structure that with a single tunable delay satisfies these conditions. Using an asynchronous clock distribution network benefits from both the natural robustness of asynchronous structures and the advantageous performance of synchronous clocking. Ramy N. Tadros, Peter A. Beerel |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2019 | Automatic Retiming of Two-Phase Latch-Based Resilient CircuitsabstractTiming resilient design has shown significant promise in mitigating the excess margins associated with rare worst-case data and increased process, voltage, and temperature variations. However, resilient circuits need error detecting sequential logic (EDL) to detect timing errors which incur area and power overhead. This paper proposes two alternatives to reduce the overhead in two-phase latch-based resilient circuits. The first is a new resiliency-aware graph-based approach to solve the retiming problem. The second uses a virtual resynthesis library to enable commercial synthesis tools to recognize the EDL overhead and optimize total area during retiming. We compare both approaches to a commercially standard retiming approach, which ignores the resiliency overheads, on a wide variety of benchmarks. Our experimental results show that our methods are computationally efficient and reduce the total circuit area by an average of up to 10%-15% when compared to traditional retiming. Huimei Cheng, Hsiao-Lun Wang, Minghe Zhang, Dylan Hand, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | Opportunities for Machine Learning in Electronic Design AutomationabstractThe rise of machine learning (ML) has introduced many opportunities for computer-aided-design, VLSI design, and their intersection. Related to computer-aided design, we review several classical CAD algorithms which can benefit from ML, outline the key challenges, and discuss promising approaches. In particular, because some of the existing ML accelerators have used asynchronous design, we review the state-of-the-art in asynchronous CAD support, and identify opportunities for ML within these flows. Peter A. Beerel, Massoud Pedram |
ISCAS | 1 |
| 2018 | A Robust and Self-Adaptive Clocking Technique for RSFQ Circuits - The ArchitectureabstractIn the beyond-CMOS era, many researchers are investigating technologies that have the potential of complementing, if not overtaking, silicon-based electronics. One of which is rapid single flux quantum (RSFQ) superconductive technology. Due to timing and uncertainties challenges, the promised benefits of three orders of magnitude lower power at an order of magnitude higher performance does not look attainable. In this paper, we propose an innovative self-adaptive clocking technique which is designed to be robust against an unprecedented variability. Whereas the traditional zero-skew clocking is not reliable for large scale RSFQ designs, our proposed hierarchical chains of homogeneous clover-leaves clocking inherits its robustness from spatially correlated cell delays and from the timing robustness of the RSFQ traditional counter-flow clocking. Our simulations on ISCAS'85 benchmark circuits show our robust clocking costs can be on average within 34% as fast as ideal zero-skew jitter-free clock trees with an average area overhead of 58%. Moreover, lower area overhead is also possible at the cost of performance, demonstrating the possible trade-off between performance and area. We assert that these overheads are acceptable given the benefits of much higher functionality and feasibility, as quantified by yield, as well as improved scalability. In particular, previous efforts [1] reported that a similar, but more limited, clocking structure leads to up to 93% higher yield than zero-skew trees. Ramy N. Tadros, Peter A. Beerel |
ISCAS | 2 |
| 2018 | Area Optimization of Timing Resilient Designs Using ResynthesisabstractTiming resilient designs can remove variation margins by adding error detecting logic (EDL) that detects timing errors when execution completes within a resiliency window. Speeding up near-critical-paths during logic synthesis can reduce the amount of EDL needed but at the cost of increasing logic area. This creates a logic optimization strategy called resynthesis. This paper proposes four alternatives to optimize resilient designs through resynthesis. The first is a brute force approach that explores speeding up all combinations of near-critical paths and produces good results but is computationally impractical for complex circuits. The second is a naive brute-force (BF) approach in which near-critical paths are sped up one end-point at a time. It is much faster than the BF approach because it does not explore the benefits of speeding up multiple end-points simultaneously and thus provides a quick-and-dirty lower bound for the benefits of resynthesis. The third is a geometric program-based iterative algorithm (GPIA) that achieves area reductions that compare favorably across all four approaches. The GPIA algorithm completes within 24 h for all examples and the average area reduction is up to 16%. Because the run-time required to solve this mathematical model can still be long; however, we propose a fourth approach that involves creating a virtual resynthesis cell library that tries to trick the synthesis tool to understand EDL overhead and optimize total area quickly and automatically. This approach obtains an average of approximately two-third of the area reductions of the GPIA approach with fast run times associated with only a single synthesis run. Hsin-Ho Huang, Huimei Cheng, Chris C. N. Chu, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Retiming of Two-Phase Latch-Based Resilient CircuitsabstractTiming resilient design has shown significant promise in mitigating the excess margins associated with rare worst-case data and increased process, voltage, and temperature (PVT) variations. However, resilient circuits need error detecting sequential logic (EDL) to detect timing errors which represents area and power overhead. This article proposes a new network-simplex-based retiming method for two-phase latch-based resilient circuits to reduce the overhead of the combination of normal and error detecting latches. Our experimental results show that our method is computationally efficient and reduces the cost of the sequential elements by an average of up to 20% when compared to traditional methods. Hsiao-Lun Wang, Minghe Zhang, Peter A. Beerel |
DAC | 3 |
| 2017 | Accelerating Training of Deep Neural Networks via Sparse Edge Processing
Sourya Dey, Yinan Shao, Keith M. Chugg, Peter A. Beerel |
ICANN (1) | 4 |
| 2017 | Reconditioning: A Framework for Automatic Power Optimization of QDI CircuitsabstractThis paper introduces reconditioning: a novel systematic technique for reducing unnecessary power consumption of asynchronous gate-level netlists, which involves the optimal reordering of conditional communication and logic primitives. Our technique is applicable to asynchronous circuits with handshaking protocols that encode data and control together, in particular, quasi delay insensitive and 1-of-N handshaking circuits. Both an optimal integer linear program (ILP) and a fast heuristic algorithm are presented. We show that our ILP is feasible for moderate size circuits and our heuristic algorithm scales to much larger circuits, completing in seconds on circuits with tens of thousands of gates. Our experimental results show power improvement highly depends on the structure of the circuit but can often be above 26% with typically less than 5% area overhead. Arash Saifhashemi, Hsin-Ho Huang, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Area optimization of resilient designs guided by a mixed integer geometric programabstractTiming resilient designs can remove variation margins by adding error detecting logic (EDL) that detects timing errors when execution completes within a resiliency window. Speeding up near-critical-paths during logic synthesis can reduce the amount of EDL needed but at the cost of increasing logic area. This creates a logic optimization strategy called resynthesis. This paper proposes a gate-sizing based mixed integer geometric programming framework to analytically model and optimize paths during resynthesis. We evaluate our approach on a set of ISCAS89 benchmarks and compare the overall area improvement after resynthesis guided by our mathematical model versus a previously published naive brute-force approach. Our experimental results demonstrate that our approach achieves up to 11% larger average area improvement. Hsin-Ho Huang, Huimei Cheng, Chris C. N. Chu, Peter A. Beerel |
DAC | 4 |
| 2016 | Low Area, Low Power, Robust, Highly Sensitive Error Detecting Latch for Resilient ArchitecturesabstractOperating at lower supply voltages to meet ever-increasing demands for power-efficiency unfortunately aggravates process, voltage, and temperature (PVT) variability. Resilient architectures have emerged as a promising way to mitigate widening worst-case margins at these voltages. In particular, timing resilient architectures use extra circuitry to detect timing violations and recover to its normal operation. The error detecting latch (EDL) is an efficient circuit that helps perform this task. This paper proposes two EDL architectures that achieve as much as 11.2% less power consumption, 20.8% less leakage, 7.8% smaller area, and 18.2% better sensitivity to glitches compared to state-of-the-art EDLs. The paper offers two different flavors trading off robustness for lower power and vice versa. The paper also proposes a comprehensive power metric encapsulating many of the various energy aspects discussed in the literature. Weizhe Hua, Ramy N. Tadros, Peter A. Beerel |
ISLPED | 3 |
| 2016 | A Fine-Grain, Uniform, Energy-Efficient Delay Element for 2-Phase Bundled-Data CircuitsabstractContemporary digitally controlled delay elements (DEs) trade off power overheads and delay quantization error (DQE). This article proposes a new programmable DE that provides a balanced design that yields low power with moderate DQE even under process, voltage, and temperature variations. The element employs and leverages the advantages offered by a 28nm fully depleted silicon on insulator technology, using back body biasing to add an extra dimension to its programmability. To do so, a novel generic delay shift block is proposed, which enables incorporating both fine and coarse delays in a single DE that can be easily integrated into digital systems, which is an advantage over hybrid DEs that rely on analog design. Ajay Singhvi, Matheus T. Moreira, Ramy N. Tadros, Ney Laert Vilar Calazans, Peter A. Beerel |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2015 | Logical equivalence checking of asynchronous circuits using commercial tools
Arash Saifhashemi, Hsin-Ho Huang, Priyanka Bhalerao, Peter A. Beerel |
DATE | 4 |
| 2014 | Stochastic analysis of Bubble RazorabstractBubble Razor has been proposed to eliminate required timing margins in synchronous design caused by increasing delay variation due to process variation and aging. However, the theoretical analysis of its performance under variability is unknown. This paper presents a Markov Chain model to describe the behavior of Bubble Razor. Using this model, we analyze its performance and provide an optimizing strategy to maximize its benefits. Peter A. Beerel |
DATE | 2 |
| 2014 | Asynchronous circuit placement by lagrangian relaxationabstractRecent asynchronous VLSI circuit placement approach tries to leverage synchronous placement tools as much as possible by manual loop-breaking and creation of virtual clocks. However, this approach produces an exponential number of explicit timing constraints which is beyond the ability of synchronous placement tools to handle. Thus, synchronous placer can only produce suboptimal results. Also, it can be very costly in terms of runtime. This paper proposed a new placement approach for asynchronous VLSI circuits. We formulated the asynchronous timing-driven placement problem and transform this problem into a weighted wirelength minimization problem based on a Lagrangian relaxation framework. The problem can then be efficiently solved using any standard wirelength-driven placement engine that can handle net weights. We demonstrate our approach on QDI PCHB asynchronous circuit with a state-of-art quadratic placer. The experimental results show that our algorithm can effectively improve the asynchronous circuits performance at placement stage. In addition, the runtime of our algorithm is shown to be more scalable to large-scale circuits compared with the loop-breaking approach. Gang Wu 0002, Tao Lin 0007, Hsin-Ho Huang, Chris C. N. Chu, Peter A. Beerel |
ICCAD | 5 |
| 2014 | Performance-Driven Clustering of Asynchronous CircuitsabstractThis paper proposes the method of generating asynchronous circuits from hardware description language specifications by clustering the synthesized gates into asynchronous pipeline stages while preserving liveness, meeting throughput and latency constraints, and minimizing area. This method provides a form of automatic pipelining in which the throughput of the overall design is not limited to the clock frequency or the level of pipelining in the original Register-Transfer Level (RTL) specification. The method is design-style agnostic and is thus applicable to many asynchronous design styles. Georgios D. Dimou, Peter A. Beerel, Andrew Lines |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Area-Efficient Asynchronous Multilevel Single-Track Pipeline TemplateabstractThis paper presents a new asynchronous design theory and a novel template for single-track handshaking that targets medium-to high-performance applications. Unlike other single-track templates, the proposed work supports multiple levels of logic per pipeline stage, improving area efficiency by sharing the control logic among more computation logic while at the same time providing higher robustness to timing variability. The proposed template also yields higher throughput than most four-phase templates and lower latency than bundled-data templates. The template was incorporated into the asynchronous ASIC flow Proteus, and experiments on ISCAS benchmarks show significant improvement in achievable throughput per area. Pankaj Golani, Peter A. Beerel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Slack matching mode-based asynchronous circuits for average-case performanceabstractThis paper addresses the problem of slack matching conditional asynchronous circuits for average-case performance. The behavior of the circuit is modeled using a Markov chain which governs switching between distinct modes of operations with potentially different performance requirements. Given the probability of mode switchings and desired cycle times for each mode, a minimum number of slack-matching buffers is inserted into the circuit such that an upper bound on the overall average cycle time is achieved. The problem is formulated as a Mixed Integer Linear Program and solved through relaxation. Experimental results on a new benchmark of circuits show a significant savings of slack matching buffers compared with the traditional approach and illuminate the type of circuits for which this new formulation is most beneficial. Mehrdad Najibi, Peter A. Beerel |
ICCAD | 2 |
| 2012 | A polynomial time flow for implementing free-choice Petri-netsabstractFSM and PTnet control models are pertinent in both software and hardware applications as both specification and implementation models. The state-based, monolithic FSM model is directly implementable in software or hardware, but cannot model concurrency without state explosion. Interacting FSM models have so far lacked the formal rigor for expressing the synchronising interactions between different FSMs. The event-based, PTnet model is able to model both concurrency and choice within the same model, however lacks a polynomial time flow to implementation, as current methods of exposing the event state space require a potentially exponential number of states. In this work, we present a polynomial complexity flow for transforming a Free-Choice PTnet into a new formalism for Interacting FSMs, i.e Multiple, Synchronised FSMs (MSFSMs), a compact Interacting FSMs model, potentially implementable using any existing monolithic FSM implementation method. We believe that such a flow can in the long term bridge the event and state-based models. We present execution time and state space results of exercising our flow on 25 large PTnet specifications, describing asynchronous control circuits, and contrast our results to the popular Petrify tool for PTnet state space exploration and circuit implementation. Our results indicate a very significant reduction in both state space size and execution time. Pavlos M. Mattheakis, Christos P. Sotiriou, Peter A. Beerel |
ICCD | 3 |
| 2011 | An area-efficient multi-level single-track pipeline templateabstractThis paper presents a new asynchronous design template using single-track handshaking that targets medium-to-high performance applications. Unlike other single-track templates, the proposed template supports multiple levels of logic per pipeline stage, improving area efficiency by sharing the control logic among more logic while at the same time providing higher robustness to timing variability. The template also yields higher throughput than most four-phase templates and lower latency than bundled-data templates. The template has been incorporated into the asynchronous ASIC flow Proteus and experiments on ISCAS benchmarks show significant improvement in achievable throughput per area. Pankaj Golani, Peter A. Beerel |
DATE | 2 |
| 2011 | Energy and Performance Models for Synchronous and Asynchronous CommunicationabstractCommunication costs, which have the potential to throttle design performance as scaling continues, are mathematically modeled and compared for various pipeline methodologies. First-order models are created for common pipeline protocols, including clocked flopped, clocked time-borrowing latch, asynchronous two-phase, four-phase, delay-insensitive, single-track, and source synchronous. The models are parameterized for throughput, energy, and bandwidth. The models share common parameters for different pipeline protocols and implementations to enable a fair apple-to-apple comparison. The accuracy of the models are demonstrated for complete implementations of a subset of the protocols by applying 65-nm process simulated parameter values against the SPICE simulation of full pipeline implementations. One can determine when asynchronous communication is superior at the physical level to synchronous communication in terms of energy for a given bandwidth by applying actual or expected values of the parameters to various design targets. Comparisons between protocols at fixed targets also allow designers to understand tradeoffs between implementations that have a varying process, timing, and design requirements. Kenneth S. Stevens, Pankaj Golani, Peter A. Beerel |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2006 | Pipeline optimization for asynchronous circuits: complexity analysis and an efficient optimal algorithmabstractThis paper addresses the problem of identifying the minimum pipelining needed in an asynchronous circuit (e.g., number/size of pipeline stages/latches required) to satisfy a given performance constraint, thereby implicitly minimizing area and power for a given performance. The paper first shows that the basic pipeline optimization problem for asynchronous circuits is NP-complete. Then, it presents an efficient branch and bound algorithm that finds the optimal pipeline configuration. The experimental results on a few scalable system models demonstrate that this algorithm is computationally feasible for moderately sized models. Sangyun Kim 0001, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | An Asynchronous Low-Power High-Performance Sequential Decoder Implemented With QDI TemplatesabstractThis paper presents the design of a channel-based asynchronous sequential decoder implemented with quasi-delay-insensitive templates. The Powermill simulation results in TSMC 0.25-CMOS technology show that the circuit runs at 430 MHz and consumes 32 mW. Techniques to effectively partition and implement the top level design, the implementation of fast shift registers, memories, and various other structures are discussed. Compared to a previously designed synchronous Fano decoder, the asynchronous version consumes 1/3 the power and runs at 2.15 times the speed assuming standard process normalization. The design also highlights the introduction of a standard-cell library and back-end design flow for asynchronous designs based on precharged half buffer (PCHB) templates Recep O. Ozdag, Peter A. Beerel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Efficient asynchronous bundled-data pipelines for DCT matrix-vector multiplicationabstractThis paper demonstrates the design of efficient asynchronous bundled-data pipelines for the matrix-vector multiplication core of discrete cosine transforms (DCTs). The architecture is optimized for both zero and small-valued data, typical in DCT applications, yielding both high average performance and low average power. The proposed bundled-data pipelines include novel data-dependent delay lines with integrated control circuitry to efficiently implement speculative completion sensing. The control circuits are based on a novel control-circuit template that simplifies the design of such nonlinear pipelines. Extensive post-layout back-end timing analysis was performed to gain confidence in the timing margins as well as to quantify performance and energy. Comparison with a synchronous counterpart suggests that our best asynchronous design yields 30% higher average throughput with negligible energy overhead. Sunan Tugsinavisut, Youpyo Hong, Daewook Kim, Kyeounsoo Kim, Peter A. Beerel |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2003 | Voltage-pulse driven harmonic resonant rail drivers for low-power applicationsabstractWe describe a new design technique for efficient harmonic resonant rail drivers. The proposed circuit implementation is coupled to a standard pulse source and uses only discrete passive components and no external dc power supply. It can thus be externally tuned to minimize the consumed power in the target IC. A new design technique based on current-fed voltage pulse-forming network theory is proposed to find the value of each discrete component for a target frequency and a given load capacitance. The proposed circuit topology can be used to generate any desired periodic 50% duty-cycle waveform by superimposing multiple harmonics of the desired waveform, however, this paper focuses on the generation of trapezoidal-wave clock signals. We have tested the driver with a capacitive load between 38.3 and 97.8 pF with clock frequency ranging between 0.8 and 15 MHz. The overall power dissipation for our second-order harmonic rail driver is 19% of fC/sub L/V/sup 2/ at 15 MHz and 97.8 pF load. Joong-Seok Moon, William C. Athas, Sigfrid D. Soli, Jeffrey T. Draper, Peter A. Beerel |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2002 | Single-Track Asynchronous Pipeline Templates Using 1-of-N EncodingabstractThis paper presents a new fast and templatized family of fine-grain asynchronous pipeline stages based on the single-track protocol. No explicit control wires are required outside of the datapath and the data is 1-of-N encoded. With a forward latency of 2 transitions and a cycle time of 6 for most configurations, the new family can run at 1.6 GHz using MOSIS TSMC 0.25 /spl mu/m process. This is significantly faster than all known quasi-delay-insensitive templates and has less timing assumptions than the recently proposed ultra-high-speed GasP bundled-data circuits. Marcos Ferretti, Peter A. Beerel |
DATE | 2 |
| 2002 | High-Speed Non-Linear Asynchronous PipelinesabstractMany approaches recently proposed for high-speed asynchronous pipelines are applicable only to linear datapaths. However, real systems typically have non-linearities in their datapaths, i.e. stages may have multiple inputs ('joins') or multiple outputs ('forks'). This paper presents several new pipeline templates that extend existing high-speed approaches for linear dynamic logic pipelines, by providing efficient control structures that can accommodate forks and joins. In addition, constructs for conditional computation are also introduced. Timing analysis and SPICE simulations show that the performance overhead of these extensions is fairly low (5% to 20%). Recep O. Ozdag, Peter A. Beerel, Montek Singh, Steven M. Nowick |
DATE | 2 |
| 2002 | Control Circuit Templates for Asynchronous Bundled-Data PipelinesabstractThis paper proposes the use of templatized asynchronous control circuits with single-rail datapaths to create low-power bundled-data non-linear pipelines. First, we adapt an existing templatized control style for 1-of-N rail pipelines, the Pre-Charged Full Buffer PCFB, to bundled-data pipelines. Then, we present a novel true 4-phase template (T4PFB) that has lower control overhead. Simulation results indicate 12%-44% higher throughput for the pipeline stage equivalent to 8 to 40 gates. Sunan Tugsinavisut, Peter A. Beerel |
DATE | 2 |
| 2001 | Theory and practical implementation of harmonic resonant rail driverabstractArticle Share on Theory and practical implementation of harmonic resonant rail driver Authors: Joong-Seok Moon University of Southern California, Los Angeles, CA University of Southern California, Los Angeles, CAView Profile , William Athas Apple Computer Inc., Cupertino, CA Apple Computer Inc., Cupertino, CAView Profile , Peter Beerel University of Southern California, Los Angeles, CA University of Southern California, Los Angeles, CAView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 153–158https://doi.org/10.1145/383082.383122Published:06 August 2001Publication History 2citation169DownloadsMetricsTotal Citations2Total Downloads169Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Joong-Seok Moon, William C. Athas, Peter A. Beerel |
ISLPED | 3 |
| 2001 | A low latency SISO with application to broadband turbo decodingabstractThe standard algorithm for computing the soft-inverse of a finite-state machine [i.e., the soft-in/soft-out (SISO) module] is the forward-backward algorithm. These forward and backward recursions can be computed in parallel, yielding an architecture with latency /spl Oscr/(N), where N is the block size. We demonstrate that the standard SISO computation may be formulated using a combination of prefix and suffix operations. Based on well-known tree-structures for fast parallel prefix computations in the very large scale integration (VLSI) literature (e.g., tree adders), we propose a tree-structured SISO that has latency /spl Oscr/(log/sub 2/N). The decrease in latency comes primarily at a cost of area with, in some cases, only a marginal increase in computation. We discuss how this structure could be used to design a very high throughput turbo decoder or, more generally, an iterative detector. Various subwindowing and tiling schemes are also considered to further improve latency. Peter A. Beerel, Keith M. Chugg |
IEEE J. Sel. Areas Commun. | 1 |
| 2000 | Pipeline Optimization for Asynchronous Circuits: Complexity Analysis and an Efficient Optimal AlgorithmabstractThis paper addresses the problem of identifying the minimal pipelining needed in an asynchronous circuit (e.g., number/size of pipeline stages/latches required) to satisfy a given performance constraint, thereby implicitly minimizing area and power for a given performance. In contrast to the somewhat analogous problem of retiming in the synchronous domain, we first show that the basic pipeline optimization problem for asynchronous circuits is NP-complete. This paper then presents an efficient branch and bound algorithm that can find the optimal pipeline configuration for moderately-sized problems. Our experimental results on a few scalable system models demonstrate that our novel branch and bound solver can find the optimal pipeline configuration for models that have up to 2/sup 35/ possible pipeline configurations. Sangyun Kim 0001, Peter A. Beerel |
ICCAD | 2 |
| 2000 | An asynchronous matrix-vector multiplier for discrete cosine transformabstractThis paper proposes an efficient asynchronous hardwired matrix-vector multiplier for the rwo-dimensional discrete cosine transform and inverse discrete cosine transform (DCT/IDCT). The design achieves low power and high performance by taking advantage of the typically large fraction of zero and small-valued data in DCT and IDCT applications. In particular, it skips multiplication by zero and dynamically activates/deactivates required bit-slices of fine-grain bit-partitioned adders using simplified, static-logic-based speculative completion sensing. The results extracted by both bit-level analysis and HSPICE simulations indicate significant improvements compared to traditional designs. Kyeounsoo Kim, Peter A. Beerel, Youpyo Hong |
ISLPED | 2 |
| 2000 | Sibling-substitution-based BDD minimization using don't caresabstractIn many computer-aided design tools, binary decision diagrams (BDDs) are used to represent Boolean functions. To increase the efficiency and capability of these tools, many algorithms have been developed to reduce the size of the BDDs. This paper presents heuristic algorithms to minimize the size of the BDDs representing incompletely specified functions by intelligently assigning don't cares to binary values. Experimental results show that new algorithms yield significantly smaller BDDs compared with existing algorithms yet still require manageable run-times. These algorithms are particularly useful for synthesis application where the structure of the hardware/software is derived from the BDD representation of the function to implement because the minimization quality is more critical than the minimization speed in these applications. Youpyo Hong, Peter A. Beerel, Jerry R. Burch, Kenneth L. McMillan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Implicit enumeration of strongly connected components and anapplication to formal verificationabstractThis paper first presents a binary decision diagram-based implicit algorithm to compute all maximal strongly connected components (SCCs) of directed graphs. The algorithm iteratively applies reachability analysis and sequentially identifies SCCs. Experimental results suggest that the algorithm dramatically outperforms the only existing implicit method which must compute the transitive closure of the adjacency-matrix of the graphs. This paper then applies this SCC algorithm to solve the bad cycle detection problem encountered in formal verification. Experimental results show that our new bad cycle detection algorithm is typically significantly faster than the state-of-the-art, sometimes by more than a factor of ten. Aiguo Xie, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Symbolic Reachability Analysis of Large Finite State Machines Using Don't CaresabstractReachability analysis of finite state machines is essential to many computer-aided design applications. We present new techniques to improve both approximate and exact reachability analysis using don't cares. First, we propose an iterative approximate reachability analysis technique in which don't care sets derived from previous iterations are used in subsequent iterations for better approximation. Second, we propose new techniques to use the final approximation to enhance the capability and efficiency of exact reachability analysis. Experimental results show that the new techniques can improve reachability analysis significantly. Youpyo Hong, Peter A. Beerel |
DATE | 2 |
| 1999 | Implicit enumeration of strongly connected componentsabstractThis paper presents a binary decision diagram (BDD) based implicit algorithm to compute all maximal strongly connected components (SCCs) of directed graphs. The algorithm iteratively applies reachability analysis and sequentially identifies SCCs. Experiments suggest that the algorithm dramatically outperforms the only existing implicit method which must compute the transitive closure of the adjacency matrix of the graphs. Aiguo Xie, Peter A. Beerel |
ICCAD | 2 |
| 1999 | Statistically optimized asynchronous barrel shifters for variable length codecsabstract: This paper presents low-power asynchronous barrel shifters for variable length encoders and decoders useful in portable applications using multimedia standards. Our approach is to create multi-level asynchronous barrel shifters optimized for the skewed shift control statistics often found in these codecs. For common shifts, data passes through one level, whereas for rare shifts, data passes though multiple levels. We compare our optimized designs with the straight-forward asynchronous and synchronous designs. Both pre- and post-layout HSPICE simulation results indicate that, compared to their synchronous counterparts, our designs provide over a 40% savings in average energy consumption for a given average performance. 1 Introduction Asynchronous circuits sometimes can consume very low average energy for a given average performance partly because of their ability to adapt to variations in chip temperature and voltage supply level. To achieve this goal, however, the asynchronous circu... Peter A. Beerel, Sangyun Kim 0001, Pei-Chuan Yeh, Kyeounsoo Kim |
ISLPED | 1 |
| 1999 | Average-case technology mapping of asynchronous burst-mode circuitsabstractThis paper presents a technology mapper that optimizes the average performance of asynchronous burst-mode control circuits. More specifically, the mapper can be directed to minimize either the average latency or the average cycle time of the circuit. The input to the mapper is a burst-mode specification and its NAND-decomposed unmapped network. The mapper preprocesses the circuit's specification using stochastic techniques to determine the relative frequency of occurrence of each state transition. Then, it maps the NAND-decomposed network using a given library of gates. Of many possible mappings, the mapper selects a solution that minimizes the sum of the delays (latency or cycle time) of all state transitions, weighted by their relative frequencies, thereby optimizing for average performance. We present experimental results on a large set of benchmark circuits, which demonstrate that our mapped circuits have significantly lower average latency and cycle time than comparable circuits mapped with a leading conventional mapping technique which minimizes the worst case delay. Moreover, these performance improvements can be achieved with manageable run-times and significantly smaller area. Wei-Chun Chou, Peter A. Beerel, Kenneth Y. Yun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Accelerating Markovian analysis of asynchronous systems using state compressionabstractThis paper presents a methodology to speed up the stationary analysis of large Markov chains that model asynchronous systems. Instead of directly working on the original Markov chain, we propose to analyze a smaller Markov chain obtained via a novel technique called state compression. Once the smaller chain is solved, the solution to the original chain is obtained via a process called expansion. The method is especially powerful when the Markov chain has a small feedback vertex set, which happens often in asynchronous systems that contain mostly bounded-delay components. Our experimental results show that the method can yield reductions of more than an order of magnitude in CPU time and facilitate the analysis of larger systems than possible using traditional techniques. Aiguo Xie, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Don't Care-Based BDD Minimization for Embedded SoftwareabstractThis paper explores the use of don’t cares in software synthesis for embedded systems. Embedded systems have extremely tight realtime and code/data size constraints, that make expensive optimizations desirable. We propose applying BDD minimization techniques in the presence of a don’t care set to synthesize code for extended Finite State Machines from a BDD-based representation of the FSM transition function. The don’t care set can be derived from local analysis (such as unused state codes or don’t care inputs) as well as from external information (such as impossible input patterns). We show experimental results, discuss their implications, the interaction between BDD-based minimization and dynamic variable reordering, and propose directions for future work. 1 Youpyo Hong, Peter A. Beerel, Luciano Lavagno, Ellen Sentovich |
DAC | 2 |
| 1998 | Efficient State Classification of Finite State Markov ChainsabstractThis paper presents an efficient method for state classification of finite state Markov chains using BDDbased symbolic techniques. The method exploits the fundamental properties of a Markov chain and classifies the state space by iteratively applying reachability analysis. We compare our method with the state-of-the-art technique which requires the transitive closure of the transition relation of a Markov chain. Experiments in over a dozen synchronous, asynchronous systems and queueing networks demonstrate that our method dramatically reduces the CPU time needed, and solves much larger problems because of the reduced memory requirements. I. Introduction Markov chains are an important class of stochastic processes that can effectively model randomly evolving systems. They are conceptually simple because they have short memory, i.e., their future evolution depends only on the current outcome and not on further history. Nonetheless, Markov chains have found vast applications in engineer... Aiguo Xie, Peter A. Beerel |
DAC | 2 |
| 1998 | Checking Combinational Equivalence of Speed-Independent Circuits
Peter A. Beerel, Jerry R. Burch, Teresa H. Meng |
Formal Methods Syst. Des. | 1 |
| 1998 | Covering conditions and algorithms for the synthesis of speed-independent circuitsabstractThis paper presents theory and algorithms for the synthesis of standard C-implementations of speed-independent circuits. These implementations are block-level circuits which may consist of atomic gates to perform complex functions in order to ensure hazard freedom. First, we present Boolean covering conditions that guarantee that the standard C-implementations operate correctly. Then, we present two algorithms that produce optimal solutions to the covering problem. The first algorithm is always applicable, but does not complete on large circuits. The second algorithm, motivated by our observation that our covering problem can often be solved with a single cube, finds the optimal single-cube solution when such a solution exists. When applicable, the second algorithm is dramatically more efficient than the first, more general algorithm. We present results for benchmark specifications which indicate that our single-cube algorithm is applicable on most benchmark circuits and reduces run times by over an order of magnitude. The block-level circuits generated by our algorithms are a good starting point for tools that perform technology mapping to obtain gate-level speed-independent circuits. Peter A. Beerel, Chris J. Myers, Teresa H. Meng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Efficient state classification of finite-state Markov chainsabstractThis paper presents an efficient method for state classification of finite-state Markov chains using binary-decision diagram-based symbolic techniques. The method exploits the fundamental properties of a Markov chain and classifies the state space by iteratively applying reachability analysis. We compare our method with the state-of-the-art technique, which requires the transitive closure of the transition relation of a Markov chain. Experiments in over a dozen synchronous and asynchronous systems and queueing networks demonstrate that our method dramatically reduces the CPU time needed and solves much larger problems because of the reduced memory requirements. Aiguo Xie, Peter A. Beerel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | The design and verification of a high-performance low-control-overhead asynchronous differential equation solverabstractThis paper describes the design and verification of a high-performance asynchronous differential equation solver benchmark circuit. The design has low-control-overhead which allows its average-case speed (tested at 22/spl deg/C and 3.3 V) to be 48% faster than any comparable synchronous design (designed to operate at 100/spl deg/C and 3 V for the slow process corner). The techniques to reduce completion sensing overhead and hide control overhead at the circuit, architectural, and protocol levels are discussed. In addition, symbolic model checking techniques are described that were used to gain higher confidence in the correctness of the timed distributed control. Kenneth Y. Yun, Peter A. Beerel, Vida Vakilotojar, Ayoob E. Dooply, Julio Arceo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1997 | RTL verification of timed asynchronous and heterogeneous systems using symbolic model checkingabstractThis paper describes a tool-supported methodology for the register-transfer-level formal verification of a growing hardware design paradigm-timed asynchronous systems. These systems are a network of communicating asynchronous and synchronous components and have correctness constraints that depend on specified bounded delays. This paper formalizes the verification problem and demonstrates how time-discretization, abstraction, and non-determinism can lead to a system model comprised of communicating finite state machines composed synchronously. The paper then describes a translator that accepts structural VHDL system description along with controller specifications and generates the input to a symbolic model checker (SMV). Finally, we describe two case studies in which concurrent verification and design led to the correction of many errors not easily found using simulation. Vida Vakilotojar, Peter A. Beerel |
ASP-DAC | 2 |
| 1997 | Safe BDD Minimization Using Don't CaresabstractIn many computer-aided design tools, binary decision diagrams(BDDs) are used to represent Boolean functions. To increase theefficiency and capability of these tools, many algorithms have beendeveloped to reduce the size of BDDs. This paper presents heuristicalgorithms that minimize the size of BDDs representing incompletelyspecified functions by intelligently assigning don't cares tobinary values. The traditional algorithm, restrict [Verification of Synchronous Sequential Machines Based on Symbolic Execution], is often effectivein BDD minimization, but can increase the BDD size. We proposenew algorithms based on restrict which are guaranteed neverto increase the size of the BDD, thereby significantly reducing peakmemory requirements. Experimental results show that our techniquestypically yield significantly smaller BDDs than restrict. Youpyo Hong, Peter A. Beerel, Jerry R. Burch, Kenneth L. McMillan |
DAC | 2 |
| 1997 | RTL verification of timed asynchronous and heterogeneous systems using symbolic model checking
Vida Vakilotojar, Peter A. Beerel |
Integr. | 2 |
| 1996 | Estimation of energy consumption in speed-independent control circuitsabstractWe describe a technique to estimate the energy consumed by speed-independent asynchronous (clock-less) control circuits. Because speed-independent circuits are hazard-free under all possible combinations of gate delays, we prove that an accurate estimate of their energy consumption is independent of relative component gate delays and can be determined by simulating only a small number of input patterns proportional to the size of the circuit's Signal Transition Graph specification. Specifically, we calculate the average energy per external signal transition consumed by a circuit. This can be used to compare the energy consumption between two different circuit implementations of the same specification, to calculate average energy for a given high-level operation, and to provide average circuit power when combined with delay information. Peter A. Beerel, Cheng-Ta Hsieh, Suhrid A. Wadekar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1995 | Estimation and bounding of energy consumption in burst-mode control circuitsabstractThis paper describes two techniques to quantify energy consumption of burst-mode asynchronous (clock-less) control circuits. The circuit specifications considered are extended burst-mode specifications, and the implementations are multi-level logic implementations whose outputs are guaranteed to be free of any voltage glitches (hazards). Both techniques use stochastic analysis to combine a small number of simulations in order to quantify average energy per external signal transition. The first technique uses N-valued simulation to derive mathematically tight upper and lower bounds of energy consumption. Using this technique we bound the effect of hazards under all possible operating conditions and environments for a given circuit. Additionally, to drive synthesis tools for low-power we propose a second technique that uses fixed-delay simulation to derive a realistic estimate of energy consumption within our derived upper and lower bounds. We demonstrate the feasibility of both these techniques on a variety of burst-mode control circuits used in an industrial-quality chip. Our preliminary results indicate that less than 5% of the power of typical multi-level burst-mode circuits can be attributed to hazards. Peter A. Beerel, Kenneth Y. Yun, Steven M. Nowick, Pei-Chuan Yeh |
ICCAD | 1 |
| 1993 | Efficient verification of determinate speed-independent circuitsabstractWe present sufficient conditions for the correctness of speed-independent circuits with respect to their state graph (SG) specification, which can be tested in linear-time with respect to the size of the SG. Our correctness conditions consist of one safety condition and one progress condition. The progress condition detects deadlock conditions that are not present in the specification. The SG specifications considered are determinate, allowing input choice (conditionals) but not output choice (arbitration). The circuits considered are a network of basic gates; arbiters and mutual-exclusion elements are not allowed. We present an efficient algorithm to test the correctness conditions, in which false positives are not possible, but false negatives are possible. We have implemented the algorithm and present a table of run-time comparisons between our verification tool and the tool AVER by D. Dill on a large benchmark of asynchronous circuits. The results demonstrate run-times of over an order of magnitude faster than AVER and no false negatives were found. Our speedup is achieved by avoiding the state explosion problem caused by explicitly examining the behavior of internal signals. Peter A. Beerel, Jerry R. Burch, Teresa H. Meng |
ICCAD | 1 |
| 1992 | Automatic gate-level synthesis of speed-independent circuitsabstractA CAD tool for the synthesis of asynchronous control circuits using basic gates such as AND gates and OR gates is presented. The synthesized circuits are speed-independent-that is, they work correctly regardless of individual gate delays. Synthesis results for a variety of specifications taken from industry and previously published examples are presented. The speed-independent circuits are compared with those non-speed-independent circuits synthesized using previously described algorithms, in which delay elements are added to remove circuit hazards. These synthesis results show that the new circuits are on average approximately 25% faster with an area penalty of only 15%. This work demonstrates that direct synthesis of gate-level speed-independent circuits is not only feasible, but also produces robust and relatively efficient circuits compared to those synthesized with timing constraints.> Peter A. Beerel, Teresa H. Meng |
ICCAD | 1 |
| 1992 | Semi-modularity and testability of speed-independent circuits
Peter A. Beerel, Teresa H. Meng |
Integr. | 1 |
| 1991 | Testability of Asynchronous Timed Control Circuits with Delay AssumptionsabstractThis paper addresses the testability of “timed” asynchronous control circuits built of standard logic cells, in which each gate’s rise and fall time is associated with a pair of minimum and maximum values. The circuit that we analyze is assumed to be hazardfie. given that the gate delay assumptions are met in its implementation. We first give sufficient and necessary conditions under which the affects of single stuck-at-faults (SAFs) can be characterized. Then we provide sufficient conditions which ensure that these faults can be detected without access to the memory elements. Finally we present an automated tool which analyzes the testability of “timed” asynchronous circuits based on the derived conditions. Peter A. Beerel, Teresa H. Meng |
DAC | 1 |