Arash Ardakani

dblp:133/4133 · DBLP profile ↗
← Back
17ranked-venue papers
17as first author
4since 2021 · last 2025
0000-0003-3274-2394ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 11 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 DEMOTIC: A Differentiable Sampler for Multi-Level Digital Circuits
abstract
Efficient sampling of satisfying formulas for circuit satisfiability (CircuitSAT), a well-known NP-complete problem, is essential in modern front-end applications for thorough testing and verification of digital circuits. Generating such samples is a hard computational problem due to the inherent complexity of digital circuits, size of the search space, and resource constraints involved in the process. Addressing these challenges has prompted the development of specialized algorithms that heavily rely on heuristics. However, these heuristic-based approaches frequently encounter scalability issues when tasked with sampling from a larger number of solutions, primarily due to their sequential nature. Different from such heuristic algorithms, we propose a novel differentiable sampler for multi-level digital circuits, called Demotic, that utilizes gradient descent (GD) to solve the CircuitSAT problem and obtain a wide range of valid and distinct solutions. Demotic leverages the circuit structure of the problem instance to learn valid solutions using GD by re-framing the CircuitSAT problem as a supervised multi-output regression task. This differentiable approach allows bit-wise operations to be performed independently on each element of a tensor, enabling parallel execution of learning operations, and accordingly, GPU-accelerated sampling with significant runtime improvements compared to state-of-the-art heuristic samplers. We demonstrate the superior runtime performance of Demotic in the sampling task across various CircuitSAT instances from the ISCAS-85 benchmark suite. Specifically, Demotic outperforms the state-of-the-art sampler by more than two orders of magnitude in most cases.
Arash Ardakani, Kevin He, Qijing Huang 0001, Vighnesh M. Iyer, Suhong Moon, John Wawrzynek
ASP-DAC1
2025 High-Throughput SAT Sampling
abstract
In this work, we present a novel technique for GPU-accelerated Boolean satisfiability (SAT) sampling. Unlike conventional sampling algorithms that directly operate on conjunctive normal form (CNF), our method transforms the logical constraints of SAT problems by factoring their CNF representations into simplified multilevel, multi-output Boolean functions. It then leverages gradient-based optimization to guide the search for a diverse set of valid solutions. Our method operates directly on the circuit structure of refactored SAT instances, reinterpreting the SAT problem as a supervised multi-output regression task. This differentiable technique enables independent bit-wise operations on each tensor element, allowing parallel execution of learning processes. As a result, we achieve GPU-accelerated sampling with significant runtime improvements ranging from 33.6x to 523.6x over state-of-the-art heuristic samplers. We demonstrate the superior performance of our sampling method through an extensive evaluation on 60 instances from a public domain benchmark suite utilized in previous studies.
Arash Ardakani, Kevin He, Qijing Huang 0001, John Wawrzynek
DATE1
2024 Late Breaking Results: Differential and Massively Parallel Sampling of SAT Formulas
abstract
Diverse solutions to the Boolean satisfiability (SAT) problem are essential for thorough testing and verification of software and hardware designs, ensuring reliability and applicability to real-world scenarios. We introduce a novel differentiable sampling method, called DiffSampler, which employs gradient descent (GD) to learn diverse solutions to the SAT problem. By formulating SAT as a supervised multi-output regression task and minimizing its loss function using GD, our approach enables performing the learning operations in parallel, leading to GPU-accelerated sampling and comparable run time performance w.r.t. heuristic samplers. We demonstrate that DiffSampler can generate diverse uniform-like solutions similar to conventional samplers.
Arash Ardakani, Kevin He, Vighnesh M. Iyer, Suhong Moon, John Wawrzynek
DAC1
2024 SlimFit: Memory-Efficient Fine-Tuning of Transformer-based Models Using Training Dynamics
abstract
Arash Ardakani, Altan Haan, Shangyin Tan, Doru Thom Popovici, Alvin Cheung, Costin Iancu, Koushik Sen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Arash Ardakani, Altan Haan, Shangyin Tan, Doru-Thom Popovici, Alvin Cheung, Costin Iancu, Koushik Sen
NAACL-HLT1
2020 A Regression-Based Method to Synthesize Complex Arithmetic Computations on Stochastic Streams
abstract
In stochastic computing, values are represented as sequences of random bits and arithmetic computations are computed on the bit streams. Since bit-wise operations are performed on random bit streams, stochastic computing offers low-cost error-tolerant architectures for its hardware implementations. In stochastic computing, complex arithmetic operations can be computed using linear finite state machines (FSMs). However, the synthesis of a linear FSM for a given target function is nontrivial. In this paper, we exploit linear regression and demonstrate a general approach to synthesize linear FSMs for stochastic computations. We show that our approach outperforms traditional numerical synthesis methods in terms of mean-squared error. We also demonstrate that fault-tolerance of FSMs synthesized using linear regression can be improved by injecting noise during the synthesis phase, allowing the synthesized functions to tolerate up to 35% of random bit flips.
Arash Ardakani, Amir Ardakani, Warren J. Gross
ISCAS1
2020 Training Linear Finite-State Machines
abstract
A finite-state machine (FSM) is a computation model to process binary strings in sequential circuits. Hence, a single-input linear FSM is conventionally used to implement complex single-input functions , such as tanh and exponentiation functions, in stochastic computing (SC) domain where continuous values are represented by sequences of random bits. In this paper, we introduce a method that can train a multi-layer FSM-based network where FSMs are connected to every FSM in the previous and the next layer. We show that the proposed FSM-based network can synthesize multi-input complex functions such as 2D Gabor filters and can perform non-sequential tasks such as image classifications on stochastic streams with no multiplication since FSMs are implemented by look-up tables only. Inspired by the capability of FSMs in processing binary streams, we then propose an FSM-based model that can process time series data when performing temporal tasks such as character-level language modeling. Unlike long short-term memories (LSTMs) that unroll the network for each input time step and perform back-propagation on the unrolled network, our FSM-based model requires to backpropagate gradients only for the current input time step while it is still capable of learning long-term dependencies. Therefore, our FSM-based model can learn extremely long-term dependencies as it requires 1/l memory storage during training compared to LSTMs, where l is the number of time steps. Moreover, our FSM-based model reduces the power consumption of training on a GPU by 33% compared to an LSTM model of the same size.
Arash Ardakani, Amir Ardakani, Warren J. Gross
NeurIPS1
2020 Fast and Efficient Convolutional Accelerator for Edge Computing
abstract
Convolutional neural networks (CNNs) are a vital approach in machine learning. However, their high complexity and energy consumption make them challenging to embed in mobile applications at the edge requiring real-time processes such as smart phones. In order to meet the real-time constraint of edge devices, recently proposed custom hardware CNN accelerators have exploited parallel processing elements (PEs) to increase throughput. However, this straightforward parallelization of PEs and high memory bandwidth require high data movement, leading to large energy consumption. As a result, only a certain number of PEs can be instantiated when designing bandwidth-limited custom accelerators targeting edge devices. While most bandwidth-limited designs claim a peak performance of a few hundred giga operations per second, their average runtime performance is substantially lower than their roofline when applied to state-of-the-art CNNs such as AlexNet, VGGNet and ResNet, as a result of low resource utilization and arithmetic intensity. In this work, we propose a zero-activation-skipping convolutional accelerator (ZASCA) that avoids noncontributory multiplications with zero-valued activations. ZASCA employs a dataflow that minimizes the gap between its average and peak performances while maximizing its arithmetic intensity for both sparse and dense representations of activations, targeting the bandwidth-limited edge computing scenario. More precisely, ZASCA achieves a performance efficiency of up to 94 percent over a set of state-of-the-art CNNs for image classification with dense representation where the performance efficiency is the ratio between the average runtime performance and the peak performance. Using its zero-skipping feature, ZASCA can further improve the performance efficiency of the state-of-the-art CNNs by up to 1.9× depending on the sparsity degree of activations. The implementation results in 65-nm TSMC CMOS technology show that, compared to the most energy-efficient accelerator, ZASCA can process convolutions from 5.5× to 17.5× faster, and is between 2.1× and 4.5× more energy efficient while occupying 2.1× less silicon area.
Arash Ardakani, Carlo Condo, Warren J. Gross
IEEE Trans. Computers1
2019 Learning to Skip Ineffectual Recurrent Computations in LSTMs
abstract
Long Short-Term Memory (LSTM) is a special class of recurrent neural network, which has shown remarkable successes in processing sequential data. The typical architecture of an LSTM involves a set of states and gates: the states retain information over arbitrary time intervals and the gates regulate the flow of information. Due to the recursive nature of LSTMs, they are computationally intensive to deploy on edge devices with limited hardware resources. To reduce the computational complexity of LSTMs, we first introduce a method that learns to retain only the important information in the states by pruning redundant information. We then show that our method can prune over 90% of information in the states without incurring any accuracy degradation over a set of temporal tasks. This observation suggests that a large fraction of the recurrent computations are ineffectual and can be avoided to speed up the process during the inference as they involve noncontributory multiplications/accumulations with zero-valued states. Finally, we introduce a custom hardware accelerator that can perform the recurrent computations using both sparse and dense states. Experimental measurements show that performing the computations using the sparse states speeds up the process and improves energy efficiency by up to 5.2× when compared to implementation results of the accelerator performing the computations using dense states.
Arash Ardakani, Zhengyun Ji, Warren J. Gross
DATE1
2019 Learning Recurrent Binary/Ternary Weights
Arash Ardakani, Zhengyun Ji, Sean C. Smithson, Brett H. Meyer, Warren J. Gross
ICLR (Poster)1
2019 The Synthesis of XNOR Recurrent Neural Networks with Stochastic Logic
abstract
The emergence of XNOR networks seek to reduce the model size and computational cost of neural networks for their deployment on specialized hardware requiring real-time processes with limited hardware resources. In XNOR networks, both weights and activations are binary, bringing great benefits to specialized hardware by replacing expensive multiplications with simple XNOR operations. Although XNOR convolutional and fully-connected neural networks have been successfully developed during the past few years, there is no XNOR network implementing commonly-used variants of recurrent neural networks such as long short-term memories (LSTMs). The main computational core of LSTMs involves vector-matrix multiplications followed by a set of non-linear functions and element-wise multiplications to obtain the gate activations and state vectors, respectively. Several previous attempts on quantization of LSTMs only focused on quantization of the vector-matrix multiplications in LSTMs while retaining the element-wise multiplications in full precision. In this paper, we propose a method that converts all the multiplications in LSTMs to XNOR operations using stochastic computing. To this end, we introduce a weighted finite-state machine and its synthesis method to approximate the non-linear functions used in LSTMs on stochastic bit streams. Experimental results show that the proposed XNOR LSTMs reduce the computational complexity of their quantized counterparts by a factor of 86x without any sacrifice on latency while achieving a better accuracy across various temporal tasks.
Arash Ardakani, Zhengyun Ji, Amir Ardakani, Warren J. Gross
NeurIPS1
2018 A Convolutional Accelerator for Neural Networks With Binary Weights
abstract
Parallel processors and GP-GPUs have been routinely used in the past to perform the computations of convolutional neural networks (CNNs). However, their large power consumption has pushed researchers towards application-specific integrated circuits and on-chip accelerators implement neural networks. Nevertheless, within the Internet of Things (IoT) scenario, even these accelerators fail to meet the power and latency constraints. To address this issue, binary-weight networks were introduced, where weights are constrained to -1 and 1. Therefore, these networks facilitate hardware implementation of neural networks by replacing multiply-and-accumulate units with simple accumulators, as well as reducing the weight storage. In this paper, we introduce a convolutional accelerator for binary-weight neural networks. The proposed architecture only consumes 128 mW at a frequency of 200 MHz and occupies 1.2 mm2when synthesized in TSMC 65 nm CMOS technology. Moreover, it achieves a high area-efficiency of 176 Gops/MGC and performance efficiency of 89%, outperforming the state-of-the-art architecture for binary-weight networks by 1.8× and 3.2×, respectively.
Arash Ardakani, Carlo Condo, Warren J. Gross
ISCAS1
2017 Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks
Arash Ardakani, Carlo Condo, Warren J. Gross
ICLR (Poster)1
2017 A low-complexity fully scalable interleaver/address generator based on a novel property of QPP interleavers
abstract
5-th generation mobile networks aim the peak data rates in excess of few Gbs, which may appear to be challenging to achieve due to the existence of some blocks such as the turbo decoder. In fact, the interleaver is known to be a major challenging part of the turbo decoder due to its need to the parallel interleaved memory access. LTE uses Quadratic Permutation Polynomial (QPP) interleaver, which makes it suitable for the parallel decoding. In this paper, a new property of the QPP interleaver, called the correlated shifting property, is theoretically proved, leading to a fully scalable interleaver and a low-complexity address generator for an arbitrary order of parallelism. The proposed interleaver reduces the required addresses in half. Moreover, the scalability of the proposed interleaver proves up to 51% lower power consumption compared to the best reported interleaver to-date.
Arash Ardakani, Mahdi Shabany
ISCAS1
2017 VLSI Implementation of Deep Neural Network Using Integral Stochastic Computing
abstract
The hardware implementation of deep neural networks (DNNs) has recently received tremendous attention: many applications in fact require high-speed operations that suit a hardware implementation. However, numerous elements and complex interconnections are usually required, leading to a large area occupation and copious power consumption. Stochastic computing (SC) has shown promising results for low-power area-efficient hardware implementations, even though existing stochastic algorithms require long streams that cause long latencies. In this paper, we propose an integer form of stochastic computation and introduce some elementary circuits. We then propose an efficient implementation of a DNN based on integral SC. The proposed architecture has been implemented on a Virtex7 field-programmable gate array, resulting in 45% and 62% average reductions in area and latency compared with the best reported architecture in the literature. We also synthesize the circuits in a 65-nm CMOS technology, and we show that the proposed integral stochastic architecture results in up to 21% reduction in energy consumption compared with the binary radix implementation at the same misclassification rate. Due to fault-tolerant nature of stochastic architectures, we also consider a quasi-synchronous implementation that yields 33% reduction in energy consumption with respect to the binary radix implementation without any compromise on performance.
Arash Ardakani, François Leduc-Primeau, Naoya Onizawa, Takahiro Hanyu, Warren J. Gross
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Hardware implementation of FIR/IIR digital filters using integral stochastic computation
abstract
Stochastic computing (SC) has received much recent attention due to its inherent fault-tolerance and low implementation cost compared to binary radix representations. SC has been proposed for various signal processing applications such as digital filters. The prior art in stochastic FIR filters can accurately implement the desired filtering function for low-order filters, however, their accuracy degrades as the filter order increases. Moreover, stochastic IIR filters demonstrate high hardware complexity and degraded accuracy. In this paper, we propose an architecture for high-order FIR filters with negligible accuracy loss compared to fixed-point implementation. The proposed architecture requires fewer random number generators. We also describe a novel cascaded second-order direct-form II structure for IIR filters. The implementation results of the proposed design show an improvement in latency and hardware complexity compared to the stochastic architectures reported to date.
Arash Ardakani, François Leduc-Primeau, Warren J. Gross
ICASSP1
2015 An efficient max-log MAP algorithm for VLSI implementation of turbo decoders
abstract
Long term evolution (LTE)-advanced aims the peak data rates in excess of 3 Gbps for the next generation wireless communication systems. Turbo codes, the specified channel coding scheme in LTE, suffers from a low-decoding throughput due to its iterative decoding algorithm. One efficient approach to achieve a promising throughput is to use multiple Maximum a Posteriori (MAP) cores in parallel, resulting in a large area overhead, a big drawback. The scaled Max-log MAP algorithm is a common approach to implement the MAP algorithm due to its efficient architecture with its acceptable performance. Although many works have been reported to reduce the area of the MAP unit, an efficient VLSI architecture with minimum silicon area is still missing. To address this challenge, in this paper, a novel method based on a novel form of computation is introduced, removing a single bit from computations of the max-log MAP algorithm compared to the conventional architectures. The proposed method can be applied to any max-log MAP architecture reported to-date.
Arash Ardakani, Mahdi Shabany
ISCAS1
2013 An efficient VLSI architecture of QPP interleaver/deinterleaver for LTE turbo coding
abstract
Long Term Evolution (LTE) supports peak data rates in excess of 300 Mb/s. A good approach to achieve such rates is by parallelizing the required processing in turbo decoders. An interleaver is an important part of a turbo decoder. LTE uses the Quadratic Permutation Polynomial (QPP) interleaver, which makes it suitable for parallel decoding. In this paper, we propose an efficient architecture for the QPP interleaver, called the Add-Compare-Select (ACS) permuting network. A unique feature of the proposed architecture is that it can be used both as the interleaver and deinterleaver leading to a high-speed low-complexity hardware interleaver/deinterleaver for turbo decoding. The proposed design requires no memory or QPP inverse to perform deinterleaving and has been fully implemented and tested both on a Virtex-6 FPGA as well as in a 0.18 um CMOS process.
Arash Ardakani, Mahdi Shabany
ISCAS1