EDBT 2026 Demo / reviewers in the wild / expert
Gaurav Singh 0008
dblp:07/2202-8
· DBLP profile ↗
6ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-5232-8145ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Compressing Neural Networks using Learnable 1D Non-Linear FunctionsabstractAs deep learning models grow in size to achieve state-of-the-art accuracy, there is a pressing need for compact models. To address this challenge, we introduce a novel operation called Personal Self-Attention (PSA). It is specifically designed to learn non-linear 1D functions, enhancing existing spline-based methods while remaining compatible with gradient backpropagation. By integrating these non-linear functions with linear transformations, we can achieve the accuracy of larger models but with significantly smaller hidden dimensions, which is crucial for FPGA implementations. We evaluate PSA by implementing it in a Multi-Layer Perceptron (MLP)-based vision model, ResMLP, and testing it on the CIFAR-10 classification task. MLP is gaining increasing popularity due to its widespread use in large-language models. Our results confirm that PSA achieves equivalent accuracy with a 2 \(\times\) smaller hidden size compared to conventional MLPs. Furthermore, by quantizing our non-linear function into a simple Lookup Table (LUT), we reduce the number of operations required by 45–28%, which offers significant benefits for hardware accelerators. To showcase this, we design an end-to-end unrolled streaming accelerator for ResMLP, demonstrating that our compressed model maintains an 88% accuracy while reducing LUT \(+\) DSP resource requirements by 25%, and doubling throughput to 32 kFPS. Additionally, we implement a fixed-size SIMD accelerator for the same compressed model that achieves a 62.1% improvement in throughput while only consuming 3.5% extra LUTs. Gaurav Singh 0008, Kia Bazargan |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2024 | SimBU: Self-Similarity-Based Hybrid Binary-Unary Computing for Nonlinear FunctionsabstractUnary computing is a relatively new method for implementing arbitrary nonlinear functions that uses unpacked thermometer number encoding, enabling much lower hardware costs. In its original form, unary computing provides no trade-off between accuracy and hardware cost. In this work, we propose a novel self-similarity-based method to optimize the previous hybrid binary-unary work and provide it with the trade-off between accuracy and hardware cost by introducing controlled levels of approximation. Looking for self-similarity between different parts of a function allows us to implement a very small subset of core unique subfunctions and derive the rest of the subfunctions from this core using simple linear transformations. We compare our method to previous works such as FloPoCo-LUT (lookup table), HBU (hybrid binary-unary) and FloPoCo-PPA (piecewise polynomial approximation) on several 8–12-bit nonlinear functions including Log, Exp, Sigmoid, GELU, Sin, and Sqr, which are frequently used in neural networks and image processing applications. The area$\times$delay hardware cost of our method is on average 32%–60% better than previous methods in both exact and approximate implementations. We also extend our method to multivariate nonlinear functions and show on average 78%–92% improvement over previous work. Alireza Khataei, Gaurav Singh 0008, Kia Bazargan |
IEEE Trans. Computers | 2 |
| 2023 | Optimizing Hybrid Binary-Unary Hardware Accelerators Using Self-Similarity MeasuresabstractUnary computing is a relatively new method for implementing non-linear functions using few hardware resources compared to binary computing. In its original form, unary computing provides no trade-off between accuracy and hardware cost. In this work, we propose a novel self-similarity-based method to optimize the previous hybrid binary-unary method and provide it with the trade-off between accuracy and hardware cost by introducing controlled levels of approximation. Given a target maximum error, our method breaks a function into sub-functions and tries to find the minimum set of unique sub-functions that can derive all the other ones through trivial bit-wise transformations. We compare our method to previous works such as HBU (hybrid binary-unary) and FloPoCo-PPA (piece-wise polynomial approximation) on a number of non-linear functions including Log, Exp, Sigmoid, GELU, Sin, and Sqr, which are used in neural networks and image processing applications. Without any loss of accuracy, our method can improve the area-delay-product hardware cost of HBU on average by 7% at 8-bit, 20% at 10-bit, and 35% at 12-bit resolutions. Given the approximation of the least significant bit, our method reduces the hardware cost of HBU on average by 21% at 8-bit, 49% at 10-bit, and 60% at 12-bit resolutions, and using the same error budget as given to FloPoCo-PPA, it reduces the hardware cost of FloPoCo-PPA on average by 79% at 8-bit, 58% at 10-bit, and 9% at 12-bit resolutions. We finally show the benefits of our method by implementing a 10-bit homomorphic filter, which is used in image processing applications. Our method can implement the filter with no quality loss at lower hardware cost than what the previous approximate and exact methods can achieve. Alireza Khataei, Gaurav Singh 0008, Kia Bazargan |
FCCM | 2 |
| 2023 | Approximate Hybrid Binary-Unary Computing with Applications in BERT Language Model and Image ProcessingabstractWe propose a novel method for approximate hardware implementation of univariate math functions with significantly fewer hardware resources compared to previous approaches. Examples of such functions include exp(x) and the activation function GELU(x), both used in transformer networks, gamma(x), which is used in image processing, and other functions such as tanh(x), cosh(x), sq(x), and sqrt(x). The method builds on previous works on hybrid binary-unary computing. The novelty in our approach is that we break a function into a number of sub-functions such that implementing each sub-function becomes cheap, and converting the output of the sub-functions to binary becomes almost trivial. Our method also uses self-similarity in functions to further reduce the cost. We compare our method to the conventional binary, previous stochastic computing, and hybrid binary-unary methods on several functions at 8-, 12-, and 16-bit resolutions. While preserving high accuracy, our method outperforms previous works in terms of hardware cost, e.g., tolerating less than 0.01 mean absolute error, our method reduces the (area x latency) cost on average by 5, 7, and 2 orders of magnitude, compared to the conventional binary, stochastic computing, and hybrid binary-unary methods, respectively. Ultimately, we demonstrate the potential benefits of our method for natural language processing and image processing applications. We deploy our method to implement major blocks in an encoding layer of BERT language model, and also the Roberts Cross edge detection algorithm. Both include non-linear functions. Alireza Khataei, Gaurav Singh 0008, Kia Bazargan |
FPGA | 2 |
| 2020 | HBUCNNA: Hybrid Binary-Unary Convolutional Neural Network AcceleratorabstractConvolutional layers account for 90% of the total computational power of Convolutional Neural Networks (CNNs). Field programmable gate arrays (FPGAs) have shown great potential for accelerating inference tasks in CNNs. However, it is harder for FPGA platforms to deliver the best performance due to high computational power and high memory bandwidth requirements of today's CNNs. In this paper, we propose a reconfigurable parallel-pipelined Hybrid Binary-Unary CNN Accelerator (HBUCNNA) to implement low-cost, high-performance convolutional layers of a ResNet-18 architecture. We use the hybrid binary-unary method to implement banks of constant-coefficient multipliers, which are used to implement convolutional kernels. Moreover, we propose hybrid binary-unary batch normalization units to further improve the total hardware costs. These two units reduce {area, area×delay} costs by {50%, 30%} and {44%, 65%} on average compared to their conventional binary counterparts, respectively. The proposed accelerator stores control signals for reconfigurability instead of the numeric value of weights, which in turn reduces the memory footprint on average by 20%. Overall, the proposed HBUCNNA architecture reduces the {area, latency, power, energy, area×delay} costs on average by {25.5%, 40%, 15%, 40%, 47%} and {53%, 61%, 47%, 62%, 67%} compared to the constant-coefficient multiplier-based and variable size multiplier-based binary architectures, respectively. Moreover, the proposed accelerator improves the throughput by about 1.4 × compared to both of the mentioned architectures. S. Rasoul Faraji, Pierre Abillama, Gaurav Singh 0008, Kia Bazargan |
ISCAS | 3 |
| 2019 | HBUNN - Hybrid Binary-Unary Neural Network: Realizing a Complete CNN on an FPGAabstractMost modern neural networks use the basic Multiply-Accumulate (MAC) Operation in some form or another, and as networks get larger, the computational needs for these larger networks grow rapidly. Typically neural networks implemented on FPGAs use variable multipliers so that any weight can be used in the MAC, but this also forces the accelerator designer to store the weights of the model in off-chip memory (DRAM). Moreover, because of high computational power and high memory bandwidth requirements of today's CNNs, it is harder for FPGA platforms to deliver the best performance. In this paper, we propose a fully parallel-pipeline Hybrid Binary-Unary Neural Network (HBUNN) architecture to implement a low-cost and high-performance ResNet-18 convolutional neural network. We use a hybrid binary-unary method to implement constant-coefficient multipliers and batch normalization units. These two units reduce hardware cost by 30.7% and 47.97% on average compared to the conventional binary equivalent, respectively. Moreover, we propose a novel training scheme using our hardware cost-aware regularizers that not only improves the area cost of the proposed architecture and the conventional binary architecture by 59.3% and 76.7% respectively, but also maintains the same accuracy. Finally, we have implemented three trained networks using different regularizers. The proposed HBUNN architectures reduce the area cost by 30%, and the area × delay cost by 69% on average compared to the conventional binary architectures. The error rate of the proposed work is 12.93%, while its throughput is 278 Kfps. Sayed Abdolrasouol Faraji, Gaurav Singh 0008, Kia Bazargan |
ICCD | 2 |