EDBT 2026 Demo / reviewers in the wild / expert
Haroon Waris
dblp:235/0721
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-4670-3919ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Algorithm-Aided Design and Verification for Multiple-Node-Upset-Recovery Latches
Zhiyuan Pei, Haroon Waris, M. Shahzad Younis, Aibin Yan, Jiahui Deng, Yilin Gui, Xiaoqing Wen |
ISCAS | 3 |
| 2026 | Approximate Computing-Based Framework for Low-Cost Runtime Hardware Trojan DefenseabstractWith the growing reliance on third-party intellectual property (3PIP) in integrated circuit design, the threat of functional failures induced by hardware Trojans embedded within these components has become increasingly critical. The high stealthiness and sophistication of hardware Trojans often render existing detection techniques insufficient for comprehensive coverage. As a result, runtime detection and recovery techniques have emerged and been proposed as a crucial last line of defense. However, these solutions typically introduce substantial hardware overhead, limiting their practicality in resource-constrained applications such as edge computing. To address this critical challenge, this work leverages approximate computing to develop low-cost runtime recovery strategies against hardware Trojans. Specifically, it begins by analyzing the challenges and potential opportunities introduced into existing security schemes when approximate computing is applied. Based on this analysis, targeted solutions and optimization techniques are proposed. These are then integrated into a unified framework that explores the trade-off between hardware resource usage and computational accuracy while maintaining circuit-level security. Experimental results across several commonly used applications demonstrate that the proposed framework can achieve over 20% of hardware resource savings with only a 5% reduction in accuracy. To the best of our knowledge, this work is the first to systematically integrate approximate computing with runtime hardware Trojan recovery, providing a new cost-effective direction for circuit-level security in resource-constrained systems. Yuqin Dou, Yang Wang 0142, Shiquan Liu, Haroon Waris, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | FPGA-Based Low-Power Signed Approximate Multipliers for Diverse Error-Resilient ApplicationsabstractThe Booth algorithm is widely used for efficient signed multiplication due to its ability to reduce partial products. A higher radix Booth multiplier generates fewer partial products, while it also increases hardware complexity in the generator, diminishing the advantage of fewer accumulators. Previous optimizations of generators and accumulators were designed for application-specific integrated circuits (ASICs), but their performance gains cannot be comparably translated to field-programmable gate arrays (FPGAs) due to differences in architecture. This article proposes FPGA-friendly approximate Booth multipliers that combine approximate hybrid-radix partial product generation with resource-efficient accumulation techniques. Initially, to improve generation efficiency, an look-up table (LUT)-reused exact radix-8 generator is introduced through logical partitioning to integrate two types of partial products into a single LUT. In addition, approximate adjacent-compensation radix-8 and radix-16 generators are developed based on the Booth encoding bit-repetition principle. Later, to speed up partial product accumulation, an overlap-parallel accumulation scheme and various accumulators are proposed, reducing compression steps and enhancing resource utilization. Last, performance-configurable hybrid radix-8/-16 approximate Booth multipliers are designed to meet the needs of different error-resilient applications. The most hardware-efficient configuration of the proposed 16-bit multiplier reduces power–delay product (PDP) and LUT consumption by 38.31% and 35.66%, respectively, compared with the exact multiplier. Furthermore, the proposed designs offer a better balance between accuracy and hardware complexity than existing approximate multipliers. The practicality of these multipliers is demonstrated in both joint photographic experts group (JPEG) image compression and finite impulse response (FIR) filtering applications. An open-source library of the proposed multipliers is available athttps://github.com/YnuGuoLab/FPGA_Signed_Approx_Multo support further research. Xuetao Li, Heming Sun, Haroon Waris, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible AcceleratorabstractLarge Language Models (LLMs) excel in natural language processing tasks but pose significant computational and memory challenges for edge deployment due to their intensive resource demands. This work addresses the efficiency of LLM inference by algorithm-hardwaredataflow tri-optimizations. We propose a novel voting-based KV cache eviction algorithm, balancing hardware efficiency and algorithm accuracy by adaptively identifying unimportant kv vectors. From a dataflow perspective, we introduce a flexible-product dataflow and a runtime reconfigurable PE array for matrix-vector multiplication. The proposed approach effectively handles the diverse dimensional requirements and solves the challenges of incrementally varying sequence lengths. Additionally, an element-serial scheduling scheme is proposed for nonlinear operations, such as softmax and layer normalization (layernorm). Results demonstrate a substantial reduction in latency, accompanied by a significant decrease in hardware complexity, from $O(N)$ to $O(1)$. The proposed solution is realized in a custom-designed accelerator, VEDA, which outperforms existing hardware platforms. This research represents a significant advancement in LLM inference on resource-constrained edge devices, facilitating real-time processing, enhancing data privacy, and enabling model customization. Zhican Wang, Hongxiang Fan, Haroon Waris, Gang Wang 0063, Jianfei Jiang 0001, Yanan Sun 0003, Guanghui He 0002 |
DAC | 3 |
| 2024 | Design Co-Processor Based on Partially Homomorphic Encryption Execution Using Open-Source ToolabstractExploiting a weakness in OpenSSL's TLS protocol implementation, the Heartbleed vulnerability was found in April 2014 and exposed sensitive data. This study describes a co-processor architecture that reduces the danger of data leaking by applying partly homomorphic encryption. Using secure multiplexers and non-traditional arithmetic units as challenges, our fully functioning device is created in 130nm CMOS technology. Through UART, it exchanges data with the primary CPU. This paper describes an ASIC development process utilizing open-source methodologies, from RTL design to GDS-II. Mujahid Bilal, M. Kamran Bhatti, Muhammad Kahsif Minhas, Haroon Waris |
VLSI-SoC | 4 |
| 2024 | FVDCLS: Functional Verification of RISCV Based Dual-Core Lockstep Feature Using Fault Injection MechanismabstractThis paper presents the framework to verify RISCV based Dual-Core Lockstep (DCLS), a feature that enhances the fault tolerance in processors. The DCLS feature, when two identical processors run in parallel, effectively detects the error and rolls back the core operations to a previously known state. The purpose-built fault injection mechanism is developed using universal verification methodology (UVM) and contains two modules: i) The fault injector (FI), ii) The virtual comparator (VC). The FI systematically injects the error in one of the core's internal states by flipping one or multiple bits. While the VC stores the internal states of both the RISCV cores on each clock cycle and reports the anomaly, if detected. The SweRV EH1 (an open-source RISCV core) is customized to add the DCLS functionality and later Google RISC-V DV test suite is used to validate the proposed FVDCLS. In total, nine different verification scenarios are developed consisting of transient and permanent errors using FI module. For all these test scenarios the VC compared hundred different internal state signals/registers. The result shows that the FVDCLS successfully detected and reported all the mismatches. Moreover, the FVDCLS incorporates an error debugging module that generates the summary of conducted test particularly, the type of fault injected, the number of errors detected, and the injected fault coverage. The proposed FVDCLS is portable and configurable; therefore, can be seamlessly integrated with any processor (with DCLS feature) under test. Muhammad Kashif Minhas, Haroon Waris, Sajid Baloch |
VLSI-SoC | 2 |
| 2024 | FPAX: A Fast Prior Knowledge-Based Framework for DSE in Approximate ConfigurationsabstractCurrent artificial intelligence and data science applications typically require complex computations and massive amounts of data handling, presenting unprecedented challenges for embedded platforms. Approximate computing has emerged as the most promising design technique to address this issue, by providing a potential performance increase, while sacrificing accuracy within an acceptable range. Approximate arithmetic units require the creation of design space exploration techniques that can swiftly and automatically form an approximate configuration in fault-tolerant systems. Existing methods, however, use iterative design space sampling, resulting in a large amount of redundant computation. In this work, we propose the efficient FPAX automatic search framework which can learn from prior knowledge regarding the exploration process of known applications and use it to guide design exploration. This avoids excessive redundant computation and quickly provides an impressive approximate configuration. Compared with the Jump Search algorithm known for its efficiency, FPAX can also achieve faster convergence speed and better exploration quality. Even compared to our previous ENAP framework, it exhibits an 18x faster performance while achieving almost identical exploration quality for several commonly used fault-tolerant applications. Yuqin Dou, Chenghua Wang, Haroon Waris, Roger F. Woods, Weiqiang Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Exact and Approximate Squarers for Error-Tolerant ApplicationsabstractApproximate computing is considered an innovative paradigm with wide applications to high performance and low power systems. These applications have relaxed requirements for accuracy, so they can tolerate errors in results and achieve high performance. In approximate computing, multipliers have been widely studied, but squarers (as similar schemes) have not received much attention. In this paper, an accurate squarer is designed based on a Radix-8 Booth-folding square algorithm to reduce the number of partial products and the depth of the partial product array. Several approximate squarers (R8AS1, R8AS2 and R8AS3) are proposed based on the exact squarer to reduce power and delay. Two approximate partial product generators are also designed to simplify the Radix-8 Booth square encoder in R8AS1 and R8AS2. In addition, approximate compressors with compensation are used in the partial product compression stage to reduce additional area and power consumption in R8AS3. Synthesis results for power, area, and delay at 28 nm CMOS technology are presented. Compared with designs in the technical literature with the same accuracy, the proposed 16-bit designs reduce the PDP by 37%; in general, the PDP is decreased by up to 51%. Finally, the proposed approximate squarers are implemented in a square-law detector as a communication application and achieve an SNR close to 30 dB. Also, the three proposed approximate squarers are applied to the k-means clustering algorithm for machine learning to accomplish high performance in classification. Ke Chen 0018, Chenyu Xu, Haroon Waris, Weiqiang Liu 0001, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2023 | Approximate Softmax Functions for Energy-Efficient Deep Neural NetworksabstractApproximate computing has emerged as a new paradigm that provides power-efficient and high-performance arithmetic designs by relaxing the stringent requirement of accuracy. Nonlinear functions (such as softmax, rectified linear unit (ReLU), Tanh, and Sigmoid) are extensively used in deep neural networks (DNNs). However, they incur significant power dissipation due to the high circuit complexity. As DNNs are error-tolerant, the design of approximation-linear functions is possible and desired. In this article, the design of an approximate softmax function (AxSF) is proposed. AxSF is based on a double hybrid structure (DHS). AxSF divides the input of the softmax function into two parts for different processing methods. The most significant bits (MSBs) are processed with lookup tables (LUTs) and an exact restoring array divider (EXDr). Taylor’s expansion and a logarithmic divider are used for the less significant bits (LSBs). An improved DHS (IDHS) is also proposed to reduce the hardware complexity. In IDHS, a novel Booth multiplier is utilized for the hybrid scheme to improve the partial product generation and compression, while the truncated implementation is applied to the divider unit. The proposed DHS and IDHS are compared with existing softmax designs. The results show that the proposed approximate softmax design reduces hardware by 48% and delay by 54% while retaining a high accuracy. Ke Chen 0018, Haroon Waris, Weiqiang Liu 0001, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | Design and Analysis of Energy-Efficient Dynamic Range Approximate Logarithmic Multipliers for Machine LearningabstractApproximate computing provides an emerging approach to design high performance and low power arithmetic circuits. The logarithmic multiplier (LM) converts multiplication into addition and has inherent approximate characteristics. In this article, dynamic range approximate LMs (DR-ALMs) for machine learning applications are proposed; they use Mitchell’s approximation and a dynamic range operand truncation scheme. The worst case (absolute and relative) errors for the proposed DR-ALMs are analyzed. The accuracy and the hardware overhead of these designs are provided to select the best approximate scheme according to different metrics. The proposed DR-ALMs are compared with the conventional LM with exact operands and previous approximate multipliers; the results show that the power-delay product (PDP) of the best proposed DR-ALM (DR-ALM-6) are decreased by up to 54.07 percent with the mean relative error distance (MRED) decreasing by 21.30 percent compared with 16-bit conventional design. Case studies for three machine learning applications show the viability of the proposed DR-ALMs. Compared with the exact multiplier and its conventional counterpart, the back-propagation classifier with DR-ALMs with a truncation length larger than 4 has a similar classification result for the three datasets; the K-means clustering application with all DR-ALMs has a similar clustering result for four datasets; and the handwritten digit recognition application with DR-ALM-5 or DR-ALM-6 for LeNet-5 achieves similar or even slightly higher recognition rate. Peipei Yin, Chenghua Wang, Haroon Waris, Weiqiang Liu 0001, Yinhe Han 0001, Fabrizio Lombardi |
IEEE Trans. Sustain. Comput. | 3 |