VLDB 2026 Research / reviewers in the wild / expert
Yousra Al-Kabani
dblp:223/9933 · also Yousra Alkabani
· DBLP profile ↗
15ranked-venue papers
8as first author
2since 2021 · last 2025
0000-0001-8806-8146ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 7 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Electronic design automation · 35% Energy-efficient computing · 20% Embedded and real-time systems · 13% | |
| Network and information security
2 papers |
Hardware security and side channels · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware security and side channels
intellectual property protection |
0.2 | 2 | 2008 | N-variant IC design: methodology and applications · DAC 2008 Active Hardware Metering for Intellectual Property Protection and Security · USENIX Security Symposium 2007 |
Electronic design automation › hardware verification and test › functional verification
emulation |
0.1 | 1 | 2010 | Real time emulations: foundation and applications · DAC 2010 |
Electronic design automation › hardware verification and test
hardware verification |
0.1 | 1 | 2010 | Real time emulations: foundation and applications · DAC 2010 |
Embedded and real-time systems
real-time emulation |
0.1 | 1 | 2010 | Real time emulations: foundation and applications · DAC 2010 |
Energy-efficient computing › leakage power reduction
input vector control |
0.1 | 1 | 2008 | Input vector control for post-silicon leakage current minimization in the presence of manufacturing variability · DAC 2008 |
Energy-efficient computing
leakage power reduction |
0.1 | 1 | 2008 | Input vector control for post-silicon leakage current minimization in the presence of manufacturing variability · DAC 2008 |
Hardware reliability and fault tolerance
process variation |
0.1 | 1 | 2008 | Input vector control for post-silicon leakage current minimization in the presence of manufacturing variability · DAC 2008 |
Integrated circuit design › digital circuit design
sequential circuit design |
0.1 | 1 | 2008 | N-variant IC design: methodology and applications · DAC 2008 |
Electronic design automation
intellectual property protection |
0.1 | 1 | 2007 | Active Hardware Metering for Intellectual Property Protection and Security · USENIX Security Symposium 2007 |
High-performance computing › supercomputing
petascale computing |
0.0 | 1 | 2010 | Real time emulations: foundation and applications · DAC 2010 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2010 | Real time emulations: foundation and applications · DAC 2010 |
Distributed systems
fault tolerance |
0.0 | 1 | 2008 | N-variant IC design: methodology and applications · DAC 2008 |
Methods — techniques the papers use, named apart from their topics
finite state machine extension · 0.2physical prototyping · 0.12d/3d silicon emulation · 0.1input vector control · 0.1gate-level characterization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Software acceleration of multi-user MIMO uplink detection on GPUabstractThis paper presents the exploration of GPU-accelerated block-wise decompositions for zero-forcing (ZF) based QR and Cholesky methods applied to massive multiple-input multiple-output (MIMO) uplink detection algorithms. Three algorithms are evaluated: ZF with block Cholesky decomposition, ZF with block QR decomposition (QRD), and minimum mean square error (MMSE) with block Cholesky decomposition. The latter was the only one previously explored, but it used standard Cholesky decomposition. Our approach achieves an 11% improvement over the previous GPU-accelerated MMSE study. Through performance analysis, we observe a trade-off between precision and execution time. Reducing precision from FP64 to FP32 improves execution time but increases bit error rate (BER), with ZF-based QRD reducing execution time from 2 . 04 μ s to 1 . 24 μ s for a 128 × 8 MIMO size. The study also highlights that larger MIMO sizes, particularly 2048 × 32, require GPUs to fully utilize their computational and memory capabilities, especially under FP64 precision. In contrast, smaller matrices are compute-bound. Our results recommend GPUs for larger MIMO sizes, as they offer the parallelism and memory resources necessary to efficiently handle the computational demands of next-generation networks. This work paves the way for scalable, GPU-based massive MIMO uplink detection systems. Ali Nada, Hazem Ismail Abdel Aziz Ali, Liang Liu 0002, Yousra Al-Kabani |
Parallel Comput. | 4 |
| 2022 | A Deep Neural Network Accelerator using Residue Arithmetic in a Hybrid Optoelectronic SystemabstractThe acceleration of Deep Neural Networks (DNNs) has attracted much attention in research. Many critical real-time applications benefit from DNN accelerators but are limited by their compute-intensive nature. This work introduces an accelerator for Convolutional Neural Network (CNN) , based on a hybrid optoelectronic computing architecture and residue number system (RNS) . The RNS reduces the optical critical path and lowers the power requirements. In addition, the wavelength division multiplexing (WDM) allows high-speed operation at the system level by enabling high-level parallelism. The proposed RNS compute modules use one-hot encoding, and thus enable fast switching between the electrical and optical domains. We propose a new architecture that combines residue electrical adders and optical multipliers as the matrix-vector multiplication unit. Moreover, we enhance the implementation of different CNN computational kernels using WDM-enabled RNS based integrated photonics. The area and power efficiency of the proposed accelerator are 0.39 TOPS/s/mm 2 and 3.22 TOPS/s/W, respectively. In terms of computation capability, the proposed chip is 12.7× and 4.02× better than other optical implementation and memristor implementation, respectively. Our experimental evaluation using DNN benchmarks illustrates that our architecture can perform on average more than 72 times faster than GPU under the same power budget. Yousra Al-Kabani, Krunal Puri, Volker J. Sorger, Tarek A. El-Ghazawi |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | DNNARA: A Deep Neural Network Accelerator using Residue Arithmetic and Integrated PhotonicsabstractDeep Neural Networks (DNNs) are currently used in many fields, including critical real-time applications. Due to its compute-intensive nature, speeding up DNNs has become an important topic in current research. We propose a hybrid opto-electronic computing architecture targeting the acceleration of DNNs based on the residue number system (RNS). In this novel architecture, we combine the use of Wavelength Division Multiplexing (WDM) and RNS for efficient execution. WDM is used to enable a high level of parallelism while reducing the number of optical components needed to decrease the area of the accelerator. Moreover, RNS is used to generate optical components with short optical critical paths. In addition to speed, this has the advantage of lowering the optical losses and reducing the need for high laser power. Our RNS compute modules use one-hot encoding and thus enable fast switching between the electrical and optical domains. Yousra Al-Kabani, Volker J. Sorger, Tarek A. El-Ghazawi |
ICPP | 2 |
| 2019 | Photonic Processor for Fully Discretized Neural NetworksabstractMachine learning is now moving towards, and will become prevalent in, fog-computing and real-time computing environments. To this end, much machine-learning-at-the-edge research has focused on efficient neural network architectures, giving rise to efficient approximations of fixed-point neural networks, called discretized neural networks. While higher performing than their fixed and floating-point counterparts, discretized neural networks still have an existing bottleneck at the neuron's accumulation of products, called the popcount. This bottleneck sets an upper bound on performance regardless of neural network architecture. We address the popcount bottleneck by introducing a photonic discretized neural network processor. This processor minimizes the popcount bottleneck, thereby maximizing neural network computational throughput. Additionally, it offers potential for performance enhancement through simultaneous convolution operations enabled by wavelength division multiplexing. We show that the photonic architecture is capable of increasing performance by 700% and 100% when compared to state-of-the-art digital and analog architectures, respectively. Jeff Anderson, Yousra Al-Kabani, Volker J. Sorger, Tarek A. El-Ghazawi |
ASAP | 3 |
| 2015 | Finite element emulation-based solver for electromagnetic computationsabstractElectromagnetic (EM) computations are the cornerstone in the design process of several real-world applications, such as radar systems, satellites, and cell-phones. Unfortunately, these computations are mainly based on numerical techniques that require solving millions of linear equations simultaneously. Software-based solvers do not scale well as the number of equations-to-solve increases. FPGA solver implementations were used to speed up the process. However, using emulation technology is more appealing as emulators overcome the FPGA memory and area constraints. In this paper, we present a scalable design to accelerate the finite element solver of an EM simulator on a hardware emulation platform. Experimental results show that our optimized solver achieves 101.05x speed-up over the same pure software implementation on MATLAB and 35.29x over the best iterative software solver from ALGLIB C++ package in case of solving 2,002,000 equations. M. Tarek Ibn Ziad, Mohamed Hossam, Mohamad A. Masoud, Mohamed Nagy, Hesham A. Adel, Yousra Al-Kabani, M. Watheq El-Kharashi, Khaled Salah 0001, Mohamed Abdel Salam |
ISCAS | 6 |
| 2013 | Hardware Trojan Protection for Third Party IPsabstractHardware Trojan detection is a very important topic especially as parts of critical systems which are designed and/or manufactured by untrusted third parties. Most of the current research concentrates on detecting Trojans at the testing phase by comparing the suspected circuit to a golden (trusted) one. However, these attempts do not work in the case of third party IPs, which are black boxes with no golden IPs to trust. In this work, we present novel methods for system protection that alleviate the need for a golden chip. Protection against injected Trojan is done using simple blockage method. We show the practicality of the introduced schemes by providing a proof of concept implementation of the proposed methodology on FPGA. We showed that the overhead is low enough in the simple blockage method. The delay overhead is negligible while the power overhead does not exceed 2%. Amr Al-Anwar 0001, Yousra Al-Kabani, M. Watheq El-Kharashi, Hassan Bedour |
DSD | 2 |
| 2012 | Trojan Immune Circuits Using DualityabstractThe problem of hardware Trojan detection has been recently studied extensively. The use of traditional testing strategies to detect hardware Trojans is not effective because the probability of triggering a hardware Trojan during testing is very low. Moreover, the small size of the Trojan compared to the overall size of the chip reduces the impact of the Trojan on side channels such as static and dynamic power. Process variations in modern technologies will even distort this impact and reduce the efficiency of such methods to detect the Trojan. In this work, we propose the development of new Trojan immune circuits where Trojans are easier to detect using traditional testing methods. The main idea is that each circuit has a designed dual. By testing the dual, one can easily detect a Trojan embedded in the original circuit. On the other hand, if an attacker tries to hide the Trojan in the dual, the Trojan will be detectable in the original circuit. We study an example dual design using gates with dual function. We also present attacks and countermeasures on the dual to provide guidelines for the low level implementation of the proposed architecture. Experimental results show that a Trojan hidden in a circuit can be detected by applying a few random test inputs on the dual. In addition, the estimated average area overhead to construct the example dual is 5.5%. Yousra Al-Kabani |
DSD | 1 |
| 2010 | Real time emulations: foundation and applicationsabstractThe mesoscopic properties of the state-of-the-art nanoscale devices and the emerging petascale computing and storage systems have one thing in common: they function at scales that are orders of magnitude larger than what can be simulated in standard industry and academic laboratory settings. For many decades, CAD and verification communities have successfully developed and used emulations to overcome and complement the shortcomings of simulations for logic verification. Physical prototyping and 2D/3D silicon emulation of the increasingly complex systems holds a significant promise to overcome the limitations of computer modeling and simulations. While the potential opportunities are plenty, much research is required for prototyping and building effective, relevant and indicative emulation platforms. Azalia Mirhoseini, Yousra Al-Kabani, Farinaz Koushanfar |
DAC | 2 |
| 2009 | Consistency-based characterization for IC Trojan detectionabstractA Trojan attack maliciously modifies, alters, or embeds unplanned components inside the exploited chips. Given the original chip specifications, and process and simulation models, the goal of Trojan detection is to identify the malicious components. This paper introduces a new Trojan detection method based on nonintrusive external IC quiescent current measurements. We define a new metric called consistency. Based on the consistency metric and properties of the objective function, we present a robust estimation method that estimates the gate properties while simultaneously detecting the Trojans. Experimental evaluations on standard benchmark designs show the validity of the metric, and demonstrate the effectiveness of the new Trojan detection. Yousra Al-Kabani, Farinaz Koushanfar |
ICCAD | 1 |
| 2009 | N-version temperature-aware scheduling and bindingabstractTechnology scaling to nanometer nodes causes growing increase in power density and especially leakage that in turn result in locally hot regions on the chip. In this paper, we introduce a novel methodology for temperature-aware design. The methodology embeds N-versions of the scheduler and binder such that the thermal profiles of the versions are distant from each other. Next, instead of using only one version of the scheduler and binder, a rotation of N-versions of the scheduler and binder is constructed for balancing the thermal profile of the chip. We propose a linear programming framework that takes the multiple versions as the input, and constructs the thermal-aware rotational scheduling and binding by selecting the N most efficient versions and by determining the duration of each version. Our experimental evaluation shows a very low overhead and an average 5% decrease in the steady-state peak temperature produced on the benchmark designs compared to using a schedule that balances the amount of usage of different modules. Yousra Al-Kabani, Farinaz Koushanfar, Miodrag Potkonjak |
ISLPED | 1 |
| 2008 | Active control and digital rights management of integrated circuit IP coresabstractWe introduce the first approach that can actively control multiple hardware intellectual property (IP) cores used in an integrated circuit (IC). The IP rights owner(s) can remotely monitor, control, enable, or disable each individual IP on each chip. The approach introduces a paradigm shift in the microelectronic business model, nurturing smaller businesses, and supporting the design-reuse paradigm. The IPs can be controlled by the original designer or by the designers who reuse them. Each IP has a built-in functional lock that pertains to the unique unclonable ID of the chip. A control structure that coordinates the locking and unlocking of the IPs is embedded within the IC. We introduce a trusted third party approach for issuing certificates of authenticity, in case it is required for the applications. We present methods for safeguarding the approach against two attack sources: the foundry (fab), and the reuser. Experimental results show that our approach can be implemented with low area, power, and delay overheads making it suitable for embedded systems. The introduced control method is also low overhead in terms of the added steps to the current design and manufacturing flow. Yousra Al-Kabani, Farinaz Koushanfar |
CASES | 1 |
| 2008 | N-variant IC design: methodology and applicationsabstractWe propose the first method for designing N-variant sequential circuits. The flexibility provided by the N-variants enables a number of important tasks, including IP protection, IP metering, security, design optimization, self-adaptation and fault-tolerance. The method is based on extending the finite state machine (FSM) of the design to include multiple variants of the same design specification. The state transitions are managed by added signals that may come from various triggers depending on the target application. We devise an algorithm for implementing the N-variant IC design. We discuss the necessary manipulations of the added signals that would facilitate the various tasks. The key advantage to integrating the heterogeneity in the functional specification of the design is that we can configure the variants during or post-manufacturing, but removal, extraction or deletion of the variants is not viable. Experimental results on benchmark circuits demonstrate that the method can be automatically and efficiently implemented. Because of its lightweight, N-variant design is particularly well-suited for securing embedded systems. As a proof-of-concept, we implement the N-variant method for content protection in portable media players, e.g., iPod. We discuss how the N-variant design methodology readily enables new digital rights management methods. Yousra Al-Kabani, Farinaz Koushanfar |
DAC | 1 |
| 2008 | Input vector control for post-silicon leakage current minimization in the presence of manufacturing variabilityabstractWe present the first approach for post-silicon leakage power reduction through input vector control (IVC) that takes into account the impact of the manufacturing variability (MV). Because of the MV, the integrated circuits (ICs) implementing one design require different input vectors to achieve their lowest leakage states. We address two major challenges. The first is the extraction of the gatelevel characteristics of an IC by measuring only the overall leakage power for different inputs. The second problem is the rapid generation of input vectors that result in a low leakage for a large number of unique ICs that implement a given design, but are different in the post-manufacturing phase. Experimental results on a large set of benchmark instances demonstrate the efficiency of the proposed methods. For example, the leakage power consumption could be reduced in average by more than 10.4%, when compared to the previously published IVC techniques that did not consider MV. Yousra Al-Kabani, Tammara Massey, Farinaz Koushanfar, Miodrag Potkonjak |
DAC | 1 |
| 2007 | Remote activation of ICs for piracy prevention and digital right managementabstractWe introduce a remote activation scheme that aims to protect integrated circuits (IC) intellectual property (IP) against piracy. Remote activation enables designers to lock each working IC and to then remotely enable it. The new method exploits inherent unclonable variability in modern manufacturing for unique identification (ID) and integrate the IDs into the circuit functionality. The objectives are realized by replication of a few states of the finite state machine (FSM) and adding control to the state transitions. On each chip, the added control signals are a function of the unique IDs and are thus unclonable. On standard benchmark circuits, the experimental results show that the novel activation method is stable, unclonable, attack-resilient, while having a low overhead and a unique key for each IC. Yousra Al-Kabani, Farinaz Koushanfar, Miodrag Potkonjak |
ICCAD | 1 |
| 2007 | Active Hardware Metering for Intellectual Property Protection and Security
Yousra Al-Kabani, Farinaz Koushanfar |
USENIX Security Symposium | 1 |