Dhamin Al-Khalili

dblp:17/5168 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0002-0079-8417ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Electronic design automation · 51% Integrated circuit design · 28% Performance modeling and evaluation · 21%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.132005
Delay analysis of CMOS gates using modified logical effort model · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Technology-portable analytical model for DSM CMOS inverter transition-time estimation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
A module generator for optimized CMOS buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1990
Performance modeling and evaluation › analytical modeling
analytical delay modeling
0.122005
Delay analysis of CMOS gates using modified logical effort model · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Technology-portable analytical model for DSM CMOS inverter transition-time estimation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Electronic design automation
physical design
0.122005
Simultaneous adaptive wire adjustment and local topology modification for tuning a bounded-skew clock tree · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
A Module Generator for Optimized CMOS Buffers · DAC 1989
Electronic design automation › physical design › clock network synthesis
clock skew optimization
0.112005
Simultaneous adaptive wire adjustment and local topology modification for tuning a bounded-skew clock tree · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › physical design › clock network synthesis
clock tree synthesis
0.112005
Simultaneous adaptive wire adjustment and local topology modification for tuning a bounded-skew clock tree · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Integrated circuit design › digital circuit design › logic gate design
CMOS gate
0.112005
Delay analysis of CMOS gates using modified logical effort model · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Performance modeling and evaluation
delay analysis
0.112005
Delay analysis of CMOS gates using modified logical effort model · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Integrated circuit design › digital circuit design › logic gate design
CMOS inverter
0.012003
Technology-portable analytical model for DSM CMOS inverter transition-time estimation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Electronic design automation › hardware verification and test
fault modeling
0.012003
IC Bridge Fault Modeling for IP Blocks Using Neural Network-Based VHDL Saboteurs · IEEE Trans. Computers 2003
Electronic design automation › hardware verification and test
fault simulation
0.012003
IC Bridge Fault Modeling for IP Blocks Using Neural Network-Based VHDL Saboteurs · IEEE Trans. Computers 2003
Electronic design automation
hardware verification and test
0.012003
IC Bridge Fault Modeling for IP Blocks Using Neural Network-Based VHDL Saboteurs · IEEE Trans. Computers 2003
Electronic design automation
timing analysis
0.022005
Delay analysis of CMOS gates using modified logical effort model · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Technology-portable analytical model for DSM CMOS inverter transition-time estimation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Electronic design automation › physical design › layout optimization
wire length minimization
0.012005
Simultaneous adaptive wire adjustment and local topology modification for tuning a bounded-skew clock tree · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › physical design
module generation
0.021990
A module generator for optimized CMOS buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1990
A Module Generator for Optimized CMOS Buffers · DAC 1989
Electronic design automation › physical design
buffer optimization
0.011990
A module generator for optimized CMOS buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1990
Electronic design automation › physical design › buffer optimization
buffer sizing
0.011990
A module generator for optimized CMOS buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1990

Methods — techniques the papers use, named apart from their topics

SPICE simulation · 0.1modified logical effort · 0.1local topology modification · 0.1deferred-merge embedding · 0.1adaptive wire adjustment · 0.1Greedy-DME · 0.1neural network · 0.0analytical modeling · 0.0VHDL saboteur · 0.0process and layout parameter variation study · 0.0
YearPublicationVenuePosition
2023 FinFET 6T-SRAM Compute-in-Memory Targeting Low Power Neural Networks Operations
abstract
Modern computing relies on the artificial intelligence (AI) for making independent decisions. AI utilizes neural networks (NN) to handle complex tasks execution such as image recognition, text processing, and language interpretation. However, NNs are extremely data centric and pose performance bottleneck while moving data between memory and processing unit. Compute-in-memory (CIM) addresses this issue but limited attention has been given to the low power neural network operations. This paper proposes a novel CIM solution for multiply and accumulate (MAC) operations without incorporating any dedicated arithmetic units. Proposed CIM makes use of FinFET 6T-SRAM cells, sense amplifiers, write drivers, and NOR gates to achieve the power efficient in-memory computations. Simulations in 12nm FinFET shows that proposed CIM for 2-bit input and 2-bit weights: consumes$411.2\mu \mathrm{W}$, exhibits latency of 36.1 ns, and shows maximum power efficiency of 6.51 KOP/W.
Waqas Gul, Maitham Shams, Dhamin Al-Khalili
ISCAS3
2018 Gate Oxide Short Defect Model in FinFETs
Roya Dibaj, Dhamin Al-Khalili, Maitham Shams
J. Electron. Test.2
2017 Comprehensive investigation of gate oxide short in FinFETs
abstract
Manufacturing complexities due to FinFET's three-dimensional structure and reduced critical dimensions have caused new challenges in achieving reliable device testing. Gate oxide short (GOS) is one of the defects that requires a thorough investigation due to its complexity in 3D transistors and its significant impact on circuit reliability. In this paper, we present a comprehensive study on the transistor defect characteristics as we introduce pinholes in the gate oxide of rectangular fin and trapezoidal fin shape structures. The pinholes are represented by small cuboid cuts of various sizes located along the fin height and channel length. Our analysis is performed with the aid of Synopsys' Sentaurus TCAD tools. The results presented in this paper can lead to the development of more realistic analytical GOS defect model for circuit level simulation.
Roya Dibaj, Dhamin Al-Khalili, Maitham Shams
VTS2
2008 Efficient FPGA implementation of complex multipliers using the logarithmic number system
abstract
In many real-time DSP applications, high performance is a prime target. However, achieving this may be done at the expense of area, power dissipation and accuracy. Attempts have been made to use alternative number systems to optimize the realization of arithmetic blocks, maintaining high performance without incurring prohibitive area and power increases. This paper presents the FPGA implementation of complex multipliers based on the logarithmic number system. Synthesis results show that a design with a 10-stage pipeline can achieve a maximum clock rate of 224 MHz and 140 MHz for 16-bit and 32-bit designs, respectively. Both designs use the lowest amount of hardware in terms of gate equivalents as compared to a complex multiplier built with regular FPGA features. In particular, the proposed architecture uses 67% and 35% fewer gates to implement a 32-bit and 16-bit complex multiplier, respectively, when compared to a design realized with embedded multipliers. Simulation results based on selected test vectors show that the greatest relative error of the logarithmic-based 16-bit complex multiplier is 2.14%
Man Yan Kong, J. M. Pierre Langlois, Dhamin Al-Khalili
ISCAS3
2007 FPGA-Based Efficient Design Approach for Large-Size Two's Complement Squarers
abstract
This paper presents an optimized design approach of two's complement large-size squarers using embedded multipliers in FPGAs. The realization is based on Baugh-Wooley's algorithm, which partitions the multiplication into unsigned and signed sections. To achieve efficient implementation, a set of optimized schemes for the realization of multi-level additions of the partial products is proposed. Our approach has been evaluated through the implementation of squarers for operands with sizes ranging from 20 to 128 bits. The designs are synthesized and implemented on Xilinx' Spartan-3 with ISE 8.1 design platform and compared with the standard implementation, and with Xilinx' IP Core. The results indicate that our approach offers substantial LUT savings by up to 52% with an average delay reduction of 13%. The usage of the number of embedded multipliers is reduced by 38% compared with the standard schemes.
Shuli Gao, Noureddine Chabini, Dhamin Al-Khalili, J. M. Pierre Langlois
ASAP3
2006 Accurate Total Static Leakage Current Estimation in Transistor Stacks
abstract
In this paper, a simple model for the estimation of static leakage current in NMOS transistor stacks is introduced. The three leakage mechanisms addressed are subthreshold leakage, gate-tunneling and gate induced drain leakage (GIDL). The algorithmic description of the model can be broken down into three phases i) pre-extraction , ii) estimation and iii) width scaling. In the pre-extraction phase, data necessary for subthreshold leakage estimation is extracted a priori. This also involves characterizing voltages required for a specific set of input vector scenarios (exception vectors/voltages). In the estimation phase, unit width GIDL and gate tunneling are estimated deterministically while subthreshold leakage is estimated using the pre-extracted data. Finally, in the width scaling phase each leakage component is then width scaled and summed, to give the total static leakage exhibited by the stack. The proposed model was scripted in MatLab and compared with SPICE simulations for various scenarios. The average total error for each scenario was under 3%.
Hussam Al-Hertani, Dhamin Al-Khalili, Côme Rozon
AICCSA2
2006 An Optimized Design Approach for Squaring Large Integers Using Embedded Hardwired Multipliers
abstract
This paper presents an efficient design methodology and a systematic approach for the implementation of squaring functions with large integers, using small-size embedded multipliers. A general architecture of the squarer and a set of equations are derived to aid in the realization. The inputs of the squarer are split into several segments leading to an efficient utilization of the small-size embedded multipliers and reduced number of required addition operations. Various benchmarks were tested for different segments ranging from 2 to 5 targeting Xilinx Spartan-3 FPGA. The synthesis was performed with the aid of the Xilinx ISE 7.1 XST tool. Our approach was compared with the traditional technique using the same tool. The results illustrate that our design approach is very efficient in terms of both timing and area saving. The combinational delay is reduced by an average of 15.8%, and the area saving is about 50 % in terms of number of slices and number of 4-input LUTs. Also, the number of required embedded multipliers is reduced by an average of 32.3% compared to the traditional technique.
Shuli Gao, Noureddine Chabini, Dhamin Al-Khalili, J. M. Pierre Langlois
AICCSA3
2006 Automatic generation of defect injectable VHDL fault models for ASIC standard cell libraries
Donald B. Shaw, Dhamin Al-Khalili, Côme Rozon
Integr.2
2005 Delay analysis of CMOS gates using modified logical effort model
abstract
In this paper, modified logical effort (MLE) technique is proposed to provide delay estimation for CMOS gates. The model accounts for the behavior of series-connected MOSFET structure (SCMS), the input transition time, and internodal charges. Also, the model takes into account deep submicron effects, such as mobility degradation and velocity saturation. This model exhibits good accuracy when compared with Spectre simulations based on BSIM3v3 model. Using UMC's 0.13-/spl mu/m and TSMC's 0.18-/spl mu/m technologies, the model has an average error of 4.5% and a maximum error of 15%.
Adnan Kabbani, Dhamin Al-Khalili, Asim J. Al-Khalili
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 Simultaneous adaptive wire adjustment and local topology modification for tuning a bounded-skew clock tree
abstract
The need for incremental algorithms to implement engineering changes (ECs) in clock trees (CTs) is critical in the system-on-a-chip (SoC) design cycle. An algorithm, called adaptive wire adjustment (AWA), is proposed to minimize the clock skew iteratively to any given bound. In order to speed up AWA's convergence, a local topology-modification (LTM) technique is incorporated into AWA. Moreover, LTM incorporation into AWA results in total wire-length reduction as well. Also, the incorporation of the LTM technique into the deferred-merge embedding (DME) algorithm and Greedy-DME (GDME) helps reduce the total wire length by around 7.8% and 9.8%, respectively. Additionally, applying LTM to GDME reduces wire elongations and the standard deviation of the path lengths (SDPL) between clock pins by 96.4% and 51.5%, respectively.
Haydar Saaied, Dhamin Al-Khalili, Asim J. Al-Khalili, Mohamed Nekili
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Adaptive wire adjustment for bounded skew Clock Distribution Network
abstract
In this paper, we suggest an adaptive approach for the Clock Distribution Network (CDN) to cope with a modification in the VLSI system design. The CDN's wires are adjusted iteratively to reduce the skew that is resulting from a minor modification in the clock pins of a complex VLSI system. Such skew can be remedied by selecting a Balancing Node (BN) and adjust its edges so that the skew gets smaller. The required edge adjustments are determined using the Elmore delay model. The performance of the algorithm is investigated using different random sets of clock pins. Also, the algorithm is tested by altering some clock pins in a zero skew CDN. For small modifications in a large number of nodes in the CDN, our algorithm can achieve zero skew with less iterations than linear order algorithms.
Haydar Saaied, Dhamin Al-Khalili, Asim J. Al-Khalili, Mohamed Nekili
ASP-DAC2
2003 IC Bridge Fault Modeling for IP Blocks Using Neural Network-Based VHDL Saboteurs
abstract
This paper presents a new bridge fault model, suitable for IP blocks, that is based on a multiple layer feedforward neural network and implemented within the framework of a VHDL saboteur cell. Empirical evidence and experimental results show that it satisfies a prescribed set of bridge fault model criteria better than any existing approach. The new model computes bridged node voltages and propagation delay times with due attention to surrounding circuit elements. This is especially significant since, with the exception of full analog defect simulation, no other technique even attempts to model the delay effects of bridge defects. Yet, compared to these analog simulations, the new approach is several orders of magnitude faster and, for a 0.35u cell library, is able to compute bridged node voltages with an average error near 0.006 volts and propagation delay times with an average error near 14 ps. Furthermore, dealing with a concept that has not previously been considered in related research, the new model is validated with respect to deep-submicron technologies for limited gate-count circuit modules.
Donald B. Shaw, Dhamin Al-Khalili, Côme Rozon
IEEE Trans. Computers2
2003 Technology-portable analytical model for DSM CMOS inverter transition-time estimation
abstract
In this paper, we propose a new analytical model to estimate the transition time of CMOS inverters, taking into account the main effects of deep submicron (DSM) such as velocity saturation and mobility degradation. The relationship between the input and output transitions is discussed and captured by a closed-form expression. Also, this model has been formulated to depend only on SPICE parameters that are usually provided with the given technology. In other words, neither presimulated extracted parameters nor fitting parameters are required. Thus, the developed model is technology-portable. The proposed model is validated by comparing its results with Spectre simulation results. To ensure the model's robustness, a wide range of output loading and input transition times have been considered. Also, to ensure the model's portability and accuracy for DSM devices, 0.25, 0.18, and 0.13-/spl mu/m technologies have been used to conduct our comparison. Considering the mentioned technologies, the proposed model achieved less than 10% error when it is compared to Specter level 11 (BSIM3v3) simulation results.
Adnan Kabbani, Dhamin Al-Khalili, Asim J. Al-Khalili
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 A low power direct digital frequency synthesizer with 60 dBc spectral purity
abstract
We present a low-power sine-output Direct Digital Frequency Synthesizer (DDFS) realized in 0.18 μm CMOS that achieves 60 dBc spectral purity from DC to the Nyquist frequency. No ROM or multipliers are used, but an external DAC is required if an analog output is desired. Power consumption is 10 mW for a 100 MHz clock, which is significantly less than figures reported previously. System complexity is greatly reduced by using an efficient linear interpolation scheme to approximate a sinusoid function. This has resulted in silicon area utilization of 0.011 mm2. The design would be suitable as an IP core in a low power digital RF transceiver ASIC.
J. M. Pierre Langlois, Dhamin Al-Khalili
ACM Great Lakes Symposium on VLSI2
2002 Fault security analysis of CMOS VLSI circuits using defect-injectable VHDL models
Donald B. Shaw, Dhamin Al-Khalili, Côme Rozon
Integr.2
2001 Accurate CMOS Bridge Fault Modeling with Neural Network-Based VHDL Saboteurs
abstract
This paper presents a new bridge fault model that is based on a multiple layer feedforward neural network and implemented within the framework of a VHDL saboteur cell. Empirical evidence and experimental results show that it satisfies a prescribed set of bridge fault model criteria better than existing approaches. The new model computes exact bridged node voltages and propagation delay times with due attention to surrounding circuit elements. This is significant since, with the exception of full analog simulation, no other technique attempts to model the delay effects of bridge defects. Yet, compared to these analog simulations, the new approach is orders of magnitude faster and achieves reasonable accuracy; computing bridged node voltages with an average error near 0.006 volts and propagation delay times with an average error near 14 ps.
Donald B. Shaw, Dhamin Al-Khalili, Côme Rozon
ICCAD2
2000 Comprehensive defect analysis and testability of current-mode logic circuits
abstract
This paper presents comprehensive defect analysis of digital CML circuits using detailed defect models at the device level. The circuits are based on Nortel's BiCMOS technology. A defect model macro has been used to analyze fourteen physical defects for each bipolar transistor and where all defects are being qualified as hard or soft. Defects are activated individually within CML gates which are exhaustively simulated and the responses compared to the response of equivalent gold circuits. Extracted data allow us to obtain both defect and fault coverages, and to generate statistics to assess the effectiveness of a given testing approach. For the CML gates studied, logical testing alone is not adequate as the coverage remained between 33% and 79%. By using combined testing the coverage was increased to 100%.
Saman Adham, Dhamin Al-Khalili, Côme Rozon, Douglas Racz
ISCAS2
1999 Synthesis of low-power CMOS circuits using hybrid topologies
Michael Gallant, Dhamin Al-Khalili
Integr.2
1998 VHDL Modelling and Analysis of Fault Secure Systems
abstract
This paper presents an analysis process targeted for the verification of fault secure systems during their design phase. This process deals with a realistic set of micro-defects at the device level which are mapped into mutant and saboteur based VHDL fault models in the form of logical and/or performance degradation faults. Automatic defect injection and simulation are performed through a VHDL test bench. Extensive post processing analysis is performed to determine defect coverage, figure of merit for fault secureness, and MTTF.
Jason Coppens, Dhamin Al-Khalili, Côme Rozon
DATE2
1997 A Low Power Approach to Floating Point Adder Design
abstract
We present a new architecture of a low power floating point adder, that is fast and has low latency. The functional partitioning of the adder into three distinct, controlled data paths allows activity reduction. During any given operation cycle, only one of the data paths is active, during which time, the logic assertion status of the circuit nodes of the other data paths are held at their previous states. Critical path delay and latency are reduced by incorporating speculative rounding and pseudo leading zero anticipation logic as well as data path simplifications. The proposed scheme offers a 10/spl times/ reduction in power consumption in comparison to that of conventional high speed floating point adders that use leading zero anticipation logic, for IEEE single precision floating point data format. The reduction in power delay product is about 16/spl times/. The corresponding figures for double precision units are around 40/spl times/ and 66/spl times/ respectively.
R. V. K. Pillai, Dhamin Al-Khalili, Asim J. Al-Khalili
ICCD2
1997 Energy delay measures of barrel switch architectures for pre-alignment of floating point operands for addition
abstract
Article Free Access Share on Energy delay measures of barrel switch architectures for pre-alignment of floating point operands for addition Authors: R. V. K. Pillai Concordia University, Montreal, Canada Concordia University, Montreal, CanadaView Profile , D. Al-Khalili Concordia University, Montreal, Canada Concordia University, Montreal, CanadaView Profile , A. J. Al-Khalili View Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 235–238https://doi.org/10.1145/263272.263341Published:01 August 1997Publication History 6citation138DownloadsMetricsTotal Citations6Total Downloads138Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
R. V. K. Pillai, Dhamin Al-Khalili, Asim J. Al-Khalili
ISLPED2
1996 Energy delay analysis of partial product reduction methods for parallel multiplier implementation
abstract
This paper examines the energy delay implications of partial product reduction methods employed in parallel multiplier implementations.Radix 4 Modified Booth Algorithm (MBA) is currently the most popular choice for partial product reduction in parallel multipliers although 4:2 compressors can also produce equivalent results.Our energy delay analysis of these two schemes taking into account the architectural as well as circuit implementation issues suggests the superiority of the 4:2 compressor based partial product reduction technique as far as circuit delays, power consumption and architectural regularity are concerned.SPICE simulations of partial product generation using these schemes for an 8 bit multiplier suggest a worst case energy delay advantage of the order of 36% and 15% respectively for the 4:2 compressor based scheme in comparison with two different implementations of MBA.The corresponding figures for power reduction are of the order of 26% and 11% respectively.
R. V. K. Pillai, Dhamin Al-Khalili, Asim J. Al-Khalili
ISLPED2
1993 Fault Characterization and Testability Analysis of Emitter Coupled Logic and Comparison with CMOS & BiCMOS Circuits
Michael Ogbonna Esonu, Dhamin Al-Khalili, Côme Rozon
ISCAS2
1992 Testability analysis and fault modeling of BiCMOS circuits
Dhamin Al-Khalili, Côme Rozon, B. Stewart
J. Electron. Test.1
1990 A module generator for optimized CMOS buffers
abstract
The theory and implementation of a module generator for CMOS buffers are presented. The generator is written in the C language, and outputs optimal buffer designs in respect to a preselected objective function and layout. The user has the choice of minimizing delay, power, and area, or a combination of these, plus the choice of layout configuration. The research concentrates mainly on theoretical analysis, where variations of process, design, and layout parameters with respect to each objective function are studied in detail.>
Asim J. Al-Khalili, Dhamin Al-Khalili
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1989 A Module Generator for Optimized CMOS Buffers
abstract
A module generator for CMOS buffers have been written in C. The generator optimizes buffer design with respect to a user specified objective function both in terms of performance and layout.Speed, area, power consumption, power-delay, AT and AT" are selectively optimized before the layout is produced.Such layout is generated in various configurations depending on load size.Technology file is easily updatable.
Asim J. Al-Khalili, Dhamin Al-Khalili
DAC3
1988 An algorithm for polygon conversion to boxes for VLSI layouts
Asim J. Al-Khalili, Dhamin Al-Khalili, Khaled Ammar
Integr.2