Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mark Zwolinski

dblp:99/2370 · DBLP profile ↗
← Back
54ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-2230-625XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 14Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 3Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Electronic design automation · 51% High-performance computing · 24% Reconfigurable computing and FPGAs · 12%
Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 79% Deep learning architectures and training · 21%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
circuit simulation
0.222015
Parallel Sparse Matrix Solution for Circuit Simulation on FPGAs · IEEE Trans. Computers 2015
Confidence in mixed-mode circuit simulation · Comput. Aided Des. 1992
Reconfigurable computing and FPGAs
FPGA accelerator
0.212015
Parallel Sparse Matrix Solution for Circuit Simulation on FPGAs · IEEE Trans. Computers 2015
High-performance computing
sparse linear algebra
0.212015
Parallel Sparse Matrix Solution for Circuit Simulation on FPGAs · IEEE Trans. Computers 2015
High-performance computing
sparse linear solver
0.212015
Parallel Sparse Matrix Solution for Circuit Simulation on FPGAs · IEEE Trans. Computers 2015
Electronic design automation › circuit simulation › analog circuit simulation
SPICE simulation
0.212015
Parallel Sparse Matrix Solution for Circuit Simulation on FPGAs · IEEE Trans. Computers 2015
Electronic design automation
high-level synthesis
0.232008
Symbolic noise analysis approach to computational hardware optimization · DAC 2008
An Integrated High-Level On-Line Test Synthesis Tool · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
On the Design of Self-Checking Controllers with Datapath Interactions · IEEE Trans. Computers 2006
Electronic design automation
hardware verification and test
0.122006
An Integrated High-Level On-Line Test Synthesis Tool · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Applying a robust heteroscedastic probabilistic neural network toanalog fault detection and classification · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Electronic design automation › high-level synthesis › arithmetic-level optimization
bit-width optimization
0.112008
Symbolic noise analysis approach to computational hardware optimization · DAC 2008
Hardware reliability and fault tolerance
error modeling
0.112008
Symbolic noise analysis approach to computational hardware optimization · DAC 2008
Integrated circuit design › digital circuit design
secure circuit design
0.112008
Divided Backend Duplication Methodology for Balanced Dual Rail Routing · CHES 2008
Hardware reliability and fault tolerance
self-checking circuits
0.112006
On the Design of Self-Checking Controllers with Datapath Interactions · IEEE Trans. Computers 2006
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.012001
Mutual Information Theory for Adaptive Mixture Models · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Information theory › information measures
mutual information
0.012001
Mutual Information Theory for Adaptive Mixture Models · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Electronic design automation › hardware verification and test › analog circuit testing
analog fault detection
0.012000
Applying a robust heteroscedastic probabilistic neural network toanalog fault detection and classification · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Electronic design automation › hardware test
fault classification
0.012000
Applying a robust heteroscedastic probabilistic neural network toanalog fault detection and classification · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Hardware security and side channels
side-channel countermeasures
0.012008
Divided Backend Duplication Methodology for Balanced Dual Rail Routing · CHES 2008
Electronic design automation › hardware verification and test
hardware verification
0.011992
Confidence in mixed-mode circuit simulation · Comput. Aided Des. 1992
Electronic design automation
logic synthesis
0.011992
Interleaving: an additional topological compaction technique for Weinberger array generation · Comput. Aided Des. 1992
Machine learning › Deep learning architectures and training › feedforward neural network › shallow neural networks
probabilistic neural network
0.012000
Applying a robust heteroscedastic probabilistic neural network toanalog fault detection and classification · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Electronic design automation › physical design › routing
global routing
0.011990
Lee router modified for global routing · Comput. Aided Des. 1990
Electronic design automation
physical design
0.011990
Lee router modified for global routing · Comput. Aided Des. 1990
Electronic design automation › physical design
routing
0.011990
Lee router modified for global routing · Comput. Aided Des. 1990

Methods — techniques the papers use, named apart from their topics

symbolic analysis · 0.2static pivoting · 0.2dual rail routing · 0.2backend duplication · 0.2symbolic noise analysis · 0.1probabilistic error modeling · 0.1parity checking · 0.1mutual information · 0.1inversion testing · 0.1duplication-based self-checking · 0.11-out-of-n checking · 0.1statistical classification · 0.0heteroscedastic probabilistic neural network · 0.0
YearPublicationVenuePosition
2025 Design of a Single-Event Upset Tolerant Low-Power Double-Tail Comparator
abstract
Comparators have been optimized for quick decision-making and reduction of dynamic power consumption through years of research by incorporating strong positive feedback latches. However, these advancements can make double-tail dynamic comparators more susceptible to single-event effects (SEEs). In this work, we present a new comparator design that is hardened against these radiation effects. The proposed design is a modified version of a low-voltage, low-power double-tail comparator, designed in a 180 nm technology node, and it can achieve high levels of SEE tolerance with an acceptable degree of trade-off in area, delay, and power consumption. The proposed comparator is shown to have superior SEE tolerance compared to a radiation-hardened conventional double-tail comparator with a similar electrical performance designed in the same technology node, and the proposed design’s functionality is proven by post-layout simulations across extreme simulation corners, combining process corners with temperature and supply voltage variations.
Ahmet Cirakoglu, Alexander Serb, Khaled Humood, Mark Zwolinski, Themistoklis Prodromakis
ISCAS4
2022 Session details: Session 4A: Testing, Reliability and Fault Tolerance
abstract
No abstract available.
Mark Zwolinski
ACM Great Lakes Symposium on VLSI1
2019 A reliable PUF in a dual function SRAM
Mohd Syafiq Mispan, Shengyu Duan, Basel Halak, Mark Zwolinski
Integr.4
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.56
2018 Cell Flipping with Distributed Refresh for Cache Ageing Minimization
abstract
CMOS wear-out mechanisms, especially Bias Temperature Instability (BTI), have caused growing concerns about circuit reliability. For cache memories, BTI reduces the static noise margin (SNM), causing unreliable read operations. In practice, error-correction codes (ECCs) are often used to protect data from transient errors in caches, but the limited error correction capabilities are not always enough to overcome BTIinduced read failures. In this paper, we propose a cell flipping technique with distributed refresh phases (CFDR) to minimize cache degradations. The CFDR method flips and refreshes each cache block at different times, minimizing the interruption time and balancing the degradation rate, even for infrequently replaced cache blocks. We evaluate the CFDR technique on an instruction cache in a 32-bit ARM architecture and show our method reduces the number of error bits by 58.86% and 13.59%, compared with an ECC scheme and a traditional cell flipping technique. The cache lifetime can be improved by 125% by using CFDR with less than 1% area overhead, which is not only more effective but also more cost-efficient than the existing techniques.
Shengyu Duan, Basel Halak, Mark Zwolinski
ATS3
2018 Cost-efficient design for modeling attacks resistant PUFs
abstract
Physical Unclonable Functions (PUFs) exploit the intrinsic manufacturing process variations to generate a unique signature for each silicon chip; this technology allows building lightweight cryptographic primitive suitable for resource-constrained devices. However, the vast majority of existing PUF design is susceptible to modeling attacks using machine learning technique, this means it is possible for an adversary to build a mathematical clone of the PUF that have the same challenge/response behavior of the device. Existing approaches to solve this problem include the use of hash functions, which can be prohibitively expensive and render PUF technology as the suitable candidate for lightweight security. This work presents a challenge permutation and substitution techniques which are both area and energy efficient. We implemented two examples of the proposed solution in 65-nm CMOS technology, the first using a delay-based structure design (an Arbiter-PUF), and the second using sub-threshold current design (two-choose-one PUF or TCO-PUF). The resiliency of both architectures against modeling attacks is tested using an artificial neural network machine learning algorithm. The experiment results show that it is possible to reduce the predictability of PUFs to less than 70% and a fractional area and power costs compared to existing hash function approaches.
Mohd Syafiq Mispan, Haibo Su, Mark Zwolinski, Basel Halak
DATE3
2018 Early detection of system-level anomalous behaviour using hardware performance counters
abstract
Embedded systems suffer from reliability issues such as variations in temperature and voltage, single event effects and component degradation, as well as being exposed to various security attacks such as control hijacking, malware, reverse engineering, eavesdropping and many others. Both reliability problems and security attacks can cause the system to behave anomalously. In this paper, we will present a detection technique that is able to detect a change in the system before the system encounters a failure, by using data from Hardware Performance Counters (HPCs). Previously, we have shown how HPC data can be used to create an execution profile of a system based on measured events and any deviation from this profile indicates an anomaly has occurred in the system. The first step in developing a detector is to analyse the HPC data and extract the features from the collected data to build a forecasting model. Anomalies are assumed to happen if the observed value falls outside a given confidence interval, which is calculated based on the forecast values and prediction confidence. The detector is designed to provide a warning to the user if anomalies that are detected occur consecutively for a certain number of times. We evaluate our detection algorithm on benchmarks that are affected by single bit flip faults. Our initial results show that the detection algorithm is suitable for use for this kind of univariate time series data and is able to correctly identify anomalous data from normal data.
Elena Lai Leng Woo, Mark Zwolinski, Basel Halak
DATE2
2018 Lifetime Reliability-Aware Digital Synthesis
Shengyu Duan, Mark Zwolinski, Basel Halak
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Fault analysis in analog circuits through language manipulation and abstraction
abstract
Each year automotive systems are becoming smarter thanks to their enhancement with sensing, actuation and computation features. The recent advancements in the field of autonomous driving have increased even more the complexity of the electronic components used to provide such services. ISO 26262 represents the natural response to the growing concerns in terms of the functional safety of electrical safety-related systems in this area. However, if the functional safety analysis of digital devices is quite a stable methodology, the same analysis for analog components is still in its infancy. This paper aims to explore the problem of fault analysis in analog circuits and how it can be integrated into the design processes with minimum effort. The methodology is based on analog language manipulation, analog fault instrumentation and automatic abstraction. An efficient and comprehensive flow for performing such an activity is proposed and applied to complex case studies.
Enrico Fraccaroli, Francesco Stefanni, Franco Fummi, Mark Zwolinski
FDL4
2017 A cost-efficient delay-fault monitor
abstract
Delay-fault monitoring sensors are widely used for Dynamic Voltage and Frequency Scaling (DVFS) to compensate for intrinsic Process, Voltage, Temperature and Ageing (PVTA) variations. Such techniques are generally based on monitoring the circuit’s critical paths. This paper presents a new delay-fault monitoring circuit, which is able to monitoring multiple paths simultaneously. The proposed circuitry has been designed and verified in a 32 bit MIPS processor using a 65nm technology. Our results indicate that the use of the proposed sensor for delay monitoring can lead to a significant saving in area and power overheads of two-thirds and one-third, respectively, compared to a canary flip-flop.
Gaole Sai, Basel Halak, Mark Zwolinski
ISCAS3
2016 The influence of hysteresis voltage on single event transients in a 65nm CMOS high speed comparator
abstract
Hysteresis in a comparator improves the input noise immunity, but can also cause analogue single event transients (ASETs) to be captured. For example, compared to a hysteresis-free comparator, a comparator with a hysteresis voltage of 8 mV, takes an additional 40 ns to recover. As the requirement for noise immunity increases, the vulnerability of a comparator with hysteresis to ASETs worsens. The reliability also worsens for higher sampling frequencies and lower differential input voltage amplitudes. This paper investigates the trade-off between noise immunity and reliability in a 65nm CMOS comparator.
Illani Mohd Nawi, Basel Halak, Mark Zwolinski
ETS3
2016 High accuracy implementation of Adaptive Exponential integrated and fire neuron model
abstract
It is expensive to simulate large-scale neural networks on hardware while ensuring a high resemblance to the original neurons' behavior. This paper introduces a novel technique to facilitate digital implementation and computer simulation of neuron models that contain an exponential term. This technique is applied to a biologically realistic neuron model called Adaptive Exponential integrated and fire (AdEx). Hardware synthesis and physical implementations show that the resulting model can reproduce precise neural behavior with high performance and considerably lower implementation costs compared with the original AdEx model.
Aliasghar Makhlooghpour, Hamid Soleimani, Arash Ahmadi, Mark Zwolinski, Mehrdad Saif
IJCNN4
2016 NBTI aging evaluation of PUF-based differential architectures
abstract
Silicon Physical Unclonable Functions (PUFs) have emerged as novel cryptographic primitives, with the ability to generate unique chip identifiers and cryptographic keys by exploiting intrinsic manufacturing process variations. The “Two Choose One” PUF (TCO-PUF) has recently been proposed. It is based on a differential architecture and exploits the non-linear relationship between current and voltage in the subthreshold operating region. As CMOS technology scales down, aging-induced Negative Bias Temperature Instability (NBTI) is becoming more pronounced, resulting in reliability issues for the PUF response. Differential design techniques can be useful for mitigating and canceling out first-order environmental dependencies such as aging, temperature and supply voltage. In this study, we investigate the robustness of PUFs with differential architectures, such as TCO-PUF and Arbiter-PUF, under the influence of NBTI. Our results indicate PUFs with differential architectures are less vulnerable to aging-related degradation compared to other PUF designs such as RO-PUF and SRAM-PUF. We show that the reliability of TCO-PUF and Arbiter-PUF only degrades by about 4.5% and 2.41%, respectively, after 10 years, while RO-PUFs and SRAM-PUFs degrade by about 12.76% in 10 years and 7% in 4.5 years, respectively.
Mohd Syafiq Mispan, Basel Halak, Mark Zwolinski
IOLTS3
2016 A Low-Cost, Radiation-Hardened Method for Pipeline Protection in Microprocessors
abstract
The aggressive scaling of semiconductor technology has significantly increased the radiation-induced soft-error rate in modern microprocessors. Meanwhile, due to the increasing complexity of modern processor pipelines and the limited error-tolerance capabilities that previous radiation hardening techniques can provide, the existing pipeline protection mechanisms cannot achieve complete protection. This paper proposes a complete and cost-effective pipeline protection mechanism using a self-checking architecture. The radiation-hardened pipeline is achieved by incorporating soft-error- and timing-error-tolerant flip-flop (SETTOFF)-based self-checking cells into the sequential cells of the pipeline. A replay recovery mechanism is also developed at the architectural level to recover the detected errors. The proposed pipeline protection technique is implemented in an OpenRISC microprocessor in a 65-nm technology. A gate-level transient fault-injection and analysis technique is used to evaluate the error-tolerance capability of the proposed hardened pipeline design. The results show that compared with the techniques such as triple modular redundancy, the SETTOFF-based self-checking technique requires over 30% less area and 80% less power overheads. Meanwhile, the error-tolerant and self-checking capabilities of the register allow the proposed pipeline protection technique to provide a noticeably higher level of reliability for different parts of the pipeline compared with the previous pipeline protection techniques.
Mark Zwolinski, Basel Halak
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Conservative behavioural modelling in systemc-AMS
abstract
SystemC has recently been extended with the Analogue and Mixed Signal (AMS) library, with the ultimate goal of providing simulation support to analogue electronics and continuous time behaviours. SystemC-AMS allows modelling of systems that are either conservative and extremely low level or continuous time and behavioural, which is limited compared to other AMS HDLs. This work faces up this challenge, by extending SystemCAMS support to a new level of abstraction, called Analogue Behavioural Modelling (ABM), covering models that are both behavioural and conservative. This leads to a methodology that uses SystemC-AMS constructs in a novel way. Full automation of the methodology allows proof of its effectiveness both in terms of accuracy and simulation performance, and application of the overall approach to a complex industrial Micro Electro- Mechanical System (MEMS) case study. The effectiveness of the proposed approach is further highlighted in the context of virtual platforms for smart systems, as adopting a C++-based language for MEMS simulation reduces the simulation time by about 2x, thus enhancing the design and integration flow.
Sara Vinco, Michele Lora, Mark Zwolinski
FDL3
2015 Parallel Sparse Matrix Solution for Circuit Simulation on FPGAs
abstract
SPICE is the de facto standard for circuit simulation. However, accurate SPICE simulations of today’s sub-micron circuits can often take days or weeks on conventional processors. A SPICE simulation is an iterative process that consists of two phases per iteration: model evaluation followed by a matrix solution. The model evaluation phase has been found to be easily parallelizable, unlike the subsequent phase, which involves the solution of highly sparse and asymmetric matrices. In this paper, we present an FPGA implementation of a sparse matrix solver, geared towards matrices that arise in SPICE circuit simulations. Our approach combines static pivoting with symbolic analysis to compute an accurate task flow-graph which efficiently exploits parallelism at multiple granularities and sustains high floating-point data rates. We also present a quantitative comparison between the performance of our hardware prototype and state-of-the-art software packages running on a general-purpose PC. We report average speed-ups of 9.65$\times$, 11.83$\times$, and 17.21$\times$against UMFPACK, KLU, and Kundert Sparse matrix packages, respectively.
Tarek Nechma, Mark Zwolinski
IEEE Trans. Computers2
2015 Resistive Open Faults Detectability Analysis and Implications for Testing Low Power Nanometric ICs
abstract
Resistive open faults (ROFs) represent common manufacturing defects in IC interconnects and result in delay faults that cause timing failures and reliability risks. The nonmonotonic dependence of ROF-induced delay faults on the supply voltage (VDD) poses a concern as to whether single-VDDtesting will suffice for low power nanometric designs. Our analysis shows multi-VDDtests could be required, depending on the test speed. This knowledge can be exploited in small delay fault testing to reduce the chances of test escapes while minimizing cost.
Mohamed Tagelsir Mohammadat, Noohul Basheer Zain Ali, Fawnizu Azmadi Hussin, Mark Zwolinski
IEEE Trans. Very Large Scale Integr. Syst.4
2014 A low-cost radiation hardened flip-flop
abstract
The aggressive scaling of semiconductor devices has caused a significant increase in the soft error rate caused by radiation hits. This has led to an increasing need for fault-tolerant techniques to maintain system reliability. Conventional radiation hardening techniques, typically used in safety-critical applications, are prohibitively expensive for non-safety-critical electronics. This work proposes a novel flip-flop architecture named SETTOFF which significantly improves circuit resilience to radiation hits over previous techniques. In addition, compared to other techniques such as a TMR latch, SETTOFF reduces the area and performance overheads by up to 50% and 80%, respectively; the power consumption overhead is also reduced by up to 85%. In addition, a novel reliability metric called radiation-induced failure rate is developed which can be a valuable tool to predict the impact of radiation hits and quantitatively compare the reliability of various radiation hardened techniques. Our analysis shows that the proposed technique can achieve zero SEU failure rate, and significantly reduce the SET failure rate.
Mark Zwolinski, Basel Halak
DATE2
2014 Efficient simulation and modelling of non-rectangular NoC topologies
abstract
With increasing chip complexity, Networks-on-Chips (NoCs) are becoming a central platform for future on-chip communications. Many regular NoC architectures have been proposed to eliminate the communication bottlenecks of traditional bus-based networks. Non-rectangular and irregular architectures have also been proposed to increase performance. However, the complexity of designing custom non-rectangular networks leads to a rapid increase in design and verification times. To alleviate the conflict between performance and efficiency, this paper proposes a novel method that efficiently constructs virtual non-rectangular topologies on a mesh network by using time-regulated models to emulate irregular patterns. Data routings on virtual hexagonal and two irregular geometries validate the proposed method. An MPEG-4 decoder is used to exemplify the proposed method for media applications. Results analysis shows the virtual topologies emulated by the proposed method can provide precise timing and energy performance.
Mark Zwolinski
DATE2
2014 Highly adaptive and congestion-aware routing for 3D NoCs
abstract
In this paper, we propose a novel highly adaptive and congestion aware routing algorithm 3D meshes which is equally applicable to 2D meshes as well. The proposed algorithm allows cyclic dependencies in channel dependency graph (CDG) providing higher degree of adaptiveness. The algorithm uses congestion-aware channel selection strategy that results balanced distribution of traffic flows across the network. A packet follows non-minimal paths only when minimal paths are congested at the neighboring channels. The deadlock avoidance methodology adopted by our algorithm remains cost-efficient as it uses one extra virtual channel along each of Y and Z dimensions to achieve deadlock freedom.
Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Seok-Bum Ko, Mark Zwolinski
ACM Great Lakes Symposium on VLSI6
2014 A cost-efficient self-checking register architecture for radiation hardened designs
abstract
The rapid development of CMOS technology has significantly increased the susceptibility of electronic systems to radiation-induced soft errors. Conventional error-tolerant techniques typically use redundancies to mitigate soft errors and increase system immunity. However they do not have self-checking capabilities, and therefore are still vulnerable to the errors in the redundant circuitry added for error-tolerance. This paper proposes a novel self-checking soft error-tolerant register based on SETTOFF, a Soft Error and Timing error Tolerant Flip-Flop. The register significantly improves the error-tolerant capability over previous techniques since it has a self-checking capability, which allows the register to tolerate both the errors in the original flip-flops and the redundant circuitry. In addition, the register can also tolerate both soft errors (SETs and SEUs) and timing errors. Compared with other previous techniques such as TMR, the proposed register reduces the power consumption overhead by 81%, and the delay overhead by 54% in 65nm technology; The area overhead is also reduced by 25%.
Mark Zwolinski
ISCAS2
2014 A novel non-minimal/minimal turn model for highly adaptive routing in 2D NoCs
abstract
Networks-on-Chip (NoCs) are emerging as a promising communication paradigm to overcome bottleneck of traditional bus-based interconnects for current micro-architectures (MCSoC and CMP). One of the current issues in NoC routing is the use of acyclic Channel Dependency Graph (CDG) for deadlock freedom. This requirement forces certain routing turns to be prohibited, thus, reducing the degree of adaptiveness. In this paper, we propose a novel non-minimal turn model which allows cycles in CDG provided that Extended Channel Dependency Graph (ECDG) remains acyclic. The proposed turn model reduces number of restrictions on routing turns, hence able to provide path diversity through additional minimal and non-minimal routes between source and destination.
Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Pankaj Kumar Srivastava, Seok-Bum Ko, Mark Zwolinski
NOCS7
2014 A novel non-minimal turn model for highly adaptive routing in 2D NoCs
abstract
Network-on-Chip (NoC) is emerging as a promising communication paradigm to overcome bottleneck of traditional bus-based interconnects for future micro-architectures (MPSoC and CMP). One of current issue in NoC routing is the use of acyclic channel dependency graph (ACDG) for deadlock freedom prohibiting certain routing turns. Thus, ACDG reduces the degree of adaptiveness. In this paper, we propose a novel nonminimal turn model which allows cycles in channel dependency graph provided that extended channel dependency graph is acyclic. Proposed turn model reduces number of restrictions on routing turns (specially on 90-degree), hence able to provide additional minimal and non-minimal routes between source and destination. We also propose a non-minimal and congestion-aware adaptive routing algorithm based on proposed turn model to demonstrate advantages. From results, we can observe that proposed method improves the network performance by distributing the traffic load in the non-congested regions.
Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Mark Zwolinski
VLSI-SoC5
2014 Multivoltage Aware Resistive Open Fault Model
abstract
Resistive open faults (ROFs) represent common interconnect manufacturing defects in VLSI designs causing delay failures and reliability-related concerns. The widespread utilization of multiple supply voltages in contemporary VLSI designs and emerging test methods poses a critical concern as to whether conventional models for resistive opens will still be effective. Conventional models do not explicitly model the VDDeffect on fault behavior and detectability. We have empirically observed that a sensitized ROF could exhibit multiple behaviors across its resistance continuum. We also observe that the detectable resistance range versus VDDvaries with test speed. We consequently propose a voltage-aware model that divides the full range of open resistances into continuous behavioral intervals and three detectability ranges. The presented model is expected to substantially enhance multivoltage test generation and fault distinction.
Mohamed Tagelsir Mohammadat, Noohul Basheer Zain Ali, Fawnizu Azmadi Hussin, Mark Zwolinski
IEEE Trans. Very Large Scale Integr. Syst.4
2013 ISPD 2013 expert designer/user session (eds)
abstract
We have newly introduced the expert designer/user session (EDS) to ISPD in 2013, which is tailor-made for designers and users of physical design (PD) tools. Since ISPD is the premium PD-centric symposium, it is a great opportunity for designers, tools developers and PD researchers to interact and learn from each other. The benefit of including EDS in ISPD's program is twofold.
Cliff C. N. Sze, Laleh Behjat, Nikhil Jayakumar, Atul Walimbe, Gregory Ford, Mark Zwolinski, Harish Dangat, Giriraj Kakol
ISPD6
2012 SETTOFF: A fault tolerant flip-flop for building Cost-efficient Reliable Systems
abstract
Conventional fault tolerance techniques either require big overheads or have limited reliability. We propose a novel fault tolerant flip-flop (SETTOFF) that addresses timing errors and soft errors in one cost-efficient architecture. In SETTOFF, most SEUs are detected by monitoring the illegal transitions at the output of a flip-flop and recovered by inverting the cell state. SETs, timing errors and the other SEUs are detected by a time redundancy-based architecture. For a 10% activity rate, SETTOFF consumes 35.8% and 39.7% more power than a library flip-flop in 120nm and 65nm technologies, respectively. It only consumes about 5.7% more power than the detection based RazorII flip-flop [1]. SETTOFF therefore provides an increased coverage of fault tolerance with only moderate increase in overhead, hence it is suitable for building highly reliable systems at lower cost than the traditional techniques.
Mark Zwolinski
IOLTS2
2011 Modelling circuit performance variations due to statistical variability: Monte Carlo static timing analysis
abstract
The scaling of MOSFETs has improved performance and lowered the cost per function of CMOS integrated circuits and systems over the last 40 years, but devices are subject to increasing amounts of statistical variability within the deca-nano domain. The causes of these statistical variations and their effects on device performance have been extensively studied, but there have been few systematic studies of their impact on circuit performance. This paper describes a method for modelling the impact of random intra-die statistical variations on digital circuit timing and power consumption. The method allows the variation modelled by large-scale statistical transistor simulations to be propagated up the design flow to the circuit level, by making use of commercial STA and standard cell characterisation tools. The method provides circuit designers with the information required to analyse power, performance and yield trade-offs when fabricating a design, while removing the large levels of pessimism generated by traditional Corner Based Analysis.
Michael Merrett, Plamen Asenov, Mark Zwolinski, Dave Reid, Campbell Millar, Scott Roy, Steve Furber, A. Asenov
DATE4
2011 Timing Vulnerability Factors of Ultra Deep-sub-micron CMOS
abstract
Soft errors are a significant reliability issue for Ultra Deep-Sub-Micron (UDSM) CMOS circuits. Therefore, an accurate assessment of the Soft Error Rate (SER) is crucial. In this paper, we argue that the conventional definitions for the Window of Vulnerability (WOV) under-estimate the risk. We propose a new method for determining the timing factors and WOV for the sequential elements from the susceptibility perspective rather than the conventional performance perspective. We also discuss that the process variation does not have any special impact on the WOV. Our methodology leads to a more realistic definition of the WOV for SER computation.
Massoud Mokhtarpour Ghahroodi, Mark Zwolinski, Richard Wong, Shi-Jie Wen
ETS2
2011 Parallelizing TUNAMI-N1 Using GPGPU
abstract
We present a high performance tsunami-prediction system using General Purpose Graphics Processing Units (GPGPU). It is based on TUNAMI-N1, a Numerical Analysis Model for Investigation of near-field tsunamis. It uses linear shallow water wave equations, commonly accepted approximation for tsunami propagation, taking the input from a bathymetry file containing a large data set. Due to the largeness of the data set, the model is more amenable to parallelization. The system maps the TUNAMI-N1 model into the massively parallel GPU architecture using Nvidia CUDA framework. It employs multiple kernels that contain inherently parallel portion of the model and uses the concepts of data and hybrid parallelism to fully exploit the hardware capabilities of the GPUs. Experimental results show that our system achieves a speed up of six times.
Harsh Gidra, Israrul Haque, Nitin P. Kumar, M. Sargurunathan, Manoj Singh Gaur, Vijay Laxmi, Mark Zwolinski, Virendra Singh
HPCC7
2011 Acceleration of packet filtering using gpgpu
abstract
Packet filtering is core functionality in many academic and corporate network systems. Firewalls use a rule database to decide which packets will be allowed from one network onto another thereby implementing a security policy. With the introduction of new types of services and applications there is a growing demand for larger bandwidth and also for improved security. Both demands are in conflict since providing security partly relies on screening packet traffic, which implies a considerable overhead. In such a scenario as LAN and WAN speeds are becoming comparable, a single firewall can become a bottleneck and reduces the overall throughput of the network. A firewall with heavy load and limited processing power, which is supposed to be a first line of defence against attacks, becomes susceptible to Denial of Service (DoS) attacks. Many research groups have proposed different methods to improve efficiency and throughput to optimize firewalls. This paper presents and analyse various parallel implementations of packet filtering running on cost effective GPGPU. We describe an approach to efficiently exploit the massively parallel capabilities of the GPGPU.
Manoj Singh Gaur, Vijay Laxmi, Lakshminarayanan V., Kamal Cahndra, Mark Zwolinski
SIN5
2010 Modelling Smart Card Security Protocols in SystemC TLM
abstract
Smart cards are an example of advanced chip technology. They allow information transfer between the card holder and the system over secure networks, but they contain sensitive data related to both the card holder and the system, that has to be kept private and confidential. The objective of this work is to create an executable model of a smart card system, including the security protocols and transactions, and to examine the strengths and determine the weaknesses by running tests on the model. The security objectives have to be considered during the early stages of systems development and design, an executable model will give the designer the advantage of exploring the vulnerabilities early, and therefore enhancing the system security. The Unified Modeling Language (UML) 2.0 is used to model the smart card security protocol. The executable model is programmed in SystemC with the Transaction Level Modeling (TLM) extensions. The final model was used to examine the effectiveness of a number of authentication mechanisms with different probabilities of failure. In addition, a number of probable attacks on the current security protocol were modeled to examine the vulnerabilities. The executable model shows that the smart card system security protocols and transactions need further improvement to withstand different types of security attacks.
Aisha Fouad Bushager, Mark Zwolinski
EUC2
2010 Design metrics for RTL level estimation of delay variability due to intradie (random) variations
abstract
A simple metric is presented for the accurate prediction of path delay variability during the automated synthesis of digital VLSI circuits. This allows circuit variability to be assessed at early stages within the design process with minimal computational effort, as extensive Monte Carlo or SSTA runs are not required. This paper introduces the metric and investigates its effectiveness. The final predictions of path delay variability are found to be within 10% of measured path delay variability, with an average error of 3%, for a series of test paths synthesised from randomised models of a 130nm technology library. These randomised models are generated from a 3D atomistic simulator and provide more accuracy than traditional Monte Carlo simulation runs.
Michael Merrett, Mark Zwolinski, Koushik Maharatna, Massimo Alioto
ISCAS3
2010 Parallel sparse matrix solver for direct circuit simulations on FPGAs
abstract
As part of our effort to parallelised SPICE simulations over multiple FPGAs, we present a parallel FPGA implementation for a sparse matrix solver optimised for execution on a single FPGA node. Our approach combines static pivoting with symbolic analysis to compute an accurate task flow-graph which efficiently exploits parallelism at multiple granularities and sustains high floating-point data rates. The sparse matrix solver is tested with circuit simulation matrices from the University of Florida matrix collection. We report a 10-30× speedup compared to a 2.4 GHz Intel Core Duo processor running UMFPACK, a state-of-the-art sparse matrix solver.
Tarek Nechma, Mark Zwolinski, Jeffrey S. Reeve
ISCAS2
2009 Variation resilient adaptive controller for subthreshold circuits
abstract
Subthreshold logic is showing good promise as a viable ultra-low-power circuit design technique for power-limited applications. For this design technique to gain widespread adoption, one of the most pressing concerns is how to improve the robustness of subthreshold logic to process and temperature variations. We propose a variation resilient adaptive controller for subthreshold circuits with the following novel features: new sensor based on time-to-digital converter for capturing the variations accurately as digital signatures, and an all-digital DC-DC converter incorporating the sensor capable of generating an operating operating Vddfrom 0 V to 1.2 V with a resolution of 18.75 mV, suitable for subthreshold circuit operation. The benefits of the proposed controller is reflected with energy improvement of up to 55% compared to when no controller is employed. The detailed implementation and validation of the proposed controller is discussed.
Biswajit Mishra, Bashir M. Al-Hashimi, Mark Zwolinski
DATE3
2009 Analytical Transient Response and Propagation Delay Model for Nanoscale CMOS Inverter
abstract
This paper presents a new analytical propagation delay model for nanoscale CMOS inverters. By using a non-saturation current model, the analytical input-output transfer responses and propagation delay model are derived. The model is used for calculating inverter delays for different input transition times, load capacitances and supply voltages. Delays predicted by the proposed model are in good agreement with those of transistor level simulation results from SPICE, with accuracy of 3% or better.
Mark Zwolinski
ISCAS2
2008 Divided Backend Duplication Methodology for Balanced Dual Rail Routing
Karthik Baddam, Mark Zwolinski
CHES2
2008 Symbolic noise analysis approach to computational hardware optimization
abstract
This paper addresses the problem of computational error modeling and analysis. Choosing different word-lengths for each functional unit in hardware implementations of numerical algorithms always results in an optimization problem of trading computational error with implementation costs. In this study, a symbolic noise analysis method is introduced for high-level synthesis, which is based on symbolic modeling of the error bounds where the error symbols are considered to be specified with a probability distribution function over a known range. The ability to combine word-length optimization with high-level synthesis parameters and costs to minimize the overall design cost is demonstrated using case studies.
Arash Ahmadi, Mark Zwolinski
DAC2
2007 Multiple-Width Bus Partitioning Approach to Datapath Synthesis
abstract
A shared bus is a suitable structure for minimizing the interconnections costs in system synthesis. It has also been shown that the word-length of functional units has a great impact on design costs. A combination of both methods is used in this paper in the form of a partitioned shared bus structure, in which every partition has a different width and all the functional units connected to a bus partition have the same input/output word-lengths. Having controlled the group binding and word-length of the FUs as well as the other synthesis parameters, a high-level synthesis tool is introduced to implement DSP algorithms in digital hardware. The tool uses a multi-objective optimization genetic algorithm to minimize the circuit area, delay, power consumption and digital noise by selecting an optimal grouping and word-length for each FU in a shared bus system. Results demonstrate that savings can be made in the overall system costs by applying this method.
Arash Ahmadi, Mark Zwolinski
ISCAS2
2006 Dynamic Voltage Scaling Aware Delay Fault Testing
abstract
The application of Dynamic Voltage Scaling (DVS) to reduce energy consumption may have a detrimental impact on the quality of manufacturing tests employed to detect permanent faults. This paper analyses the influence of different voltage/frequency settings on fault detection within a DVS application. In particular, the effect of supply voltage on different types of delay faults is considered. This paper presents a study of these problems with simulation results. We have demonstrated that the test application time increases as we reduce the test voltage. We have also shown that for newer technologies we do not have to go to very low voltage levels for delay fault testing. We conclude that it is necessary to test at more than one operating voltage and that the lowest operating voltage does not necessarily give the best fault cover.
Noohul Basheer Zain Ali, Mark Zwolinski, Bashir M. Al-Hashimi, Peter Harrod
ETS2
2006 On the Design of Self-Checking Controllers with Datapath Interactions
abstract
We consider the problem of designing self-checking controllers for controller/datapath architectures. We introduce the concept of intrinsically secure states. We present six alternative schemes based on parity checking, on 1-out-of-n checking, as well as on the observation that a self-checking sequential datapath can also be employed for control path self-checking by exploiting the concept of intrinsically secure control states. A high-level synthesis tool has been modified to automatically insert self-checking controllers and datapath units and is able to trade this self-checking property against other design objectives. We discuss the properties of each configuration and present experimental results and conclusions
Petros Oikonomakos, Mark Zwolinski
IEEE Trans. Computers2
2006 An Integrated High-Level On-Line Test Synthesis Tool
abstract
Several researchers have recently implemented on-line testability in the form of duplication-based self-checking digital system design, early in the design process. The authors consider the on-line testability within the optimization phase of iterative, cost function-driven high-level synthesis, such that self-checking resources are inserted automatically without any modification of the source behavioral hardware description language code. This is enabled by introducing a metric for the on-line testability. A new variation of duplication (namely inversion testing) is also proposed and used, providing the system with an additional degree of freedom for minimizing hardware overheads associated with test resource insertion. Considering the on-line testability within the synthesis process facilitates fast and painless design space exploration, resulting in a versatile high-level-synthesis process, capable of producing alternative realizations according to the designer's directions, for alternative target technologies. Finally, the fault escape probability of the overall scheme is discussed theoretically and evaluated experimentally
Petros Oikonomakos, Mark Zwolinski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2004 Generation and Verification of Tests for Analog Circuits Subject to Process Parameter Deviations
Stephen J. Spinks, Chris D. Chalk, Ian M. Bell, Mark Zwolinski
J. Electron. Test.4
2003 Versatile High-Level Synthesis of Self-Checking Datapaths Using an On-Line Testability Metric
Petros Oikonomakos, Mark Zwolinski, Bashir M. Al-Hashimi
DATE2
2003 Foundation of Combined Datapath and Controller Self-checking Design
abstract
We consider the problem of designing self-checking controllers for applications with sequential datapaths. Firstly we compare encoded and unencoded (one-hot) controller implementations and we argue that self-checking of encoded control signals is not sufficient in terms of testability. Subsequently, we present four alternative controller self-checking schemes, based both on parity and on the observation that a self-checking data path can be employed for control path self-checking as well, by exploiting intrinsically secure control states. We discuss the properties of each of them, and present a few experimental results.
Petros Oikonomakos, Mark Zwolinski
IOLTS2
2003 Globally convergent algorithms for DC operating point analysis of nonlinear circuits
abstract
An important objective in the analysis of an electronic circuit is to find its quiescent or dc operating point. This is the starting point for performing other types of circuit analysis. The most common method for finding the dc operating point of a nonlinear electronic circuit is the Newton-Raphson method (NR), a gradient search technique. There are known convergence issues with this method. NR is sensitive to starting conditions. Hence, it is not globally convergent and can diverge or oscillate between solutions. Furthermore, NR can only find one solution of a set of equations at a time. This paper discusses and evaluates a new approach to dc operating-point analysis based on evolutionary computing. Evolutionary algorithms (EAs) are globally convergent and can find multiple solutions to a problem by using a parallel search. At the operating point(s) of a circuit, the equations describing the current at each node are consistent and the overall error has a minimum value. Therefore, we can use an EA to search the solution space to find these minima. We discuss the development of an analysis tool based on this approach. The principles of computer-aided circuit analysis are briefly discussed, together with the NR method and some of its variants. Various EAs are described. Several such algorithms have been implemented in a full circuit-analysis tool. The performance and accuracy of the EAs are compared with each other and with NR. EAs are shown to be robust and to have an accuracy comparable to that of NR. The performance is, at best, two orders of magnitude worse than NR, although it should be noted that time-consuming setting of initial conditions is avoided.
Duncan Crutchley, Mark Zwolinski
IEEE Trans. Evol. Comput.2
2002 Using evolutionary and hybrid algorithms for DC operating point analysis of nonlinear circuits
abstract
Traditionally, the DC operating points of a nonlinear electronic circuit are found using the Newton-Raphson method, which has known problems. It is not globally convergent; it can frequently diverge; and cannot find multiple solutions in a single pass. We discuss the use of evolutionary algorithms to overcome these problems.
Duncan Crutchley, Mark Zwolinski
IEEE Congress on Evolutionary Computation2
2002 Behavioural Modelling of Operational Amplifier Faults Using VHDL-AMS
abstract
The use of behavioural modelling for operational amplifiers has been well known for many years and previous work has included modelling of specific fault conditions using a macro-model. In this paper, the models are implemented in a more abstract form using an Analogue Hardware Description Language (AHDL), VHDL-AMS, taking advantage of the ability to control the behaviour of the model using high-level fault condition states. The implementation method allows a range of fault conditions to be integrated without switching to a completely new model. The various transistor faults are categorised, and used to characterise the behaviour of the HDL models. Simulations compare the accuracy and speed of the transistor and behavioural level models under a set of representative fault conditions.
Peter R. Wilson, J. Neil Ross, Mark Zwolinski, Andrew D. Brown, Yavuz Kiliç
DATE3
2001 Mutual Information Theory for Adaptive Mixture Models
abstract
Many pattern recognition systems need to estimate an underlying probability density function (pdf). Mixture models are commonly used for this purpose in which an underlying pdf is estimated by a finite mixing of distributions. The basic computational element of a density mixture model is a component with a nonlinear mapping function, which takes part in mixing. Selecting an optimal set of components for mixture models is important to ensure an efficient and accurate estimate of an underlying pdf. Previous work has commonly estimated an underlying pdf based on the information contained in patterns. In this paper, mutual information theory is employed to measure whether two components are statistically dependent. If a component has small mutual information, it is statistically independent of the other components. Hence, that component makes a significant contribution to the system pdf and should not be removed. However, if a particular component has large mutual information, it is unlikely to be statistically independent of the other components and may be removed without significant damage to the estimated pdf. Continuing to remove components with large and positive mutual information will give a density mixture model with an optimal structure, which is very close to the true pdf.
Zheng Rong Yang, Mark Zwolinski
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Applying Mutual Information to Adaptive Mixture Models
Zheng Rong Yang, Mark Zwolinski
IDEAL2
2000 Applying a robust heteroscedastic probabilistic neural network toanalog fault detection and classification
abstract
The problem of distinguishing and classifying the responses of analog integrated circuits containing catastrophic faults has aroused recent interest. The problem is made more difficult when parametric variations are taken into account. Hence, statistical methods and techniques such as neural networks have been employed to automate classification. The major drawback to such techniques has been the implicit assumption that the variances of the responses of faulty circuits have been the same as each other and the same as that of the fault-free circuit. This assumption can be shown to be false. Neural networks, moreover, have proved to be slow. This paper describes a new neural network structure that clusters responses assuming different means and variances. Sophisticated statistical techniques are employed to handle situations where the variance tends to zero, such as happens with a fault that causes a response to be stuck at a supply rail. Two example circuits are used to show that this technique is significantly more accurate than other classification methods. A set of responses can be classified in the order of 1 s.
Zheng Rong Yang, Mark Zwolinski, Chris D. Chalk, Alan Christopher Williams
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 Fast, Robust DC and Transient Fault Simulation for Nonlinear Analog Circuits
abstract
The evaluation of analogue and mixed-signal test strategies and design for test techniques requires the fault simulation of analogue circuits. The need to reduce fault simulation time for has resulted in the research into concurrent analogue fault simulation, analogous to digital fault simulation. Concurrent simulation can reduce the simulation time by avoiding repeated construction of the circuit matrix. Fault collapsing and dropping is also desirable. A robust, fast algorithm for concurrent analogue fault simulation is presented in this paper. Three techniques for the automatic dropping of faults have been addressed: a robust closeness measurement technique; a late start rule and an early stop rule. The algorithm has been successfully applied to both DC and transient analyses. A significant increase in the speed of analogue fault simulation has been obtained.
Zheng Rong Yang, Mark Zwolinski
DATE2
1992 Interleaving: an additional topological compaction technique for Weinberger array generation
Keith Richard Baker, Mark Zwolinski
Comput. Aided Des.2
1992 Confidence in mixed-mode circuit simulation
Andrew D. Brown, Mark Zwolinski, Ken G. Nichols, Tom J. Kazmierski
Comput. Aided Des.2
1990 Lee router modified for global routing
Andrew D. Brown, Mark Zwolinski
Comput. Aided Des.2