Yvon Savaria

dblp:s/YvonSavaria · DBLP profile ↗
← Back
178ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-3404-9959ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 144 · 3 first-author · 18 since 2021Computer networks · 15 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 14 · 2 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 IMICLiVAN: an improved method to increase cluster lifetime in vehicular ad hoc networks (VANETs)
Erfan Moghadam, Elham Asghari, Seyed Amir Asghari, Mohammadreza Binesh Marvasti, Yvon Savaria
J. Supercomput.5
2025 A Prototyping Framework for P4-Programmable Traffic Managers
abstract
International audience
Karl La Grassa, André Béliveau, Mathieu Léonardon, Jean-Pierre David, Matthieu Arzel, Yvon Savaria
RSP6
2025 Enabling Rank-Based P4 Programmable Schedulers: Requirements, Implementation, and Evaluation on BMv2 Switches
abstract
Software-defined networking (SDN) has revolutionized network infrastructure, offering programmability to meet evolving network demands. However, the fixed-function nature of the packet scheduler in current network equipment impedes the exploration of scheduling policies within a programmable network environment. This paper proposes a novel methodology to implement rank-based programmable schedulers in programmable BMv2 switches expressed with the network-specific programming language (P4). A proposed custom networking environment facilitates the study and evaluation of various scheduling policies. This environment is used to implement 20 different scheduling and shaping policies to identify the required language constructs and components needed to express these policies with the P4 language efficiently. Our experiments reveal that specific scheduling policies do not seamlessly align with a previously proposed architecture for rank-based scheduling policies. Thus, we propose rank-based versions for five previously reported scheduling policies, making them efficiently implementable in any rank-based schedulers and programmable network equipment. The reported results confirm that the rank-based versions of these scheduling policies accurately replicate the behavior and performance of the original policies, with a maximum error of 0.5% in the resulting flow completion times (FCTs).
Mostafa Elbediwy, Bill Pontikakis, Jean-Pierre David, Yvon Savaria
IEEE Trans. Netw.4
2024 Utilization of Noise-Shaping in Mixed-Signal Timing-Skew Mismatch Calibration of TI-ADCs
abstract
This paper proposes an approach based on oversampling and noise-shaping mechanisms to mitigate the implementation complexity of variable delay lines for mixed-signal calibration of timing-skew mismatch in time-interleaving ADCs. The proposed technique pushes the error (introduced by increasing the delay steps) outside the operating frequencies. Our method is more appropriate for noise-shaping time-interleaved ADCs since they are band-limited. However, the proposed method permits the utilization of either the high-frequency or low-frequency zones. This approach avoids utilizing complex clock routing methods. Also, it does not restrict the number of sub-ADCs. The proposed method is verified by behavioral simulations with MATLAB/Simulink. The mean of the spurious-free dynamic range (SFDR) is respectively enhanced by 11.64 dB and 17.72 dB for the low-frequency and high-frequency zones.
Hamidreza Mafi, Mohamed Amine Bensenouci, Sadok Aouini, Mohammad Honarparvar, Naim Ben-Hamida, Yvon Savaria
ISCAS6
2024 Enhancing P4 Syntax to Support Extended Finite State Machines as Native Stateful Objects
abstract
The P4 language has proven to be a powerful tool for programming packet processing, but its original design did not intend to handle stateful processing effectively. This shortcoming stems from the fact that the network switches for which the language was designed have restricted memory capacities, which makes it challenging to manage complex stateful objects. As a result, P4’s syntax was not optimized for handling such objects. With contemporary networks increasingly relying on stateful processing and abstractions like Extended Finite State Machines (EFSMs), we propose extending P4’s syntax through an EFSM construct. This work aims to grant developers the ability to create streamlined and productive P4 programs that can effortlessly deal with stateful objects. This improvement holds great promise for expanding P4’s functionality and refining it to support stateful processing.
Florent Allard, Tarek Ould Bachir, Yvon Savaria
NetSoft3
2024 Incremental reinforcement learning for multi-objective analog circuit design acceleration
Ahmed Abuelnasr, Ahmed Ragab, Mostafa Amer, Benoit Gosselin, Yvon Savaria
Eng. Appl. Artif. Intell.5
2024 Statistical Hardware Design With Multimodel Active Learning
abstract
With the rising complexity of numerous novel applications that serve our modern society comes the strong need to design efficient computing platforms. Designing efficient hardware is, however, a complex multiobjective problem that deals with multiple parameters and their interactions. Given that there is a large number of parameters and objectives involved in hardware design, synthesizing all possible combinations is not a feasible method to find the optimal solution. One promising approach to tackle this problem is statistical modeling of a desired hardware performance. Here, we propose a model-based active learning approach to solve this problem. Our proposed method uses Bayesian models to characterize various aspects of hardware performance. We also use acrlong TL and Gaussian regression bootstrapping techniques in conjunction with active learning to create more accurate models. Our proposed statistical modeling method provides hardware models that are sufficiently accurate to perform design space exploration (DSE) as well as performance prediction simultaneously. We use our proposed method to perform DSE and performance prediction for various hardware setups, such as micro-architecture design and OpenCL kernels for FPGA targets. Our experiments show that the number of samples required to create performance models significantly reduces while maintaining the predictive power of our proposed statistical models. For instance, in our performance prediction setting, the proposed method needs 65% fewer samples to create the model, and in the DSE setting, our proposed method can find the best parameter settings by exploring fewer than 50 samples.
Alireza Ghaffari, Masoud Asgharian, Yvon Savaria
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 DR-PIFO: A Dynamic Ranking Packet Scheduler Using a Push-In-First-Out Queue
abstract
Software-defined Networking (SDN) introduced the decoupling of control and data forwarding planes. Despite advances in the programmability of SDNs, there remains a strong need for a fully programmable packet scheduler in the data plane. In this context, the ability to adapt to various traffic patterns and the expressiveness of schedulers are of paramount importance. This paper introduces the Dynamic Ranking Push-In-First-Out (DR-PIFO), as an algorithmic model that can be used to develop programmable packet schedulers based on PIFO queues. The DR-PIFO is a highly expressive model, capable of expressing a wide range of work-conserving, non-work-conserving, and hierarchical scheduling algorithms. Additionally, its dynamic ranking capabilities allow for real-time updates to the packet’s priority within the scheduler. The proposed solution also performs error detection in the departure order of packets, which is essential to avoid starvation in strict priority scheduling. These features are crucial when implementing popular scheduling algorithms such as the pFabric. The DR-PIFO is evaluated through its algorithmic properties and by implementing two distinct case studies. Its performance is further evaluated by incorporating it as an external module, written in a high-level language, and integrating it with software switches implemented using the P4 language. The results illustrate the superior expressiveness of DR-PIFO over state-of-the-art models such as PIFO and PIEO and confirm that it is an algorithm-agnostic model. Thus, DR-PIFO represents a promising solution for implementing more fully programmable packet schedulers in SDNs, with the potential to improve performance and adaptability.
Mostafa Elbediwy, Bill Pontikakis, Alireza Ghaffari, Jean-Pierre David, Yvon Savaria
IEEE Trans. Netw. Serv. Manag.5
2023 BARVINN: Arbitrary Precision DNN Accelerator Controlled by a RISC-V CPU
abstract
We present a DNN accelerator that allows inference at arbitrary precision with dedicated processing elements that are configurable at the bit level. Our DNN accelerator has 8 Processing Elements controlled by a RISC-V controller with a combined 8.2 TMACs of computational power when implemented with the recent Alveo U250 FPGA platform. We develop a code generator tool that ingests CNN models in ONNX format and generates an executable command stream for the RISC-V controller. We demonstrate the scalable throughput of our accelerator by running different DNN kernels and models when different quantization levels are selected. Compared to other low precision accelerators, our accelerator provides run time programmability without hardware reconfiguration and can accelerate DNNs with multiple quantization levels, regardless of the target FPGA size. BARVINN is an open source project and it is available at https://github.com/hossein1387/BARVINN.
MohammadHossein AskariHemmat, Sean Wagner, Olexa Bilaniuk, Yassine Hariri, Yvon Savaria, Jean-Pierre David
ASP-DAC5
2023 Quark: An Integer RISC-V Vector Processor for Sub-Byte Quantized DNN Inference
abstract
In this paper, we present Quark, an integer RISC-V vector processor specifically tailored for sub-byte DNN inference. Quark is implemented in GlobalFoundries' 22FDX FD-SOI technology. It is designed on top of Ara, an open-source 64-bit RISC-V vector processor. To accommodate sub-byte DNN inference, Quark extends Ara by adding specialized vector instructions to perform sub-byte quantized operations. We also remove the floating-point unit from Quarks' lanes and use the CVA6 RISC-V scalar core for the re-scaling operations that are required in quantized neural network inference. This makes each lane of Quark 2 times smaller and 1.9 times more power efficient compared to the ones of Ara. In this paper we show that Quark can run quantized models at sub-byte precision. Notably we show that for 1-bit and 2-bit quantized models, Quark can accelerate computation of Conv2d over various ranges of inputs and kernel sizes.
MohammadHossein AskariHemmat, Théo Dupuis, Yoan Fournier, Nizar El Zarif, Matheus A. Cavalcante, Matteo Perotti, Frank K. Gürkaynak, Luca Benini, François Leduc-Primeau, Yvon Savaria, Jean-Pierre David
ISCAS10
2023 An Area-efficient Memory-based Architecture for P4-programmable Streaming Parsers in FPGAs
abstract
Moving toward software-defined networking and function virtualization, flexibility and reconfigurability of the network have become more and more critical. Packet parsing, the first processing stage of programmable switches, requires high performance and reconfigurability to allow implementing low-latency and highly flexible data networks. This paper proposes an overlay architecture for an FPGA-based P4-programmable streaming packet parser. The purpose of this architecture is to allow supporting different functionality with a fixed hardware design by changing a program stored in an embedded memory. This program is derived from the parser section of a P4 code, describing a parsing graph. This approach eliminates a pipeline of parsing blocks in favor of a single parsing block, thereby reducing the design's complexity. Our architecture offers an 11 Gb/s data rate on a Xilinx Virtex-7 XC7VX690 FPGA, while its implementation requires 312 LUTs and 1135 FFs.
Parisa Mashreghi-Moghadam, Tarek Ould Bachir, Yvon Savaria
ISCAS3
2023 Symbolic Analysis for Data Plane Programs Specialization
abstract
Programmable network data planes have extended the capabilities of packet processing in network devices by allowing custom processing pipelines and agnostic packet processing. While a variety of applications can be implemented on current programmable data planes, there are significant constraints due to hardware limitations. One way to meet these constraints is by optimizing data plane programs. Program optimization can be achieved by specializing code that leverages architectural specificity or by compilation passes. In the case of programmable data planes, to respond to the varying requirements of a large set of applications, data plane programs can target different architectures. This leads to difficulties when developers want to reuse the code. One solution to that is to use compiler optimization techniques. We propose performing data plane program specialization to reduce the generated program size. To this end, we propose to specialize in programs written in P4, a Domain Specific Language (DSL) designed for specifying network data planes. The proposed method takes advantage of key aspects of the P4 language to perform a symbolic analysis on a P4 program and then partially evaluate the program to specialize it. The approach we propose is independent of the target architecture. We evaluate the specialization technique by implementing a packet deparser on an FPGA. The results demonstrate that program specialization can reduce the resource usage by a factor of 2 for various packet deparsers.
Thomas Luinaud, J. M. Pierre Langlois, Yvon Savaria
ACM Trans. Archit. Code Optim.3
2023 Delay Mismatch Insensitive Dead Time Generator for High-Voltage Switched-Mode Power Amplifiers
abstract
The Design of efficient, safe, and reliable circuits is a prime objective in high-voltage (HV) electronic systems, such as switched-mode power amplifiers (PAs). One of the main causes of efficiency degradation and reliability problems, in these amplifiers, is the shoot-through current from the HV power supply to the ground. To eliminate such current, a dead time generator (DTG) is used to modify the signals propagating through the high-side and low-side gate drivers by adding a fixed dead time between them. However, any delay mismatch between these gate drivers can reduce the dead time to the point that it becomes negative. In this paper, an HV-DTG architecture is introduced. The architecture mitigates the effects of delay mismatch variations in gate drivers, which can result from parameters mismatch, fabrication process variations, and temperature variations. An HV switched-mode class-D power amplifier is used to illustrate the performance of the DTG. The amplifier is implemented in a low-cost$0.35~\mu m$HV CMOS process. The total area of the PA is$0.5~mm^{2}$, where the DTG covers an area of$0.066~mm^{2}$. A measured system’s efficiency of 95.14% is achieved with the shortest dead time of 10.8 ns, which is 1.38x smaller than the generated dead time in comparable state-of-the-art HV dead time generators.
Ahmed Abuelnasr, Mostafa Amer, Mohamed Ali 0001, Ahmad Hassan 0002, Benoit Gosselin, Ahmed Ragab, Yvon Savaria
IEEE Trans. Circuits Syst. I Regul. Pap.7
2022 A Templated VHDL Architecture for Terabit/s P4-programmable FPGA-based Packet Parsing
abstract
This paper proposes a templated VHDL architecture for P4-programmable packet parsing on FPGAs offering high throughput while occupying a small area footprint. The architecture comprises a multi-stage header parser unit arranged in a pipelined structure. Each header analysis unit is characterized by a set of generic parameters reflecting unique features and relations of supported protocols retrieved from the P4 code that describes each stage along the pipeline. Synthesis results of the packet parser show up to 549 Gb/s throughput on a Xilinx Virtex-7 FPGA and 1 Tb/s on a Xilinx UltraScale+ for a twelve-stage pipeline. Compared with state-of-the-art solutions, our proposed architecture performs at higher throughput with acceptable resource utilization.
Parisa Mashreghi-Moghadam, Tarek Ould Bachir, Yvon Savaria
ISCAS3
2022 Mobile-URSONet: an Embeddable Neural Network for Onboard Spacecraft Pose Estimation
abstract
Spacecraft pose estimation is an essential computer vision application that can improve the autonomy of in-orbit operations. An ESA/Stanford competition brought out solutions that seem hardly compatible with the constraints imposed on spacecraft onboard computers. URSONet is among the best in the competition for its generalization capabilities but at the cost of a tremendous number of parameters and high computational complexity. In this paper, we propose Mobile-URSONet: a spacecraft pose estimation convolutional neural network with 178 times fewer parameters while degrading accuracy by no more than four times compared to URSONet.
Julien Posso, Guy Bois, Yvon Savaria
ISCAS3
2022 An FPGA-based HW/SW Co-Verification Environment for Programmable Network Devices
abstract
Bugs in network devices translate to financial losses for the service providers and degrade the quality of experience for the users. Simulation tools cannot guarantee complete fault coverage as bugs can manifest at any time in live hardware. To mitigate these issues, we propose a novel hardware/software (HW/SW) co-verification tool that targets programmable dataplane network devices. The system integrates cycle-accurate software simulation with a hardware implementation. For the software simulation, open-source tools such as CocoTB and GHDL were used. The Design Under Test (DUT) and our test interfaces are embedded in programmable hardware. Data from the software can be inserted and then extracted in real-time from the input/output (I/O) ports of the DUT. To achieve this functionality the hardware design uses data insertion and extraction blocks which also support assertions. For the hardware implementation, reported experiments have been conducted on a NetFPGA-SUME platform. When a packet flows through the NetFPGA and triggers an assertion, the data present in the DUT at that time can be captured, and sent back to the simulator for further analysis and replay. Each of our design block consumes less than 1% of the available resources on the FPGA.
Mengyue Su, Jean-Pierre David, Yvon Savaria, Bill Pontikakis, Thomas Luinaud
ISCAS3
2022 Persistence Region Monitor With a Pheromone-Inspired Robot Swarm Sensor Network
abstract
In this article, we propose a controller that can coordinate a swarm of robots to cover a region persistently by forming a mobile sensor network. Therefore, the robot swarm can monitor a large area of interest (AOI) over a long time. The performance of the swarm can achieve high flexibility and coverage efficiency with swarm intelligence. This method is inspired by the behavior of large predators, such as lions that use liquid markers containing pheromones to mark their territories and indicate their status. Via interactions based on these clues, the ecosystem can achieve a dynamic balance. The controller inspired by this phenomenon consists of two layers, which are a region divider and a path planner. The region divider evenly splits the area into random shapes according to a pheromone clue and adapts to dynamic changes of the swarm size. Then, the path planner generates a closed patrol path for each robot. Like a random search scheme or an ant pheromone-inspired scheme, the proposed controller can adapt to the dynamic change of swarm size providing high efficiency of region coverage. The effectiveness of this controller is proved by modeling the life cycle of a robot swarm as a finite state machine. It is also verified with simulations and experiments with unmanned ground vehicles (UGVs).
Yuzhan Wu, Meng Li 0003, Yvon Savaria
IEEE Internet Things J.4
2022 A High-Sensitivity Wide Input-Power-Range Ultra-Low-Power RF Energy Harvester for IoT Applications
abstract
Radio frequency energy harvesting (RFEH) is very attractive for the Internet of things (IoT) and self-powered micro-systems such as wearable biomedical devices and wireless sensor networks. This paper proposes, analyzes, and implements a new RF-DC converter in standard 130 nm CMOS technology. The developed converter is designed and optimized for ultra-low-power IoT and wearable biomedical applications using the 900 MHz ISM band. The proposed 10-stage cross-connected rectifier compensates the transistors threshold voltage by using both dynamic and static bias compensation techniques. An analytical model of the rectifier based on the MOSFET transistor equations is presented, allowing optimization of the rectifier as a function of the number of stages and transistors sizing, improving the sensitivity and the input power range of the converter. The measurement results demonstrate a sensitivity of −25.5 dBm for 1 V output across a 5-$\text{M}\Omega $resistive load and −29 dBm for a 100$\text{M}\Omega $load, which is better than the best previously reported results. The measured peak end-to-end efficiency of the proposed harvester is 42.4% at −16 dBm input power, delivering 2.19 V to a 450$\text{k}\Omega $load.
Seyed Mohammad Noghabaei, Rafael L. Radin, Yvon Savaria, Mohamad Sawan
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Design Principles for Packet Deparsers on FPGAs
abstract
The P4 language has drastically changed the networking field as it allows to quickly describe and implement new networking applications. Although a large variety of applications can be described with the P4 language, current programmable switch architectures impose significant constraints on P4 programs. To address this shortcoming, FPGAs have been explored as potential targets for P4 applications. P4 applications are described using three abstractions: a packet parser, match-action tables, and a packet deparser, which reassembles the output packet with the result of the match-action tables. While implementations of packet parsers and match-action tables on FPGAs have been widely covered in the literature, no general design principles have been presented for the packet deparser. Indeed, implementing a high-speed and efficient deparser on FPGAs remains an open issue because it requires a large amount of interconnections and the architecture must be tailored to a P4 program. As a result, in several works where a P4 application is implemented on FPGAs, the deparser consumes a significant proportion of chip resources. Hence, in this paper, we address this issue by presenting design principles for efficient and high-speed deparsers on FPGAs. As an artifact, we introduce a tool that generates an efficient vendor-agnostic deparser architecture from a P4 program.Our design has been validated and simulated with a cocotb-based framework.The resulting architecture is implemented on Xilinx Ultrascale+ FPGAs and supports a throughput of more than 200 Gbps while reducing resource usage by almost 10x compared to other solutions.
Thomas Luinaud, Jeferson Santiago da Silva, J. M. Pierre Langlois, Yvon Savaria
FPGA4
2021 Causal Information Prediction for Analog Circuit Design Using Variable Selection Methods Based on Machine Learning
abstract
This paper proposes a methodology based on machine learning to find apparent causal relations between performance targets and design variables in analog circuits. Diversified filtering and wrapping variable selection algorithms are utilized to construct a causal graph that identifies the major circuit design parameters that can be used to optimize the performance of analog circuits. Based on the constructed causal graph, a sequence of design procedures can be extracted and followed to optimize the performance of a design. The proposed methodology is validated using a two-stage op-amp. The obtained causal graph agrees with analytical design equations published in the literature for the selected two-stage op-amp. The results also show that the proposed methodology can accelerate the circuit design process and effectively help designers understand the reasoning behind different design decisions.
Ahmed Abuelnasr, Mostafa Amer, Ahmed Ragab, Benoit Gosselin, Yvon Savaria
ISCAS5
2021 Design and Analysis of Combined Input-Voltage Feedforward and PI Controllers for the Buck Converter
abstract
This paper presents the design and analysis of combining input-voltage feedforward and proportional-integral (PI) controllers to regulate the output voltage of DC-DC Buck converter subject to input line disturbances. Non-idealities of the Buck converter such as passive and active components parasitics are included in the mathematical model obtained by the statespace averaging (SSA) technique for accurate control. The stability boundary locus approach is used to graphically analyze the system stability. It guides the design of the PI controller gains and the feedforward scaling factor to achieve desired phase and gain margins. Analysis shows that the feedforward scaling factor affects the stability regions of the closed-loop system and can limit the possible PI controller gains for certain phase and gain margins; 75oand 9.54 dB in our case. The results are verified by a Simulink model developed for the Buck converter system.
Mostafa Amer, Ahmed Abuelnasr, Ahmed Ragab, Ahmad Hassan 0002, Mohamed Ali 0001, Benoit Gosselin, Mohamad Sawan, Yvon Savaria
ISCAS8
2021 RISC-V Barrel Processor for Deep Neural Network Acceleration
abstract
This paper presents a barrel RISC-V processor designed to control a deep neural network accelerator. Our design has a 5-stage pipeline data path with 8 hardware threads (harts). Each thread is executed under a strict round robin scheduler and is responsible for providing data and control signals to a neural network processing element (PE). Each PE is capable of arbitrary precision GEneral Matrix Vector (GEMV) operations. The execution of each thread is independent of other threads and any communication between threads are sent through shared memory via software. To reduce the area required for implementation, our processor is an implementation of the RV32I plus a set of custom CSRs for controlling the PEs. Our design passes all riscv_test written in assembly and compiled with RISC-V gcc. Our 8-hart barrel processor runs at 250 MHz with CPI of 1 and consumes 0.372W. To demonstrate the capabilities of our design, we computed a GEMV operation with an input matrix size of 8 by 128 and a weight matrix size of 128 by 128 with two-bit precision in only 16 clock cycles.
MohammadHossein AskariHemmat, Olexa Bilaniuk, Sean Wagner, Yvon Savaria, Jean-Pierre David
ISCAS4
2021 Acceleration of the Secure Hash Algorithm-256 (SHA-256) on an FPGA-CPU Cluster Using OpenCL
abstract
The Secure Hash Algorithm-256 (SHA-256) is a cryptographic function used in a wide variety of applications ranging from Internet of Things micro-devices to highperformance systems. This paper studies a set of implementations of the SHA-256 on a field-programmable gate array (FPGA) elaborated using the Open Computing Language (OpenCL). These implementations apply several optimization techniques to improve their respective throughputs. Reported results show that a combination of OpenCL optimization techniques allows obtaining an implementation offering a 90x speed-up when compared to an unoptimized OpenCL implementation. Moreover, the best reported optimized implementation achieves a throughput of 3973 Mbps, which is 4.3 times higher than the best previously published HLS-based SHA-256 implementation and even higher than the previously published implementations using a hardware description language. To our knowledge, this work is the first that proposes an OpenCL-based FPGA implementation of SHA-256 and its OpenCL-based optimization methodology.
Hachem Bensalem, Yves Blaquière, Yvon Savaria
ISCAS3
2021 Power Bound Analysis of a Two-Step MASH Incremental ADC Based on Noise-Shaping SAR ADCs
abstract
Power consumption is an important limitation in designing analog-to-digital converters (ADCs) used in low-power sensing applications. This paper estimates analytically the power bound of a two-step multi-stage noise-shaping successive-approximation-register incremental ADC (two-step MASH NS-SAR IADC) proposed in our previous work. Our model considers the impacts of thermal noise, mismatch, and CMOS process (minimum feature size in CMOS technologies) on the power bounds of the proposed IADC. The analytic results show that thermal noise and CMOS process requirements determine the power consumption lower bounds in high and low resolutions, respectively. A comparison with the most competitive single-loop delta-sigma (ΔΣ) IADC shows a 3-dB higher theoretical figure-of-merit (FoM) for our proposed IADC when the resolutions are higher than 12-bit. Our proposed systematic analysis can be used to estimate the power bounds of amplifier-based NS-SAR ADCs used in either ΔΣ or incremental mode with multi-stage and multi-step topologies designed in various CMOS technologies. The reported analytic results are confirmed by experimental results of previously reported implementations.
Masoume Akbari, Mohammad Honarparvar, Yvon Savaria, Mohamad Sawan
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 RISC-V Barrel Processor for Accelerator Control
abstract
Hardware accelerators are important in the post-Moore’s law era of computing. To maximize performance of such accelerators, most of the logic resources should be allocated to their execution circuits, while control mechanisms should be kept small yet flexible. In this paper, we propose a barrel processor design based on the RISC-V instruction set architecture (ISA) [1]. To the best of our knowledge, this is the first implementation of a barrel RISC-V processor made public. The purpose of this processor is to concurrently control and coordinate a set of accelerator processing elements.
MohammadHossein AskariHemmat, Olexa Bilaniuk, Sean Wagner, Yvon Savaria, Jean-Pierre David
FCCM4
2020 Unleashing the Power of FPGAs as Programmable Switches
abstract
The P4 language and the PISA architecture have revolutionized the field of networking. Thanks to P4 and PISA, new networking applications and protocols can be rapidly evaluated on high performance switches. While P4 allows the expression of a wide range of packet processing algorithms, current programmable switch architecture limit the overall processing flexibility. To address this shortcoming recent work have proposed to implement PISA on FPGAs. However, little effort has been devoted to analyze whether FPGAs are good candidates to implement PISA. In this work, we take a step back and evaluate the micro-architecture efficiency of various PISA blocks. Using a theoretical analysis and experiments, we demonstrate that current FPGA architecture drastically limit the performance of a few PISA blocks. Thus, we explore two avenues to alleviate these shortcomings. First, we identify some network applications that are well tailored to current FPGAs. Second, to support a wider range of networking applications, we propose modifications to the FPGA architecture which can also be of interest outside the networking field.
Thomas Luinaud, Thibaut Stimpfling, Jeferson Santiago da Silva, Yvon Savaria, J. M. Pierre Langlois
FPGA4
2020 Bridging the Gap: FPGAs as Programmable Switches
abstract
The emergence of P4, a domain specific language, coupled to PISA, a domain specific architecture, is revolutionizing the networking field. P4 allows to describe how packets are processed by a programmable data plane, spanning ASICs and CPUs, implementing PISA. Because the processing flexibility can be limited on ASICs, while the CPUs performance for networking tasks lag behind, recent works have proposed to implement PISA on FPGAs. However, little effort has been dedicated to analyze whether FPGAs are good candidates to implement PISA. In this work, we take a step back and evaluate the micro-architecture efficiency of various PISA blocks. We demonstrate, supported by a theoretical and experimental analysis, that the performance of a few PISA blocks is severely limited by the current FPGA architectures. Specifically, we show that match tables and programmable packet schedulers represent the main performance bottlenecks for FPGA-based programmable switches. Thus, we explore two avenues to alleviate these shortcomings. First, we identify network applications well tailored to current FPGAs. Second, to support a wider range of networking applications, we propose modifications to the FPGA architectures which can also be of interest out of the networking field.
Thomas Luinaud, Thibaut Stimpfling, Jeferson Santiago da Silva, Yvon Savaria, J. M. Pierre Langlois
HPSR4
2020 Self-Adjusting Deadtime Generator for High-Efficiency High-Voltage Switched-Mode Power Amplifiers
abstract
In this paper, we propose a novel design methodology for a deadtime generator for high-efficiency power amplifiers. It consists of a two-phase non-overlapping clock circuit and level down shifters. A 3% improvement in efficiency is achieved with a maximum efficiency of 94% in a class-D power amplifier circuit. The proposed design generates a deadtime as low as 16.7ns and eliminates the problem of propagation delay mismatch between high side and low side gate drivers. The circuit is implemented in 50V AMS 0.35 μm CMOS technology. The deadtime generator consumes 16.5 mW, while occupying a total area of 0.068 mm2.
Ahmed Abuelnasr, Mohamed Ali 0001, Mostafa Amer, Morteza Nabavi, Ahmad Hassan 0002, Benoit Gosselin, Yvon Savaria
ISCAS7
2020 OTA-Free MASH 2-2 Noise Shaping SAR ADC: System and Design Considerations
abstract
A multi-stage noise shaping (MASH) analog to digital converter (ADC) architecture is presented in this paper. This architecture combines the features of noise shaping SAR (NS-SAR) with the MASH scheme to achieve a higher order noise shaping. This ADC does not suffer from the complexity issue of conventional MASH delta sigma (ΔΣ) structures, and it does not require operational transconductance amplifier-based analog integrators. It also exhibits a high resolution and moderate bandwidth while using a low oversampling ratio (OSR). These merits make the introduced architecture suitable for large number of applications, such as internet of things (IoT) and biomedical devices. The paper proposes MATLAB behavioral models along with macro models used to simulate the presented architecture to show the efficiency of the proposed ADC featuring a signal-to-quantization-noise ratio (SQNR) of 105 dB for an OSR of 10 at a sampling frequency of 10 MHz.
Masoume Akbari, Mohammad Honarparvar, Yvon Savaria, Mohamad Sawan
ISCAS3
2020 Analog Circuits to Accelerate the Relaxation Process in the Equilibrium Propagation Algorithm
abstract
Equilibrium Propagation (EP) is a novel biologically plausible algorithm for training deep neural networks. However, discrete time implementations of EP could not exploit the full potential of its proposed framework, because the relaxation step is slow on digital architectures. Here, we propose an analog circuit implementation for the relaxation process to accelerate its convergence. In this implementation, the optimization process is executed based on the continuous-time dynamics of EP. Our circuit is validated using a Continuous Hopfield Network prototype emulating the XOR function. This implementation improved the convergence time of EP by a factor of 250 compared to its Python counterpart. As this implementation is scalable to learn more complex functions, it has the potential to be applied on higher-dimensionality datasets.
Armin Najarpour Foroushani, Hussein Assaf, Fereidoon Hashemi Noshahr, Yvon Savaria, Mohamad Sawan
ISCAS4
2020 Indoor Localization Using Channel State Information With Regression Artificial Neural Networks
abstract
In this paper, the Channel State Information (CSI) is used to locate mobile stations in an indoor environment. The novelty of our technique is to use multiple packets of CSI for each location without feature extraction to provide a reach fingerprint. Two different mapping algorithms are investigated and compared with each other in terms of location accuracy and precision. In the first approach, the collected CSIs are fed to a multilayer perceptron (MLP) as input features and the learned artificial neural network (ANN) is used as a pattern-matching algorithm in order to predict a user's location. The second approach uses General Regression Neural Networks (GRNNs) from which, exploration is performed to find the best hidden-layer configuration and spread factors for Multilayer Perceptrons (MLPs) and General Regression Neural Networks (GRNNs), respectively. The novelty of this work partly stems from data expanding multiple CSI packets. The paper finally compares, the accuracy of our proposed method with previously reported state-of-the-art methods.
Seyed Mohsen Samadani, Yvon Savaria, Chahé Nerguizian
VTC Spring2
2020 Optimization of Small-Delay Defects Test Quality by Clock Speed Selection and Proper Masking Based on the Weighted Slack Percentage
abstract
Classical transition delay fault (TDF) model-based delay tests cannot detect small-delay defects (SDDs) properly in many circuits because they do not set the test slack properly to match the size of the tested delay fault. In this article, we propose a new method of using faster-than-at-speed clocks to enhance a classical TDF-based test pattern set and create an optimized SDD test that can detect effective SDDs. We define effective SDDs as SDDs that can immediately fail a circuit and distinguish them from those that grow over time to fail the circuit (reliability defects). The optimized SDD test developed in this article can either target only effective SDDs or also include reliability defects that are close to fail a circuit. In order to improve the SDD test quality, the proposed optimization method uses the recently proposed weighted slack percentage (WeSPer) as the SDD test quality metric to match each pattern with an appropriate test clock speed and then generates the proper masking vector. The optimization method includes a technique for reducing the final pattern count while keeping WeSPer value high. On a set of benchmark circuits, the proposed technique is able to improve the WeSPer value by up to 32.97% (a relative improvement of 60.98%) compared to classical at-speed testing using the same pattern set. This article also discusses the application of this technique on a self-timed circuit that has previously been at-speed tested. The proposed optimization technique improved the SDD test quality from 60% to 69% for preexisting set of test clocks and patterns. Post-silicon test results for the processor are included in this article.
Omar Al-Terkawi Hasib, Yvon Savaria, Claude Thibeault
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Toward In-System Monitoring of OpenCL-Based Designs on FPGA
abstract
This paper presents a new in-system circuit for monitoring and profiling OpenCL-based designs on FPGA. This circuit opens the door for improved monitoring of OpenCL-based FPGA accelerators. The monitoring approach allows designers to identify unexpected performance bottlenecks such as pipeline stalls and initiation interval in OpenCL loops. The proposed monitor enhances observability into OpenCL-based accelerators by capturing high-level hardware events and timing information at FPGA-clock accuracy. Any event or variable in an OpenCL kernel can be observed by instantiating monitor circuits specified in OpenCL. To our knowledge, it is the first FPGA clock-cycle accurate monitor that can select not only the variables, but also the data inputs to be monitored in such variables. To validate the proposed in-system monitoring circuit, the Arria10 FPGA and Intel SDK were used for the OpenCL tool-chain. The reported results show that the presented monitoring circuit introduces a frequency degradation that remains small for 8 tested monitors instances. These monitors use 2.3 times fewer logic resources than previously reported OpenCL monitors.
Hachem Bensalem, Yves Blaquière, Yvon Savaria
ISCAS3
2019 A Prediction Model for Implementing DVS in Single-Rail Bundled-Data Handshake-Free Asynchronous Circuits
abstract
This paper explores the use of dynamic voltage scaling (DVS) for a particular type of asynchronous circuits, namely the single-rail bundled-data handshake-free asynchronous circuits. With respect to their synchronous counterparts, applying DVS to the targeted type of circuits imposes additional timing constraints that must be met to ensure correct operation. A new model defining these additional constraints is proposed. Such DVS related constraints were never explicitly formulated. The proposed model considers setup and hold timing constraints and transition degradation to ensure proper timing closure. Reported simulation results show that the timing constraints can be satisfied while applying DVS, but also that adjustments might be required to ensure proper pulse propagation and efficient operations.
Maryem Benyoussef, Claude Thibeault, Yvon Savaria
ISCAS3
2019 Bit-Slicing FPGA Accelerator for Quantized Neural Networks
abstract
Deep Neural Networks (DNNs) become the state-of-the-art in several domains such as computer vision or speech recognition. However, using DNNs for embedded applications is still strongly limited because of their complexity and the energy required to process large data sets. In this paper, we present the architecture of an accelerator for quantized neural networks and its implementation on a Nallatech 385-A7 board with an Altera Stratix V GX A7 FPGA. The accelerator's design centers around the matrix-vector product as the key primitive, and exploits bit-slicing to extract maximum performance using low-precision arithmetic.
Olexa Bilaniuk, Sean Wagner, Yvon Savaria, Jean-Pierre David
ISCAS3
2019 Multi-PVT-Point Analysis and Comparison of Recent Small-Delay Defect Quality Metrics
Omar Al-Terkawi Hasib, Yvon Savaria, Claude Thibeault
J. Electron. Test.2
2019 SHIP: A Scalable High-Performance IPv6 Lookup Algorithm That Exploits Prefix Characteristics
abstract
Due to the emergence of new network applications, current IP lookup engines must support high bandwidth, low lookup latency, and the ongoing growth of IPv6 networks. However, the existing solutions are not designed to address jointly these three requirements. This paper introduces SHIP, an IPv6 lookup algorithm that exploits prefix characteristics to build a data structure designed to meet future application requirements. Based on the prefix length distribution and prefix density, prefixes are first clustered into groups sharing similar characteristics and then encoded in hybrid trie-trees. The resulting memory-efficient and scalable data structure can be stored in low-latency memories and allows the traversal process to be parallelized and pipelined in order to support high packet bandwidth in hardware. In addition, SHIP supports incremental updates. Evaluated on real and synthetic IPv6 prefix tables, SHIP has a logarithmic scaling factor in terms of the number of memory accesses and a linear memory consumption scaling. Compared with other well-known approaches, SHIP reduces the required amount of memory per prefix by 87%. When implemented on a state-of-the-art field-programmable gate array (FPGA), the proposed architecture can support processing 588 million packets per second.
Thibaut Stimpfling, Normand Bélanger, J. M. Pierre Langlois, Yvon Savaria
IEEE/ACM Trans. Netw.4
2019 A Defect-Tolerant Reusable Network of DACs for Wafer-Scale Integration
abstract
A novel defect-tolerant network of digital-to-analog converters (DACs) is presented in this paper. The architecture of this converter employs a single 2.5-V voltage reference and an unbalanced buffering technique to achieve a wide voltage range that extends from 864 mV to 2.538 V with an 8-bit resolution. The proposed converter incorporates a defect-tolerant architecture and is extremely compact, utilizing a per-bit silicon area of less than 350 μm2. Although such very small area allows for embedding in dense configurable fabrics (field-programmable gate arrays) and wafer-scale integration, the overall performance is not sacrificed as reported measurements show a signal-tonoise ratio of 51.87 dB and a spurious-free dynamic range of 42.31 dB, at 10 MS/s providing 7.6 effective bits. Moreover, the proposed architecture benefits from dynamic calibration capabilities, as any converter output can be finely adjusted over a range of 25 mV. This proposed DAC is also extensively reused in the same defect-tolerant network for a successive approximation register-analog-to-digital converter, as well as for a configurable voltage reference.
Nicolas Laflamme-Mayer, Gilbert Kowarzyk, Yves Blaquière, Yvon Savaria, Mohamad Sawan
IEEE Trans. Very Large Scale Integr. Syst.4
2018 A Low-Latency Memory-Efficient IPv6 Lookup Engine Implemented on FPGA Using High-Level Synthesis
abstract
The emergence of 5G networks and real-time applications across networks has a strong impact on the performance requirements of IP lookup engines. These engines must support not only high-bandwidth but also low-latency lookup operations. This paper presents the hardware architecture of a low-latency IPv6 lookup engine capable of supporting the bandwidth of current Ethernet links. The engine implements the SHIP lookup algorithm, which exploits prefix characteristics to build a compact and scalable data structure. The proposed hardware architecture leverages the characteristics of the data structure to support low-latency lookup operations, while making efficient use of memory. The architecture is described in C++, synthesized with a highlevel synthesis tool, then implemented on a Virtex-7 FPGA. Compared to the proposed IPv6 lookup architecture, other wellknown approaches use at least 87% more memory per prefix, while increasing the lookup latency by a factor of 2.3×.
Thibaut Stimpfling, J. M. Pierre Langlois, Normand Bélanger, Yvon Savaria
CCGrid4
2018 High-Temperature Modeling of the I-V Characteristics of GaN150 HEMT Using Machine Learning Techniques
abstract
We propose in this paper a high-temperature non-linear modeling for the I-V characteristics of GaN150 HEMT. Three different data-driven models were developed for a temperature range varying from 25°C to 250°C, by using three machine learning regression techniques namely: The Artificial Neural Network (ANN), the Support Vector Machine (SVM) and the Decision Tree (DT). Experiments were conducted on a GaN150 device with a width of 40 μm and accordingly, a set of measurements were obtained and exploited to build the device model. The three models were evaluated based on their ability to predict the I-V characteristics outside the temperature range (greater than 250°C) and their mean square error. The obtained results show that the models predict the device characteristics correctly based on the calculated mean squared error between the actual and predicted characteristics.
Ahmed Abubakr, Ahmad Hassan 0002, Ahmed Ragab, Soumaya Yacout, Yvon Savaria, Mohamad Sawan
ISCAS5
2018 High-Temperature Empirical Modeling for the I-V Characteristics of GaN150-Based HEMT
abstract
We describe in this paper a model for the I-V characteristics of AlGaN/GaN high electron mobility transistors (HEMTs) working in high-temperature environments up to 250°C. Modeling of this emerging technology is a very significant step toward incorporating the technology in harsh environment applications. An extended version of the Angelov model is modified in this paper to consider the temperature as a variable. The developed model is fitted to the experimental I-V data using MATLAB. The reported experimental data are in good agreement with the model outputs over the specified temperature range. Moreover, the model was validated using the Spectre circuit simulator.
Mostafa Amer, Ahmad Hassan 0002, Ahmed Ragab, Soumaya Yacout, Yvon Savaria, Mohamad Sawan
ISCAS5
2018 Design of a Low Latency 40 Gb/s Flow-Based Traffic Manager Using High-Level Synthesis
abstract
This paper presents a traffic manager architecture targeting to meet today's networking requirements, especially reduced latency, and to support the upcoming 5G technology in the software defined networking context. The proposed traffic manager functionalities are policing, scheduling, shaping, and queuing of incoming traffic (packets). The incoming traffic is assumed to be a set of flows in a network processing unit. Traffic management imposes constraints on packets to be sent out in such a way to meet the allowed bandwidth quotas for each flow, and enforce desired quality of service (QoS) targets. The FPGA prototyped architecture is based on the C++ language and is synthesized with the Vivado High-Level Synthesis (HLS) tool. The proposed traffic manager design supports 40 Gb/s per egress port for 64-byte sized packets, running at 80 MHz when implemented on a ZC706 Xilinx board. A throughput improvement of 4.0× over previous reported works is claimed.
Imad Benacer, François R. Boyer, Yvon Savaria
ISCAS3
2018 Implementation of a Cache-Based IPv6 Lookup System with Hashing
abstract
Due to the rapid growth of traffic on the Internet, the IP lookup process imposes ever-growing performance requirements in order to avoid that it becomes a bottleneck during packet forwarding. This complex function is often implemented by hardware accelerators that are integrated with a processor. In this paper, we use a modified cache memory as an accelerator to perform IP lookup. Hashing is used for mapping each bucket of a hash table to a set of the cache memory. We show that, in the proposed scheme, a table of 26K prefixes fits into a cache of 1MB and the throughput achieved allows processing packets at wire speed over four 40Gb links.
Bachir Fradj, Benjamin Wolff, Normand Bélanger, Yvon Savaria
ISCAS4
2018 Custom Low Power Processor for Polar Decoding
abstract
Cloud Radio Access Network is foreseen as one of the key features of the future 5G mobile communication standard. In this context, all the baseband processing is intended to be performed on CPUs in order to keep a high level of flexibility. The challenge is then to propose efficient software implementations of baseband processing algorithms that guarantee a sufficient throughput, while limiting the energy consumption. In this paper, as an alternative to general purpose processors, we propose an implementation of an Application Specific Instruction set Processor customized for the Successive Cancellation decoding of polar codes. The resulting software decoder achieves throughputs similar to state-of-the-art ARM processor implementations, while reducing the energy consumption by a factor 10.
Mathieu Léonardon, Camille Leroux, David Binet, J. M. Pierre Langlois, Christophe Jégo, Yvon Savaria
ISCAS6
2018 A High-Efficiency Ultra-Low-Power CMOS Rectifier for RF Energy Harvesting Applications
abstract
This paper presents a novel ultra-low power rectifier for RF energy harvester, designed and implemented in standard 130 nm CMOS technology. The proposed 915 MHz ISM band RF energy harvester is designed for wearable medical devices and internet of things (IoT) applications. An off-chip differential matching network passively boosts the low-level incoming AC signal generated by the antenna. Then, a novel self-compensated cross-coupled rectifier is designed to convert the AC signal into a DC output voltage. The rectifier is comprised of 10 stages and it uses both dynamic and static bias compensation to decrease the transistors forward voltage drop. The post-layout simulation results demonstrate a sensitivity of -30.5 dBm for 1 V output at a capacitive load which is lower than the current state-of-the-art. The peak end-to-end efficiency is 42.8 % at -16 dBm input power, delivering 2.32 V at 0.5 MΩ resistor load.
Seyed Mohammad Noghabaei, Rafael L. Radin, Yvon Savaria, Mohamad Sawan
ISCAS3
2018 Exploiting built-in delay lines for applying launch-on-capture at-speed testing on self-timed circuits
abstract
The application of scan-based at-speed delay testing on asynchronous circuits is not trivial. Their unorthodox design leaves them generally incompatible with traditional synchronous design and test tools, as well as standard automatic test equipment. The correct generation of at-speed test clocks and the use of conventional automatic test patterns generation (ATPG) tools are some of the problems that face the application of at-speed testing on asynchronous circuits. This paper presents a method of applying scan-based at-speed testing on single-rail bundleddata handshake-free (self-timed) asynchronous circuits by taking advantage of built-in delay lines. The proposed test method uses launch-on-capture scan-based testing with endpoint masking and generates the test patterns using conventional ATPG tools. The proposed test is applied on circuits in a self-timed microprocessor fabricated in 28nm FD-SOI CMOS technology. This method is validated by the reported test coverage and simulation results, along with post-silicon test results on a Teradyne FLEX tester.
Omar Al-Terkawi Hasib, Daniel Crepeau, Thomas Awad, Andrei Dulipovici, Yvon Savaria, Claude Thibeault
VTS5
2018 Enhanced Bloom filter utilisation scheme for string matching using a splitting approach
abstract
Bloom filters (BFs) are widely utilised to speed up string matching in crucial network applications such as real‐time intrusion detection and spam filters. This study introduces a new approach to improve the efficiency of BFs for string matching functions. The approach splits each target string into two substrings and considers the second substring for programming the BF. The objective is to minimise the false positive rate by maximising the common hash signatures from the second substring. Results show that compared to the traditional means of using BFs, the proposed approach reduces the false positive rate by averages of 76 and 88% for 32 and 64 Kb BFs, respectively. Moreover, a complete string matching architecture has been developed in hardware based on the proposed approach. Results demonstrate the advantages of this new architecture compared to similar previous works.
Shervin Vakili, J. M. Pierre Langlois, Yvon Savaria, Naraig Manjikian
IET Commun.3
2018 Diagnosis algorithms for a reconfigurable and defect tolerant JTAG scan chain in large area integrated circuits
Safa Berrima, Yves Blaquière, Yvon Savaria
Integr.3
2018 A pattern-based routing algorithm for a novel electronic system prototyping platform
Etienne Lepercq, Yves Blaquière, Yvon Savaria
Integr.3
2018 A Fast, Single-Instruction-Multiple-Data, Scalable Priority Queue
Imad Benacer, François R. Boyer, Yvon Savaria
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Electronics and Packaging Intended for Emerging Harsh Environment Applications: A Review
Ahmad Hassan 0002, Yvon Savaria, Mohamad Sawan
IEEE Trans. Very Large Scale Integr. Syst.2
2017 An FPGA Overlay Architecture for Cost Effective Regular Expression Search (Abstract Only)
Thomas Luinaud, Yvon Savaria, J. M. Pierre Langlois
FPGA2
2017 Analysis of SEU Propagation in Combinational Circuits at RTL Based on Satisfiability Modulo Theories
abstract
The vulnerability of VLSI designs to soft errors grows with technology scaling. In order to allow a cost-effective reliability aware design process, it is critical to assess soft error reliability parameters in early design stages. This paper presents a new methodology to estimate digital circuit vulnerability to soft errors of circuits described at Register Transfer Level (RTL). Single Event Upsets (SEUs) propagation through RTL bit-vector operations is modeled and analyzed based on Satisfiability Modulo Theories (SMT). For instance, the bit-vector reduction operators and arithmetic operators were modeled using SMT to include their fault propagation properties. In order to illustrate the practical utilization of our work, we have analyzed different RTL combinational circuits. Experimental results demonstrate that the proposed framework is on average about 4 times faster than other comparable contemporary techniques. Moreover, it provides more accurate and detailed results of the circuit vulnerability allowing a more efficient applicability of fault tolerance techniques.
Ghaith Kazma, Ghaith Bany Hamad, Otmane Aït Mohamed, Yvon Savaria
ACM Great Lakes Symposium on VLSI4
2017 An FPGA Coarse Grained Intermediate Fabric for Regular Expression Search
abstract
Deep Packet Inspection systems such as Snort and Bro express complex rules with regular expressions. In Snort, the search of a regular expression is performed with a Non-deterministic Finite Automaton (NFA). Traversing an NFA sequentially with a CPU is not deterministic in time, and it can be very time consuming. The sequential traversal of an NFA with a CPU is not deterministic in time consequently it can be time consuming. A fully parallel NFA implemented in hardware can search all rules, but most of the time only a small part is active. Furthermore, a string filter determines the traversal of an NFA. This paper proposes an FPGA Intermediate Fabric that can efficiently search regular expressions. The architecture is configured for a specific NFA based on a partial match of a rule found by the string filter. It can thus support all rules from a set such as Snort, while significantly reduce compute resources and power con-sumption compared to a fully parallel implementation. Multiple parameters can be selected to find the best tradeoff between resource consumption and the number and types of supported expressions. This architecture was implemented on a Xilinx R XC7VX1140 Virtex-7. The reported implementation, can sustain up to 512 regular expressions, while requiring 2% of the slices and 16% of the BRAM resources, for a throughput of 200 million characters per second.
Thomas Luinaud, Yvon Savaria, J. M. Pierre Langlois
ACM Great Lakes Symposium on VLSI2
2017 Comprehensive analysis of sequential circuits vulnerability to transient faults using SMT
abstract
Ultra-deep sub-micron technologies are more vulnerable to different types of uncertainties. In this paper, we introduce a novel methodology to estimate the vulnerability of sequential circuits to soft errors at gate level. A new probabilistic modeling of SET propagation is proposed, which reduces the complexity of unrolling sequential circuits. This approach enables a multi-cycle error propagation analysis of sequential circuits using only two copies of the circuit combinational part. The proposed probabilistic modeling is based on the proposed backward unrolling approach in conjunction with the proposed formulation of SET propagation into a Satisfability problem by utilizing satisfability modulo theories. Useful information about the SET latency in sequential circuits and the minimum unrolling required to observe the actual behavior of the circuit is generated. These results are then used to estimate the circuit soft error rate. Experimental results demonstrate the effectiveness and applicability of the proposed approach.
Ghaith Bany Hamad, Ghaith Kazma, Otmane Aït Mohamed, Yvon Savaria
IOLTS4
2017 A multi-measurements RO-TDC implemented in a Xilinx field programmable gate array
abstract
In this paper, an area efficient time to digital converter (TDC) performing measurements between multiple hit signals is proposed. Our TDC is based on a delay line configured as a ring oscillator and a round tracker to count the number of iterations through the oscillator. Lookup tables configured as distributed RAMs and shift registers are used to sample the oscillator and the round tracker states whenever a transition on a signal occurs. A theoretical study is elaborated to estimate FPGA resources required to implement the proposed TDC in comparison with a multichannel basic RO-TDC. It is shown that the gain in the number of Flip-Flops and Lookup tables can reach factors of 85 and 2.1 respectively in an architecture made of a six-stage oscillator, a 32-state round tracker and 20 input hit signals. Temporal characteristics extracted from our TDC implemented in a Xilinx ZYNQ family FPGA are reported.
Safa Berrima, Yves Blaquière, Yvon Savaria
ISCAS3
2017 Scalable memory-less architecture for string matching with FPGAs
abstract
String matching hardware engines generally utilize Ternary Content Addressable Memories (TCAMs). Although TCAM-based solutions are fast, they are expensive and power hungry. This paper proposes a high-performance memory-less architecture for string matching called Split-Bucket. It offers a performance comparable to TCAM-based solutions. Moreover, it is reconfigurable and scalable to the size of the target string set and the width of the string. The architecture is characterized using the Longest Prefix Match problem for IP address lookup and is implemented on a Virtex-7 FPGA. For a real-world routing table with 524 k IPv4 prefixes, the Split-Bucket architecture achieves a throughput of 103.4 M packets per second and consumes 23% and 22% of the Look Up Tables and Flip-Flops of a Xilinx XC7V2000T chip, respectively.
Ideh Sarbishei, Shervin Vakili, J. M. Pierre Langlois, Yvon Savaria
ISCAS4
2017 A Cache-Coherent Heterogeneous Architecture for Low Latency Real Time Applications
abstract
This paper proposes a generic hardware architecture for runtime acceleration of heterogeneous high performance computing (HPC) clusters. This runtime accelerator performs real time resource allocation and management of HPC systems with low latency on multiple time scales. One of the target applications is to perform the signal processing in wireless communication systems such as LTE and 5G over the cloud. A core part of this work is to develop and characterize algorithms that can distribute workloads to server blades in a balanced manner with the aim of maximizing processor utilization in computing clusters. Resources are also managed to guarantee bandwidth for data transfer between computing nodes and reserved cache memories to enable deterministic task execution. This paper shows how a workload distributed among several server blades can be scheduled at a finer time scale than what a normal software implementation would allow, in order to minimize the makespan required to complete execution of sets of tasks. A case study is conducted on the implementation of a resource allocator for the proposed platform. A 760-time acceleration factor of the resource allocation process has been achieved compared to a pure software implementation, while enabling data transfers at the nanosecond scale. It stands as a proof of concept that confirms the viability of CPU-FPGA platforms for wireless standards virtualization.
Michel Gemieux, Yvon Savaria, Jean-Pierre David, Guchuan Zhu
ISORC2
2017 Extensions to decision-tree based packet classification algorithms to address new classification paradigms
Thibaut Stimpfling, Normand Bélanger, Omar Cherkaoui, André Béliveau, Ludovic Béliveau, Yvon Savaria
Comput. Networks6
2017 Formal Methods Based Synthesis of Single Event Transient Tolerant Combinational Circuits
Ghaith Bany Hamad, Otmane Aït Mohamed, Yvon Savaria
J. Electron. Test.3
2017 Reliability Enhancement of Redundancy Management in AFDX Networks
abstract
Avionics Full Duplex Switched Ethernet is a safety critical network in which a redundancy management mechanism is employed to enhance the reliability of the network. However, as stated in the ARINC664-P7 standard, there still exists a potential problem, which may fail redundant transmissions due to sequence inversion in the redundant channels. In this paper, we explore this phenomenon and provide its mathematical analysis. It is revealed that the variable jitter and the transmission latency difference between two successive frames are the two main sources of sequence inversion. Thus, two methods are proposed and investigated to mitigate the effects of jitter pessimism, which can eliminate the potential risk. A case study is carried out and the obtained results confirm the validity and applicability of the developed approaches.
Meng Li 0003, Guchuan Zhu, Yvon Savaria, Michaël Lauer
IEEE Trans. Ind. Informatics3
2016 Memory-Efficient String Matching for Intrusion Detection Systems using a High-Precision Pattern Grouping Algorithm
abstract
The increasing complexity of cyber-attacks necessitates the design of more efficient hardware architectures for real-time Intrusion Detection Systems (IDSs). String matching is the main performance-demanding component of an IDS. An effective technique to design high-performance string matching engines is to partition the target set of strings into multiple subgroups and to use a parallel string matching hardware unit for each subgroup. This paper introduces a novel pattern grouping algorithm for heterogeneous bit-split string matching architectures. The proposed algorithm presents a reliable method to estimate the correlation between strings. The correlation factors are then used to find a preferred group for each string in a seed growing approach. Experimental results demonstrate that the proposed algorithm achieves an average of 41% reduction in memory consumption compared to the best existing approach found in the literature, while offering orders of magnitude faster execution time compared to an exhaustive search.
Shervin Vakili, J. M. Pierre Langlois, Bochra Boughzala, Yvon Savaria
ANCS4
2016 Efficient probabilistic fault tree analysis of safety critical systems via probabilistic model checking
abstract
The cost and complexity involved in the development of critical systems encourage the use of reliability assessment techniques as early in the design cycle as possible. Existing techniques often lack the capacity to perform a comprehensive and exhaustive analysis on complex redundant architectures, leading to less than optimal risk evaluation. This paper addresses these weaknesses by 1) proposing a new probabilistic modeling of Fault Tree gates and their composition as Markov Decision Processes; 2) developing a new formal-based technique to perform an in-depth verification of the system's reliability. This technique makes use of the expressiveness of fault trees and the power of probabilistic model checking in order to investigate the best Triple Modular Redundancy partitioning and configuration of a system. The presented approach greatly improves the overall scalability with respect to other techniques, while also improving the accuracy of the results. For example, we can provide probabilistic failure rates for a chain of 100 redundant components in little over one second.
Marwan Ammar, Ghaith Bany Hamad, Otmane Aït Mohamed, Yvon Savaria
FDL4
2016 Comprehensive non-functional analysis of combinational circuits vulnerability to single event transients
abstract
The progressive shrinking of device sizes in advanced technologies leads to miniaturization and performance improvements. However, ultra-deep sub-micron technologies are more vulnerable to different types of uncertainties, parametric variations, and interference. In this paper, we propose a methodology to model and analyze the behavior of a system in the presence of Single Event Transients (SETs). The problem of SET propagation was modeled as a satisfiability problem using different satisfiability modulo theories. The SET width and timing constraints are formulated as a difference logic constraint satisfaction formulation. This formulation utilizes concepts from static timing analysis to efficiently evaluate the required time and width for the SET to be latched. Next, the proposed model is analyzed using efficient SMT solvers for a set of nonfunctional assertions to investigate SETs propagation. Based on the results of this analysis, new fault observability estimates are computed. These values are then used to compute the soft error rate. Experimental results demonstrate that the proposed SMT approach provides better runtime then contemporary techniques.
Ghaith Bany Hamad, Ghaith Kazma, Otmane Aït Mohamed, Yvon Savaria
FDL4
2016 Efficient and accurate analysis of single event transients propagation using SMT-based techniques
abstract
This paper presents a hierarchical framework to model, analyze, and estimate digital design vulnerability to soft errors due to Single Event Transients (SETs). A new SET propagation model is proposed. This model simultaneously includes the impact of masking effects, width variation, and re-converging paths by utilizing satisfiability modulo theories. Furthermore, new metrics characterizing the soft error rate of a given design are proposed. Reported results show that the proposed methodology significantly enhances the efficiency of SET analysis in terms of: 1) accuracy as it gives accurate estimates of SET sensitivity based on gates timing extracted from layout. These results provide new insights to combinational designs vulnerability to SETs; 2) speed as it is orders of magnitude faster than contemporary techniques; 3) scalability as it can handle large and complex designs such as 128-bit multipliers, whereas contemporary techniques are unable to handle multipliers larger than 32 bits.
Ghaith Bany Hamad, Ghaith Kazma, Otmane Aït Mohamed, Yvon Savaria
ICCAD4
2016 A practical design method for prototyping self-timed processors using FPGAs
abstract
This paper describes a practical design method for protoyping self-timed processors using FPGAs. It adresses shortcomings of typical implementation strategies that use floorplanning to compensate for the lack of support of asynchronous designs in conventional FPGA tools. It is shown that a reported self-timed design technique can be used with a proposed set of constraints to make it fully compliant with standard timing analysis engines. This results in a more effective implementation strategy that makes FPGAs convenient for verifiying self-timed processor designs. The design technique and the timing-driven implementation method are validated by protoyping an 8-bit processor that achieves 11.1 MIPS performance while computing the Fibonacci sequence.
Mickaël Fiorentino, Yvon Savaria, Claude Thibeault, Pascal Gervais
ISCAS2
2016 Towards formal abstraction, modeling, and analysis of Single Event Transients at RTL
abstract
Soft errors due to Single Event Transients (SETs) have become one of the most challenging issues that impact the reliability of modern microelectronic systems at terrestrial altitudes. This is mainly due to the progressive shrinking of device sizes. Traditionally, the analysis of SETs has been carried out by simulations and experimental analysis. However, these techniques are resource hungry and require full details of the design structure and SET characteristics. This paper develops a hierarchical framework for formal analysis of SET propagation by (1) introducing Register Transfer Level (RTL) abstraction and modeling approaches of the underlying behavior of SET propagation using Multiway Decision Graphs (MDGs); and (2) investigating SET propagation conditions at RTL using a formal model checker. In order to illustrate the practical utilization of our work, e have analyzed different RTL combinational designs. Experimental results demonstrate the proposed framework is orders of magnitude faster than other comparable contemporary techniques. Moreover, for the first time, a decision graph based technique s developed to analyze multiplier designs.
Ghaith Bany Hamad, Otmane Aït Mohamed, Yvon Savaria
ISCAS3
2016 Wireless power transfer through metallic barriers enclosing a harsh environment; feasibility and preliminary results
abstract
Modern sensor networks are evolving toward wireless interfaces for both power and data transmission. Design of these devices is challenging, especially when both power and data transmission must reach a harsh environment subject to high temperature, high pressure and through thick metallic layers. This paper reports the design and simulation of an inductive power transfer (IPT) system, which characterizes achievable link efficiency. The problem formulation comes from a family of aerospace applications in which electronic must operate at high temperature (500°C) and high pressure (100 bar) and in which a metal casing (15mm thick) isolates the harsh environment from the environment. Specifically, the requirements of booster rockets in aerospace industry are considered. The proposed IPT system was modeled and simulate d using COMSOL. Reported results demonstrate the feasibility of wirelessly transferring power to a high temperature and high pressure zone through various metallic layers using custom electromagnetic link configurations. Power transfer efficiencies of 38% and 15.5% are reported with Titanium and Steel metal booster interfaces at resonance frequencies of 450Hz and 220Hz respectively.
Ahmad Hassan 0002, Aref Trigui, Umar Shafique, Yvon Savaria, Mohamad Sawan
ISCAS4
2016 A compact spatially configurable differential input stage for a field programmable interconnection network
abstract
This paper presents a spatially configurable input stage for differential-to-single ended conversion enabling signal propagation through single ended field programmable interconnection networks. The input stage uses current mode sensing for differential-to-single ended conversion. Compared to voltage mode sensing, it can support higher common-mode input voltage. Post-layout Monte-Carlo simulations show that the input stage can support data rates of up to 2 Gbps and 1 Gbps for an input common-mode voltage of 1.2-1.6 V and 1.2-2.0 V respectively. The input stage was laid out in a mature 0.13 μm CMOS technology and reported results demonstrate that the occupied silicon area is 22 times smaller than that required by a differential input stage based on unity-gain buffer multiplexers.
Wasim Hussain, Yvon Savaria, Yves Blaquière
ISCAS2
2016 Towards efficient and concurrent FFTs implementation on Intel Xeon/MIC clusters for LTE and HPC
abstract
Fast Fourier Transform (FFT) is an important part of many applications, such as in wireless communication based on OFDM (Orthogonal Frequency Division Multiplexing). With Cloud Radio Access Networks, implementing FFTs on multiprocessor clusters is a challenging task. For instance, supporting the Long Term Evolution (LTE) protocol requires processing 100 independent FFTs (with sizes ranging from 128 to 2048 points) in 66.7 μs. In this work, seven native FFT candidate implementations are compared. The considered implementation environments are: OpenMP (Open Multi-Processing) on 1 core, MPI (Message Passing Interface) on 1 core, 2 cores, and 3 cores, Hybrid OpenMP+MPI on 1 core and 3 cores, and MPI on an heterogeneous platform composed of Xeon-Phi and 3 cores. The reported experimental results show that the latter method meets the latency requirements of LTE. It is shown that the OpenMP and MPI paradigms running only on MICs (Many Integrated Cores) cannot benefit fully from the computing capability of many-core architectures. The heterogeneous combination of Xeon+MICs provides a better performance.
Mounir Khelifi, Daniel Massicotte, Yvon Savaria
ISCAS3
2016 WeSPer: A flexible small delay defect quality metric
abstract
Testing for small delay defects (SDDs) is important due to their dominance in recent technology nodes. Unfortunately, all the SDD test quality metrics in the literature limit their assessment to the size of the delay defect tested under at-speed or slower clocks, which makes their results misleading under special cases such as faster-than-at-speed testing. Moreover, those metrics are inadequate for assessing the quality of recent SDD test methods that consider the variation of delays in a circuit. In this paper, a novel flexible SDD quality metric that can be adapted according to the available information and the applied test method is proposed. The proposed metric is named Weighted Slack Percentage (WeSPer) as it is defined by a slack ratio weighted by confidence level (CL) multipliers. The flexibility comes from the ability to model test inaccuracies or delay varying effects into the CL multipliers. This paper presents the WeSPer metric, along with a CL multiplier that penalizes overtesting to give a more accurate assessment of the quality of faster-than-at-speed testing. The metric is calculated for several benchmark circuits and compared to other SDD metrics found in the literature. The results show that WeSPer is better than other metrics at representing the quality of SDD tests, especially under faster-than-at-speed testing.
Omar Al-Terkawi Hasib, Yvon Savaria, Claude Thibeault
VTS2
2016 A novel spatially configurable differential interface for an electronic system prototyping platform
Wasim Hussain, Olivier Valorge, Yves Blaquière, Yvon Savaria
Integr.4
2016 Monitoring Thermal Stress in Wafer-Scale Integrated Circuits by the Attentive Vision Method Using an Infrared Camera
abstract
This paper is dedicated to the development of a thermal monitoring system for microelectronics based on the attentive vision approach as applied to image sequence analysis using an infrared camera. The attentive vision method implements multiscale image sequence analysis by a spatiotemporal attention operator to detect feature points, which are located inside potential thermal stress regions. The attention operator is a linear aggregation of temporal change and spatial saliency filters. The monitoring process is organized in two hierarchical phases: 1) peripheral and 2) focal. The focal monitoring is mostly carried out through the tracking of stress-relevant feature-point areas and analysis of their spatiotemporal descriptors. The thermal monitoring experiments conducted with wafer-scale integrated circuits have confirmed the reliability of the proposed approach and showed its high potential in image sequence analysis for monitoring purposes.
Ahmed Lakhssassi, Roman Palenychka, Yvon Savaria, Michel Sayde, Marek B. Zaremba
IEEE Trans. Circuits Syst. Video Technol.3
2015 Towards an accurate reliability, availability and maintainability analysis approach for satellite systems based on probabilistic model checking
Khaza Anuarul Hoque, Otmane Aït Mohamed, Yvon Savaria
DATE3
2015 Efficient multilevel formal analysis and estimation of design vulnerability to Single Event Transients
abstract
The progressive shrinking of device size in advanced technologies leads to miniaturization and performance improvements. However, ultra-deep sub-micron technologies are more vulnerable to soft errors. Error analysis of a complex system with a sufficiently large sample of vulnerable nodes takes a large amount of time. In this paper we propose RASVAS, a hierarchical statistical method to model, analyze, and estimate the behavior of a system in the presence of Single Event Transients (SETs) modeled at different abstraction levels. Gate level propagation tables are developed to abstract SET propagation conditions and probabilities from gate level models. At RTL, these tables are utilized to model the underlying probabilistic behavior as Markov Decision Process (MDP) models. Experimental results demonstrate that RASVAS is orders of magnitude faster than contemporary techniques and also handle designs as large as 256-bit adders while maintaining accuracy.
Ghaith Bany Hamad, Otmane Aït Mohamed, Yvon Savaria
IOLTS3
2015 Defect diagnosis algorithms for a field programmable interconnect network embedded in a Very Large Area Integrated Circuit
abstract
Algorithms are proposed to diagnose defects in a defect tolerant field programmable interconnection network embedded in a large area integrated circuit. The proposed diagnosis algorithms use a diagonal configuration approach to reduce the cone of influence of individual tests, thus allowing parallel tests according to diagonal patterns. The proposed algorithms avoid redundant diagnosis tests. Efficiency of the proposed diagnosis algorithms are calculated in terms of the number of cycles of a JTAG FSM required to apply the test. Results show a 113-fold test time reduction in the considered interconnection network.
Gontran Sion, Yves Blaquière, Yvon Savaria
IOLTS3
2015 Modeling the faulty behaviour of digital designs using a feed forward neural network approach
abstract
Cosmic rays lead to soft errors and faulty behavior in electronic circuits. Knowing about their faulty behavior before fabrication would be helpful. This research proposes an approach for modeling the faulty behaviour of digital circuits. It could be applied in a design flow before circuit fabrication. This is achieved by extracting information about faulty behaviour of circuits from low-level models expressed in the VHDL language. Afterwards the extracted information is used to train high-level artificial neural networks models expressed in C/C++ or MATLABTMlanguages. The trained neural network models are able to replicate the behaviour of circuits in presence of faults. The methodology is based on experiments done with two benchmarks, the ISCAS-C17 and a 4-bit multiplier. Results show that the neural network approach leads to models that are more accurate than a previously reported signature generation method. For the C17, using only 30% of the dataset generated with the LIFTING fault simulator, the neural network is able to replicate the output of the circuit in presence of faults with a mean absolute modeling error below 6%.
Zeynab Mirzadeh, Jean-François Boland, Yvon Savaria
ISCAS3
2015 Analysis and characterization of data energy tradeoffs: For VLSI architectural agility in C-RAN platforms
abstract
We investigate trade-offs between traffic adaptation and VLSI architecture adaptation in C-RAN platforms. We propose a dynamic architectural scaling technique applied to interactive applications that require FFT computations. Our implementation results suggest Datapath scaling benefits applications with up to 4.89x improvement in GOPS/mW, while Fabric scaling can benefit Cloud Computing applications with up to 181.97x improvement in GOPS/mW, when compared to published methods. Improvements in network performance and energy-efficiency was achieved at a cost of built-in flexibility in the proposed VLSI architectures.
Pascal Nsame, Guy Bois, Yvon Savaria
ISCAS3
2014 Neuromuscular Representation and Synthetic Generation of Handwritten Whiteboard Notes
abstract
A fully automatic framework has been introduced recently for neuromuscular representation of complex handwriting patterns, such as gestures, signatures, and words, based on the Kinematic Theory of rapid human movements and its Sigma-Lognormal model. In this paper, we investigate the application of this framework to unconstrained whiteboard notes, taking into account a novel acquisition modality, multiple writers, natural language, and complete text lines. Although these conditions deviate strongly from the previously considered scenario of brief pen movements on tablet computers, we demonstrate that the Sigma-Lognormal model is still able to represent the handwriting accurately. In order to deal with longer handwriting patterns, we propose a robust component-wise representation of text lines that achieves a high model quality. Furthermore, we propose a stroke-wise distortion method to generate synthetic text lines from the Sigma-Lognormal representation of real specimens. For handwriting recognition on the IAM online database, it is demonstrated that the extension of the training set with the proposed synthesis method significantly increases current benchmark results achieved with recurrent neural networks.
Andreas Fischer 0002, Réjean Plamondon, Christian O'Reilly, Yvon Savaria
ICFHR4
2014 Design and validation of a novel reconfigurable and defect tolerant JTAG scan chain
abstract
In this paper, a novel technique to get a defect tolerant JTAG compliant scan chain in very large area integrated circuits (VLAIC) is presented. It was ruled that wafer-scale VLAICs require structural regularity and defect-tolerance to be cost effective. Using only one scan chain, as typically used in PCBs, would make the whole VLAIC unusable if a single defect is present in the chain. The proposed technique regularly distributes JTAG Test Access Port (TAP) controllers with test data ports linked to two or more neighbor test data ports. One TAP controller is wired as the entry point and another as the exit point of the scan chain that must be configured according to defect locations. An externally controlled wormhole like routing algorithm can be used for functional link discovery. This paper also proposes a mechanism to make defect tolerant access to test data registers, controlled from neighbor TAP controllers. Our technique has been successfully implemented and validated in a wafer-scale like integrated circuit used in a platform for electronic system prototyping. The logic area of this defect-tolerant configurable JTAG scan chain technique occupies 5% of the test logic and 0.3 % of the cell logic when links to four nearest neighbors are included.
Yves Blaquière, Yan Basile-Bellavance, Safa Berrima, Yvon Savaria
ISCAS4
2014 Abstracting Single Event Transient characteristics variations due to input patterns and fan-out
abstract
Due to shrinking feature sizes and significant reduction in noise margins, as CMOS technologies evolve toward ultra-deep sub-micron, digital circuits have become more susceptible to soft errors. Therefore, researchers have recently reported several approaches to model Single Event Transient (SET) propagation at gate or higher abstraction levels. However, contemporary techniques model only the possibility that SET pulse may be masked electrically, logically, or by time windowing. In this paper, the propagation induced pulse broadening (PIPB) phenomenon is further investigated and a new model which abstracts this phenomenon is proposed. This paper also investigates and abstracts the impact of input patterns and propagation paths on SET pulse width. Through electrical simulations, we validated our analysis.
Ghaith Bany Hamad, Syed Rafay Hasan, Otmane Aït Mohamed, Yvon Savaria
ISCAS4
2014 Probabilistic model checking based DAL analysis to optimize a combined TMR-blind-scrubbing mitigation technique for FPGA-based aerospace applications
abstract
SRAM-based FPGAs are increasingly popular in the aerospace industry for their field programmability and low cost. However, they suffer from cosmic radiation induced Single Event Upsets (SEUs), commonly known as soft errors. In safety-critical applications, the dependability of the design is a prime concern since failures may have catastrophic consequences. An early analysis of dependability of such safety-critical applications will enable designers to develop a design that meets the high availability and reliability requirements of the DO-254 standard. This paper introduces a novel methodology based on probabilistic model checking, to analyze the dependability properties of safety-critical systems and to suggest required mitigation techniques, such as Triple Modular Redundancy (TMR) or TMR with less frequent scrubs for early design decisions. Starting from a high-level description of a system, a Markov model is constructed from the Control Data Flow Graph (CDFG) expressing the functionality and from failure/mitigation parameters for the targeted FPGAs. Such an exhaustive model captures all the failures and repairs possible in the system within the radiation environment. We present a case study on a benchmark circuit to illustrate the applicability of the proposed approach to demonstrate that a wide range of useful dependability properties can be analyzed using our proposed methodology.
Khaza Anuarul Hoque, Otmane Aït Mohamed, Yvon Savaria, Claude Thibeault
MEMOCODE3
2014 A computationally efficient importance sampling tracking algorithm
Rana Farah, Qifeng Gan, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria
Mach. Vis. Appl.5
2014 Determinism Enhancement of AFDX Networks via Frame Insertion and Sub-Virtual Link Aggregation
abstract
Avionics Full Duplex Switched Ethernet (AFDX) is a standard proposed to implement deterministic networks by providing predictable performance guarantees. The determinism is enforced through the concept of Virtual Link, which defines a logical unidirectional connection between end systems. Although an upper bounded end-to-end delay can be obtained using analysis based on, e.g., network calculus, frame arrival uncertainty in destination End-System is a source of nondeterminism that introduces a problem with respect to real-time fault detection. In this paper, a mechanism based on frame insertion is proposed to enhance the determinism of frame arrival within AFDX networks. In order to mitigate network load increase due to frame insertion, a Sub-Virtual Link aggregation strategy, formulated as a multiobjective optimization problem, is introduced. In addition, a brute force algorithm, a greedy algorithm, and a greedy algorithm with preprocessing have been developed to find solutions to the optimization problem. Experiments are carried out and the obtained results confirm the validity and applicability of the developed approaches.
Meng Li 0003, Michaël Lauer, Guchuan Zhu, Yvon Savaria
IEEE Trans. Ind. Informatics4
2014 Optimizing the Parallel Tree-Search for Finding Shortest-Span Error-Correcting CDO Codes
abstract
Finding optimal/short-span Convolutional Self-Doubly Orthogonal (CDO) codes and Simplified-CDO (S-CDO) codes for a specified order J is computationally very challenging. This paper describes several optimizations that were applied to an implicitly-exhaustive search algorithm in order to reduce the time required for finding these types of codes. The resulting high-performance parallel implementation provides an impressive speedup that is greater than 16 300 (CDO,${\rm J} = 7$) and 6300 (S-CDO,${\rm J} = 8$) over the reference implicitly-exhaustive search algorithm, and greater than 2000$({\rm J} = 17)$over the fastest published CDO validation function used in high-performance pseudorandom search algorithms. These speedups are achieved through enhancements in the deterministic search-space reduction, and a vastly improved validation function that makes use of a novel data structure for enabling data-reuse and incremental computations. The resulting validation function speedup is greater than 60 000 (S-CDO,${\rm J} = 17$) and 190 000 (CDO,${\rm J} = 17$) when compared to its reference implementation. The combination of optimizations and load-balancing techniques allowed us to leverage hundreds of processor cores in order to complete an exhaustive search over a search space that is some$10^{14}$times larger than what was previously possible.
Gilbert Kowarzyk, Normand Bélanger, David Haccoun, Yvon Savaria
IEEE Trans. Parallel Distributed Syst.4
2013 An interface for the I2C protocol in the WaferBoard™
abstract
This paper presents a circuit proposed for the DreamWaferTMtechnology. This circuit can interconnect several pads, also called NanoPads, in such a way that they can imitate the behavior of a single metal line for open-drain (or open-collector) buses compliant to the I2C protocol. Thus, multiple serial data lines (SDA) and serial clock lines (SCL) from different user ICs can be connected together on the WaferboardTM. The interface can support up to 25 I2C IC pins together. It can support bidirectional data transfers at up to 100 kbit/s in the Standard-mode, up to 400 kbit/s in the Fast-mode, up to 1 Mbit/s in the Fast-mode Plus, or up to 3.4 Mbit/s in the High-speed mode. The entire interface would take less than 1% of the total area of the WaferICTM, the target system environment for which this circuit is proposed.
Wasim Hussain, Yvon Savaria, Yves Blaquière
ISCAS2
2013 A Library-Based Early Soft Error Sensitivity Analysis Technique for SRAM-Based FPGA Design
Claude Thibeault, Yassine Hariri, Syed Rafay Hasan, Christelle Hobeika, Yvon Savaria, Yves Audet, Fatima Zahra Tazi
J. Electron. Test.5
2013 Efficient Parallel Search Algorithm for Determining Optimal R=1/2 Systematic Convolutional Self-Doubly Orthogonal Codes
abstract
A novel parallel and implicitly-exhaustive search algorithm for finding, in systematic form, rate R=1/2 optimal-span Convolutional Self-Doubly Orthogonal (CDO) codes and Simplified Convolutional Self-Doubly Orthogonal (S-CDO) codes is presented. In order to obtain high-performance low-latency codecs with these codes, it is important to minimize their constraint length (or "span") for a given J number of generator connections. The proposed exhaustive algorithm uses algorithmic enhancements over the best previously published searching techniques, yielding new and improved codes: we were able to obtain new optimal-span CDO/S-CDO codes (having order J∈{9} and J∈{10,11} respectively), as well as new codes having the shortest spans published to date for higher values of J (J∈{10,12,...,17} and J∈{12,...,20} for CDO and S-CDO codes respectively). The new codes and their error performance are provided. An analysis of the evolution of the CDO/S-CDO code error performance as J increases is presented, and the shortest CDO/S-CDO code span values for each given J are compared.
Gilbert Kowarzyk, Normand Bélanger, David Haccoun, Yvon Savaria
IEEE Trans. Commun.4
2012 A novel hybrid FIFO asynchronous clock domain crossing interfacing method
abstract
Multi-clock domain circuits with Clock Domain Crossing (CDC) interfaces are emerging as an alternative to circuits with a global clock. CDC interfaces are susceptible to metastability, hence their design is very challenging. This paper presents a hybrid FIFO-asynchronous method for constructing robust CDC interfaces. The proposed design can handle arbitrary clock frequency ratios between the sender and receiver with random phase shifts. The proposed design avoids latency due to synchronizers with the asynchronous protocol modifications. Circuit simulation results confirm the operation and robustness of the design at maximum workloads, and arbitrary frequency ratios, over a temperature range of -50 to 50 degrees Celsius. The interface offers a maximum throughput of 606 million transfers per second without pausing the clock.
Zaid Al-bayati, Otmane Aït Mohamed, Syed Rafay Hasan, Yvon Savaria
ACM Great Lakes Symposium on VLSI4
2012 Two-level configuration for FPGA: A new design methodology based on a computing fabric
abstract
Large FPGAs require more and more time and expertise to efficiently target custom applications. This paper presents a new methodology based on two configuration levels. At the lowest level, the architecture is fully synthesized, placed and routed by experts to implement a 2-D mesh architecture of configurable algorithmic token machines. At the highest level, the users can program those machines to implement custom processing and routing. The architecture is data driven. The operations are triggered by the arrival of operands, leading to a large and functional pipeline spread over the whole FPGA. This methodology enables the fast implementation of data processing algorithms by people who are not experts in FPGA design, while achieving higher performances than a pure software solution. Two simple examples (FIR and FFT) illustrate the proposed methodology and demonstrate how it is possible to benefit from the expertise encapsulated at low level by just configuring the high level. Another advantage of the proposed methodology is the opportunity to dynamically reconfigure the fabric very quickly to best match the computation requirements at run time.
Mathieu Allard, Patrick Grogan, Yvon Savaria, Jean-Pierre David
ISCAS3
2012 Identification of soft error glitch-propagation paths: Leveraging SAT solvers
abstract
Increase in vulnerability to soft errors has affected the reliability of both synchronous and asynchronous circuits implemented in modern deep sub-micron technologies. Hence in such circuits, there is a growing need to identify the soft error glitch propagation possibility at an early stage in the design flow. This paper proposes a new methodology to obtain soft error glitch propagation paths in digital designs (both synchronous and asynchronous). To compute these paths, Multiway Decision Graphs (MDGs) and glitch-propagation sets (GP sets) are utilized in conjunction with Boolean Satisfiability solvers (MiniSat). The applicability of the proposed method is illustrated by implementing ISCAS89 benchmark sequential circuits, 8-bit adders, multipliers, and the Self-timed multiple-group pipeline asynchronous handshake circuits. The proposed SAT based methodology is on average 13 times faster than the best contemporary state-of-the-art techniques exhaustively analyze possible soft error glitch-propagation paths.
Ghaith Bany Hamad, Otmane Aït Mohamed, Syed Rafay Hasan, Yvon Savaria
ISCAS4
2012 Propagating analog signals through a fully digital network on an electronic system prototyping platform
abstract
The concept of sending and receiving analog signals through a digital interconnection network is presented in this paper. The proposed “analog bus” addresses limitations of a novel rapid prototyping platform called the WaferBoard™ that was initially designed to support prototyping of all digital circuits with its embedded fully digital interconnection network. This paper explores the simplest and least area consuming means of propagating analog signals through a digital interconnection network. A prototype integrated circuit based on the proposed concept was designed using the TSMC 0.18µm technology. The presented prototype is capable of sending an analog signal in the range of 0.6 V to 1.6 V with a maximum frequency of 200 kHz while consuming 68×53.4 µm2of chip area and 19.9 mW of power.
Omar Al-Terkawi Hasib, Walder Andre, Yves Blaquière, Yvon Savaria
ISCAS4
2012 A new approach for pin detection for an electronic system prototyping reconfigurable platform
abstract
A new approach for pin detection in a reconfigurable platform for electronic system prototyping is proposed. It makes use of image processing techniques to first, extract pin core regions by a two-pass process: a top-down multi-level erosion process to remove touching parts of pin regions, followed by a bottom-up pin core recovery process to recover core regions removed by the first process. Once all pin cores have been isolated, regions associated to every pin can be determined by a simple segmentation procedure based on the shortest distance principle. The proposed approach has successfully extracted the pin maps from many circuit footprint images, even in cases of touching pin regions. The results produced by the proposed method have also been compared with those obtained from the reference Watershed algorithm and this shows that our approach provides better results in terms of pin recovery and pin positioning accuracy for the type of images produced by our electronic prototyping system.
Hai H. Nguyen, Mikael Guillemot, Yvon Savaria, Yves Blaquière
RSP3
2012 Efficient Search Algorithm for Determining Optimal R=1/2 Systematic Convolutional Self-Doubly Orthogonal Codes
abstract
A novel implicitly-exhaustive search algorithm for finding, in systematic form, rate R=\frac{1}{2} optimal-span Convolutional Self-Doubly Orthogonal (CDO) codes and Simplified Convolutional Self-Doubly Orthogonal (S-CDO) codes is presented. In order to build high-performance low-latency codecs with these codes, it is important to minimize their constraint length (or "span") for a given J number of generator connections. The proposed algorithm is exhaustive in nature and its improvements over the best previously published searching techniques allowed it to yield new optimal-span CDO/S-CDO codes (having order J ∈ {6,7,8} and J ∈ {9} respectively), as well as a span reduction for codes with a higher J value (J ∈ {10,11} and J ∈ {14,15} for CDO and S-CDO respectively).
Gilbert Kowarzyk, N. Blanger, David Haccoun, Yvon Savaria
IEEE Trans. Commun.4
2012 Real-Time Computation of Local Neighborhood Functions in Application-Specific Instruction-Set Processors
abstract
This paper presents a systematic approach to the design of application-specific instruction-set processors for high speed computation of local neighborhood functions and intra-field deinterlacing. The intended application is real-time processing of high definition video. The approach aims at an efficient utilization of the available memory bandwidth by fully exploiting the data parallelism inherent to the target algorithm class. An appropriate choice of custom instructions and application-specific registers is used together with a very long instruction word architecture in order to mimic a pipelined systolic array. This leads to a processing speed close to the limit imposed by memory bandwidth constraints. For three intra-field deinterlacing algorithms and 2-D convolution with four kernel sizes, the design approach yields speedup factors between 36 and 1330, Area-Time (AT) product improvements between 12× and 243×, and energy consumption reduction factors between 13 and 262.
P. Aubertin, J. M. Pierre Langlois, Yvon Savaria
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Loop Acceleration Exploration for ASIP Architecture
abstract
Design space exploration is a delicate process whose success lays on the designers' shoulders. It is often based on a trial-and-error approach. Some basic metrics can be used to guide this process. In this paper, we explore accelerating loops from C-based specifications. We built a framework in which a design style, such as software-oriented or application-specific instruction-set processor (ASIP)-oriented design, can be specified. We also propose an exploration process that allows targeting the main aspects that limit acceleration and the actions that can be made to improve it. The process is based on new loop-oriented metrics that provide insight in key design issues. They help to determine which aspects of the design between data accesses and arithmetic logic unit (ALU)/control operations limit or allow leveraging loop acceleration opportunities. We profile some benchmarks from the signal and image processing fields, such as the Turbo Decoder and the JPEG algorithms, to illustrate how loop-oriented metrics help to point out aspects that limit or improve loop acceleration. The loop acceleration process was also used to explore design architectures that can leverage, as much as possible, the loop acceleration opportunities of the sum of absolute differences (SAD) algorithm.
Mame Maria Mbaye, Normand Bélanger, Yvon Savaria, Samuel Pierre
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Postsilicon Tuning of Standby Supply Voltage in SRAMs to Reduce Yield Losses Due to Parametric Data-Retention Failures
abstract
Lowering the supply voltage of static random access memories (SRAMs) during standby modes is an effective technique to reduce their leakage power consumption. To maximize leakage reductions, it is desirable to reduce the supply voltage as much as possible. SRAM cells can retain their data down to a certain voltage, called the data-retention voltage (DRV). Due to intra-die variations in process parameters, the DRV of cells differ within a single memory die. Hence, the minimum applicable standby voltage to a memory die (VDDLmin) is determined by the maximum DRV among its constituent cells. On the other hand, inter-die variations result in a die-to-die variation ofVDDLmin. Applying an identical standby voltage to all dies, regardless of their correspondingVDDLmin, can result in the failure of some dies, due to data-retention failures (DRFs), entailing yield losses. In this work, we first show that the yield losses can be significant if the standby voltage of SRAMs is reduced aggressively. Then, we propose a postsilicon standby voltage tuning scheme to avoid the yield losses due to DRFs, while reducing the leakage currents effectively. Simulation results in a 45-nm predictive technology show that tuning standby voltage of SRAMs can enhance data-retention yield by 10%-50%.
Afshin Nourivand, Asim J. Al-Khalili, Yvon Savaria
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Repeater insertion in power-managed VLSI systems
abstract
In this paper, design space exploration methods for interconnect repeaters in DSM power-managed VLSI are proposed. These methods guarantee that the designed interconnects are energy-optimal, while they meet their performance objectives in all the system operating states. These methods take the dynamic output resistance characteristic of the repeaters into account, when the system operating voltage and/or operating frequency requirement changes. Utilizing the proposed design methods, a multi-cycle bus is designed for some performance targets. HSPICE simulations confirm that the designed bus is energy-optimal, and it meets its performance objectives in all the system operating states.
Houman Zarrabi, Asim J. Al-Khalili, Yvon Savaria
ACM Great Lakes Symposium on VLSI3
2011 Comparative analysis of contrast enhancement algorithms in surveillance imaging
abstract
Image contrast enhancement methods play a key role in many image processing and vision applications. For surveillance applications, real-time contrast improvement over the whole image is required when videos are taken in poor lighting conditions. It is also necessary to highlight details in shadowed regions without introducing artifacts. In this paper, several state-of-the-art contrast enhancement methods are compared. Image quality is evaluated by means of objective metrics such as intensity contrast and brightness error, and by subjective assessment. Execution time is also measured. Experimental results show that a technique based on histogram modification presents a better trade-off considering both aspects.
Diana Carolina Gil, Rana Farah, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria
ISCAS5
2011 Analysis of Resistive Open Defects in Drowsy SRAM Cells
Afshin Nourivand, Asim J. Al-Khalili, Yvon Savaria
J. Electron. Test.3
2011 All digital skew tolerant synchronous interfacing methods for high-performance point-to-point communications in deep sub-micron SoCs
Syed Rafay Hasan, Normand Bélanger, Yvon Savaria, M. Omair Ahmad
Integr.3
2010 Fully integrated ultra-low-power asynchronously driven step-down DC-DC converter
abstract
This paper proposes a fully integrated asynchronous step-down switched capacitor DC-DC conversion structure. The circuit uses a fully digital asynchronous state machine as the heart of the control circuitry. To minimize the switching losses, the asynchronous controller scales the switching frequency of the converter according to the load. It also turns on additional parallel switches when needed. This circuit regulates load voltages from 300 mV to 1.1 V derived from a 1.2 V input voltage. A total of 350 pF on chip capacitance was implemented to support a maximum of 250 μW load power, while providing efficiencies up to 80%. The circuit validating the proposed concepts was implemented in 0.13 μm CMOS technology.
Omar Al-Terkawi Hasib, Mohamad Sawan, Yvon Savaria
ISCAS3
2010 An interconnect-aware Dynamic Voltage Scaling scheme for DSM VLSI
abstract
Dynamic Voltage Scaling (DVS) is a successful design solution that addresses the challenges associated with low-power/energy and high-performance design in Deep Sub Micron (DSM) CMOS. In DSM, VLSI systems have become interconnect-centric; correspondingly, the associated design solutions should be adapted to preserve their functionality. In reference to this concern, and with respect to DVS, we propose a DVS scheme that takes interconnect effects into account. The proposed DVS scheme is a generalization of existing methods that treat systems as pure logic. To support this DVS scheme, two design metrics are introduced. These metrics model the performance of system components subject to DVS, based on the proportion of their delay due to interconnects. Based on the proposed design metrics, a compact delay model and a method for supply voltage selection are proposed. The limit of scaling for hazard-free system operation in VLSI systems is further formulated. It is shown that this limit can be smaller than the one dictated by the process technology. The proposed DVS scheme is applied to a 4-section global clock distribution network. Reported results show that this scheme improves both the timing accuracy and energy consumption aspects of DVS by 25% and 30% on average, respectively.
Houman Zarrabi, Asim J. Al-Khalili, Yvon Savaria
ISCAS3
2009 An interconnect-aware delay model for dynamic voltage scaling in NM technologies
abstract
Employing microsystems with Dynamic Voltage Scaling (DVS) is an effective design solution to alleviate their energy consumption. The importance of such design technique keeps growing as both high-performance and low-energy consumption are simultaneously desirable. Existing Power Management Units (PMUs) that support DVS, mainly rely on the delay models valid for CMOS logic. In this work, we show that this may result into improper design and utilization of microsystems subject to DVS; as interconnect delay has become the dominant fraction of the total delay. In accordance with this design concern, we propose a modified delay model which encompasses the effect of interconnect parasitic components, and is suitable for accurate modeling, design and execution of DVS performed by PMUs in nanometer (nm) technologies. HSPICE simulations confirm that the proposed delay model is much more accurate when predicting the performance of a 4-section global H-Tree clock distribution network subject to voltage scaling. The error on predicted performance from true delays is reduced by up to a factor of 4.
Houman Zarrabi, Asim J. Al-Khalili, Yvon Savaria
ACM Great Lakes Symposium on VLSI3
2009 An All-digital Skew-adaptive Clock Scheduling Algorithm for Heterogeneous Multiprocessor Systems on Chips (MPSoCs)
abstract
In this work, we propose a clock scheduling algorithm that is used to mitigate the effects of clock skew that can arise from thermal run-time variations. Depending on the amount of skew, the algorithm selects a different minimum delay tolerance value in order to correct the skew problems, without the performance penalties that are associated with static worst-case scheduling policies. The design was first implemented in MATLAB to obtain the data needed for the clock scheduling. Then, it was implemented in VHDL and synthesized using Xilinx's Virtex-II Pro technology library. Back annotated simulations prove the functionality of the proposed design. The adaptive scheduling scheme achieves up to 60% latency reduction, in our implemented example, compared to a static scheduling scheme.
Syed Rafay Hasan, Bill Pontikakis, Yvon Savaria
ISCAS3
2009 Workflow for an Electronic Configurable Prototyping System
abstract
A recently proposed rapid prototyping technology for electronic systems, which is based on a WSI active configurable circuit board comprising more than one million contact, can be programmed to interconnect integrated circuit packages deposited on its surface. This technology has some similarities, but also some key distinctive constraints when compared to conventional printed circuit boards. A workflow that supports the design with such configurable circuit boards is proposed. As part of this workflow, algorithms and tools for package recognition and for routing through a multi-dimensional mesh interconnection network is proposed and implemented. Results reported in this paper confirm the feasibility of the proposed workflow and several architectural choices made with respect to the configurable circuit board technology. Using the prototype tools reported in this paper, packages are successfully recognized and netlists are routed even though they use up to 50% of the contact point resources, which corresponds to an extremely dense circuit board.
Etienne Lepercq, Yves Blaquière, Richard Norman, Yvon Savaria
ISCAS4
2009 A low-power 2GHz data conversion using delta modulation for portable application
Ali Naderi, Mohamad Sawan, Yvon Savaria
Integr.3
2008 Loop-oriented metrics for exploring an application-specific architecture design-space
abstract
Since ASIPs were introduced in the HW/SW architecture design space, application partitioning has become more complex. Designers have more ways to accelerate applications: with ASIPs of various kinds or with dedicated hardware modules. In this paper, we present loop-oriented metrics that will be used during design-space exploration for the partitioning process of C-based designs. These metrics help designers determine which aspect of loop iterations, between data memory accesses and ALU/Control operations, offers more acceleration potential. We implemented a profiler-scheduler LOOPPROF that gathers the metrics. Our tool also helps determine which optimization techniques such as data reuse are suitable for the considered code segments. We demonstrate the use of our tool by exploring the acceleration possibilities of the ELA Deinterlacer, a video processing algorithm.
Mame Maria Mbaye, Normand Bélanger, Yvon Savaria, Samuel Pierre
ASAP3
2008 Modeling and simulation of complex heterogeneous systems
abstract
Given the increasing heterogeneity and complexity of systems being developed, untimed modeling at a system level becomes more and more important for design space exploration and verification, due to its conciseness and speed. After showing inadequacies of SystemC, which is the predominant modeling environment in this area, we propose a paradigm shift from immediate notifications and coroutines in SystemC to Atomic Actions and true parallelism in an extension of Esys.NET. We exploit the introspection and attribute programming to extend the capabilities of the environment and to build the basis for heterogeneous cosimulation. This paper aims to show the main advantages of this paradigm shift, such as (1) the improvement of simulation time by exploiting the capabilities of multicore simulation hosts, (2) the reduction of modeling hazards related to parallelism and resource sharing, and (3) a more efficient design space exploration.
Amine Anane, El Mostapha Aboulhamid, Julie Vachon, Yvon Savaria
ISCAS4
2007 Modeling the Substrate Noise Injected by a DC-DC Converter
abstract
A custom substrate model is proposed to analyze the noise injected by a DC-DC converter. Even if commercial tools extracting substrate parasitic components already exist, they are based on a capacitive and resistive model that can prove inadequate to simulate noise injected by power switching modules, since unusual voltage variations can activate parasitic vertical bipolar transistors. The proposed model includes these components and shows an important impact on victim circuits in the neighborhood of the noise source. Simulations with the proposed model predict a 6 mV noise in the substrate that is a 67-fold increase compared to the 90μV obtained with a model produced bySubstrateStorm. These simulations with our custom model show that deep n-well and n-well noise isolations can be much less effective than expected. The proposed model is directly applicable to exploring the tradeoff between power efficiency and substrate noise injection in integrated DC-DC converters.
Vincent Binet, Yvon Savaria, Michel Meunier, Yves Gagnon
ISCAS2
2007 High-Voltage DMOS Integrated Circuits with Floating Gate Protection Technique
abstract
This paper presents an efficient low power protection technique for thin gate oxide of DMOS transistors. By connecting a capacitive divider structure to the floating gate node of a DMOS transistor, its effective gate oxide thickness is increased, and a protection from breakdown due to high voltages (HV) applied to its gate is achieved. Several HV circuits, including: positive voltage doubler and level-up shifter suitable for ultrasound sensing systems are built successfully around this technique. These circuits were implemented with the 0.8 μm CMOS/DMOS HV DALSA process. Experimental results prove the good functionality of the designed HV circuits using the proposed protection technique for voltages up to 120V.
Robert Chebli, Mohamad Sawan, Yvon Savaria, Kamal El-Sankary
ISCAS3
2007 Crosstalk Effects in Event-Driven Self-Timed Circuits Designed With 90nm CMOS Technology
abstract
Systems-on-chip (SoCs) designed in ultra-deep sub-micron technologies (90nm and beyond) often comprise modules in multiple clock domains (MCD), which are usually interconnected using asynchronous interfaces. At the same time, in ultra-deep sub-micron (DSM) technologies, minimum width, spacing, inter-metal dielectric lengths are reduced, as well as distances between metal layers. These trends raise the coupling capacitance resulting in more severe crosstalks. Therefore, asynchronous interfaces may be subject to crosstalk in ultra-DSM technologies. In this paper, a quantitative investigation is performed to approximate the crosstalk effects in 90nm technology, and to compare them with effects in other DSM technologies. It is found that for wire lengths of 1mm, and more, crosstalk effects in a 90nm technology are substantially higher, about 1.3 times, than in a 180nm technology. Furthermore, three well known self-timed asynchronous design methods are analyzed with regards to crosstalk and the importance of coupling capacitances is established. It is shown that glitches can cause errors in self-timed designs. To our knowledge, this paper is the first to report crosstalk sensitivity in self-timed circuits, which are notably proposed as a solution to the timing problems found in advanced SoCs.
Syed Rafay Hasan, Yvon Savaria
ISCAS2
2007 A Low-Complexity High-Speed Clock Generator for Dynamic Frequency Scaling of FPGA and Standard-Cell Based Designs
abstract
In this paper, the authors propose two high-speed variable rate clock generator circuits that can synthesize frequencies which are fractional multiples of an input clock. The designs can switch between frequencies in a glitch-free manner, within a single clock cycle. In response to an N-phase reference clock, the first circuit can generate a clock of up to N times the reference frequency, whereas the second solution can generate up to N/2 times that frequency. The available synthesized frequencies are given by frefmiddotN/M , where M can be any integer greater than or equal to 1, depending on the circuit. The solutions were coded in VHDL, synthesized, placed and routed in TSMC's 180nm CMOS technology. Simulations using the extracted layout show that the proposed designs can operate with a reference frequency of up to 400MHz, yielding a maximum output clock of 4times the reference, or 1.6GHz. The designs were also validated with an implementation on Xilinx's Spartan 3 FPGA device.
Bill Pontikakis, Hung Tien Bui, François R. Boyer, Yvon Savaria
ISCAS4
2007 Integrated Circuit Trimming Technique for Offset Reduction in a Precision CMOS Amplifier
abstract
This article presents an application of a recently reported IC trimming technique using laser diffused resistors to reduce the input referred offset voltage of a precision amplifier. A three stage precision CMOS operational amplifier topology is proposed utilizing laser-trimmable diffused resistors for post-fabrication trimming. The amplifier is designed to operate over an industrial temperature range (-40°C to +85°C) including process corners and utilizes an on-chip CMOS bias generation circuit to maintain a robust performance. The results of post-layout simulation of the complete circuit are summarized. The effects of the trimming technique on the input offset voltage of the amplifier are described. The circuit is designed using the TSMC 0.18μm CMOS process and operates from a single supply of 3.3 V.
Yves Audet, Yves Gagnon, Yvon Savaria
ISCAS4
2006 Soft-error classification and impact analysis on real-time operating systems
abstract
This paper investigates the sensitivity of real-time systems running applications under operating systems that are subject to soft-errors. We consider applications using different real-time operating system services: scheduling, time and memory management, intertask communication and synchronization. We report results of a detailed analysis regarding the impact of soft-errors on real-time operating systems cores, taking into account the application timing constraints. Our results show the extent to which soft-errors occurring in a real-time operating system's kernel impact its reliability
N. Ignat, Bogdan Nicolescu, Yvon Savaria, Gabriela Nicolescu
DATE3
2006 High speed differential pulse-width control loop based on frequency-to-voltage converters
abstract
A novel differential pulse-width control loop circuit based on high speed frequency-to-voltage converters is proposed. To demonstrate its functionality, a circuit has been designed and simulated in 0.18mm CMOS technology. Results show that the proposed circuit can correct a clock signal's duty cycle even for frequencies as high as 5 GHz. This design can be used to correct clock signal distortion due to process variations in high speed applications such as half-rate clock and data recovery systems.
Hung Tien Bui, Yvon Savaria
ACM Great Lakes Symposium on VLSI2
2006 Architecture of a hypertransport tunnel
abstract
This paper presents a technology independent and open source hypertransport (HT) tunnel. HT is a high performance and low latency chip-to-chip interconnect standard. Various aspects of the architecture are presented, ranging from how functionality is spread over clock domains to the means of implementing a packet reordering algorithm. The analysis of synthesis results provides a better insight of the complexity of various features and what limits the performance of the HT tunnel
Ami Castonguay, Yvon Savaria
ISCAS2
2006 A power planning model for implantable stimulators
abstract
This paper presents a new analytical, empirical and behavioral modular model developed for accurate evaluation of power dissipation in power conversion chains (PCC) dedicated to power up an electronic implantable device. The model is suitable for power estimation/planning in early design stages, to determine the contribution of each circuit module on the total power consumption and to estimate the input and output voltages of these modules. It is based on average power consumption model and is coded in Verilog-A. The model is verified and the results were found to be in good agreement with state-of-the-art designs for bioelectronics devices. It is flexible and robust to changes in architecture and design parameters and provides accurate and valid results in a fraction of second for a large variety of parameter values. The ease of implementing desired modules and architectures makes the model more advantageous
Saeid Hashemi, Mohamad Sawan, Yvon Savaria
ISCAS3
2006 High-voltage operational amplifier based on dual floating-gate transistors
abstract
A high-voltage operational amplifier (hvopamp) using dual-input floating-gate transistors for its feedback network is presented. The proposed hvopamp stabilizes the output DC voltage in the middle of its high-voltage power supply. Using floating-gate transistors eliminates the need for high-voltage resistor feedback networks. The integral nonlinearity (INL) of the hvopamp with floating gate feedback is 7% in the rail-to-rail output range, which is better than the performance of circuit using parasitic field-oxide MOS transistor as feedback network. The designer does not need to develop a new component, and can implement easily floating-gate transistors in most technologies, which facilitates design and improves the circuit robustness.
Yvon Savaria, Mohamad Sawan, R. Meinga
ISCAS2
2006 Design exploration with an application-specific instruction-set processor for ELA deinterlacing
abstract
Achievable performance gains, when accelerating applications using ASIPs, with a good sequence of specialized instructions, depends on the applications' available parallelism, and possibilities for optimizations and transformations. The type and number of operations, and the number of data transfers of the application are also critical factors. Much progress has been done on ASIP customized instruction-identification and selection research; they are usually based on operation clustering. In this paper, we propose to minimize the number of data transfers during execution of specialized instructions sequence by storing temporary values in user-defined registers. The method avoids costly data transfers and allows parallel processing of demanding computations. This method is applied to the design of an ASIP dedicated to edge line average deinterlacing, an algorithm used in HDTV. Experimental results show that our design method applied to this application, yields a speedup factor larger than 18.
Mame Maria Mbaye, D. Lebel, Normand Bélanger, Yvon Savaria, Samuel Pierre
ISCAS4
2006 A novel 2-GHz band-pass delta modulator dedicated to wireless receivers
abstract
This paper describes a sub-sampling delta modulator operating at giga Hertz range to capture radio frequency signals. Down-conversion to low-IF is achieved by sub-sampling with a 1-bit quantizer. It presents higher bandwidth and SNR than those of the state-of-the-art sub-sampling modulators. Input carrier frequency can be followed over a wide range by controlling the sampling rate. Center frequency of the band-pass filters, which is placed at IF, is independent of input carrier frequency. A SNR higher than 55 dB is expected for a 2 MHz bandwidth signal modulated at 2-GHz frequency when the sampling rate is set to 990 mega samples per second
Ali Naderi, Mohamad Sawan, Yvon Savaria
ISCAS3
2006 A 0.8V algorithmically defined buffer and ring oscillator low-energy design for nanometer SoCs
abstract
In this paper, an algorithmically defined buffer and ring oscillator design for low energy applications is proposed. The goal of the algorithm is to easily converge to a low energy solution while the system maintains constant speed and full swing at a given supply voltage, irrespective of the capacitive load. The experimental circuit is a 980MHz oscillator operating from a 0.8V supply, driving a 1pF load, designed using a 0.18mum TSMC CMOS process technology. A comparison to the well known minimum delay tapered buffer, using an exponential horn designed independently of the oscillator, is done. The comparison shows that our algorithm produces a 3.7-3.9 times improvement in terms of power, energy, and energy delay product (EDP) metrics, and 14.6 times improvement in terms of the energy area product (EAP)
Bill Pontikakis, François R. Boyer, Yvon Savaria
ISCAS3
2006 Zero skew differential clock distribution network
abstract
Clock uncertainty is a major concern in current high performance clock network design. A differential clocking scheme provides noise immunity and can address this challenge. In this paper, a differential line equivalent delay model is proposed to obtain zero skew differential clock networks. The method is applied to various benchmarks. On average, 97% skew reduction is obtained compared to the solution derived with the classic Elmore model. To improve performance of zero skew differential clock networks, differential buffers based on dynamic threshold transistors are proposed. Incorporation of proposed buffers to low swing zero skew differential clock network shows 25% delay improvement compared to conventional buffers. Moreover, the incorporation of proposed design methods shows 25% and 6% skew variations reduction in presence of power supply variations and crosstalk noise respectively, compared to low swing single-node clock distribution networks.
Houman Zarrabi, Haydar Saaied, Asim J. Al-Khalili, Yvon Savaria
ISCAS4
2006 A Metric for Automatic Word-Length Determination of Hardware Datapaths
abstract
A metric for the automatic determination of word lengths required for implementing DSP algorithms is proposed. The metric is capable of handling several error models computed between the fixed-point and the floating-point simulation results to model the impact of finite word lengths on the overall accuracy. It grades all the word-length combinations and guides a procedure towards the optimal solution. This metric was implemented in an automatic word-length determination tool to guide its search for better hardware implementations. It enables the creation of a framework for architecture and platform exploration.
Marc-André Cantin, Yvon Savaria, D. Prodanos, Pierre Lavoie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 The Role of Model-Level Transactors and UML in Functional Prototyping of Systems-on-Chip: A Software-Radio Application
abstract
Developing a functional prototype of a system-on-chip provides a unifying vehicle for model validation and system refinement. Keeping the prototype executable across several abstraction levels, clock domains and design tools is a key requirement to effective prototyping. This paper presents how model-level transactors address design heterogeneity by unifying event-based and cycle-based worlds from specification to implementation. Transactors are used to build a functional prototype of a software-radio component. An executable UML model is bridged to a hardware abstraction of a radio stream developed with Simulink to implement a realistic and working prototype. Model validation and performance measurements are realized through prototype execution and real-time monitoring.
Alexandre Chureau, Yvon Savaria, El Mostapha Aboulhamid
DATE2
2005 Scheduling and optimal register placement for synchronous circuits derived using software pipelining techniques
abstract
Data dependency constraints constitute a lower bound P on the minimal clock period of single-phase clocked sequential circuits. In contrast to methods based on basic retiming, clocked sequential circuits with clock period P can always be obtained using software pipelining techniques. Such circuits can be derived by any method that can be framed in the following four-step process: Step 1, determine P; Step 2, compute a valid periodic schedule of the computational elements; Step 3, place registers back to the circuit; Step 4, assign the clock signals to control registers.Methods with polynomial run-time to implement this process are proposed in the literature. They implement these steps sequentially, starting with Step 1. These methods do not know how to optimally place registers which leads to an unnecessary number of registers. In this article, we address the problem of how to simultaneously implement Steps 2 and 3 in order to minimize the total number of registers. We conjecture that the problem is NP-hard in its general form. We formulate the problem for the first time in the literature, and devise a Mixed Integer Linear Program (MILP) to solve it. From this MILP, we derive a linear program to determine approximate solutions to the problem for large general circuits. We show that the proposed approach can handle nonzero clock skew. Experimental results confirm the effectiveness of the approach and show that significant reductions of the number of registers can be obtained although register sharing is not used. When the schedule is given, the proposed approach provides solutions to the problem of how to place the minimal number of registers in Step 3.
Noureddine Chabini, El Mostapha Aboulhamid, Ismaïl Chabini, Yvon Savaria
ACM Trans. Design Autom. Electr. Syst.4
2004 Performance Evaluation and Failure Rate Prediction for the Soft Implemented Error Detection Technique
Bogdan Nicolescu, Yvon Savaria, Raoul Velazco
IOLTS2
2003 Unification of basic retiming and supply voltage scaling to minimize dynamic power consumption for synchronous digital designs
abstract
We address the problem of minimizing dynamic power consumption for single-phase synchronous digital designs, under timing constraints, using an unification of basic retiming and supply voltage scaling. We assume that the number of supply voltages and their values are known for each computation element. Our main objective is then to change the location of registers using basic retiming while maximizing the number of computation elements off critical paths that can operate under a low available supply voltage, and can lead to a maximum dynamic power saving. We address the problem at the system-level. We formulate the problem as a Mixed Integer Linear Program (MILP). The exact optimal solution for the problem is then guaranteed. We also devise an algorithm to compute bounds on the values assigned by basic retiming to each computational element. Besides helping to find the optimal solution to the problem, these bounds also allow to reduce the run-time for finding this solution. The proposed approach can produce converter-free designs and can also minimize short-circuit power consumption. Experimental results have shown that dynamic power consumption can be reduced by factors that range from 2.78% to 37.24% for single-phase designs with minimal clock period. For these experimental results, the run-time for solving the MILP is under 2min.
Noureddine Chabini, Ismaïl Chabini, El Mostapha Aboulhamid, Yvon Savaria
ACM Great Lakes Symposium on VLSI4
2003 A Pattern Reordering Approach Based on Ambiguity Detection for Online Category Learning
abstract
Pattern reordering is proposed as an alternative to sequential and batch processing for online category learning. Upon detecting that the categorization of a new input pattern is ambiguous, the input is postponed for a predefined time, after which it is reexamined and categorized for good. This approach is shown to improve the categorization performance over purely sequential processing, while yielding a shorter input response time, or latency, than batch processing. In order to examine the response time of processing schemes, the latency of a typical implementation is derived and compared to lower bounds. Gaussian and softmax models are derived from reject option theory and are considered for detecting ambiguity and triggering pattern postponement. The average latency and Rand Adjusted clustering score of reordered, sequential, and batch processing are compared through computer simulation using two unsupervised competitive learning neural networks and a radar pulse data set.
Eric Granger, Yvon Savaria, Pierre Lavoie
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Methods for minimizing dynamic power consumption in synchronous designs with multiple supply voltages
abstract
We address the problem of minimizing dynamic power consumption under performance constraints by scaling down the supply voltage of computational elements off critical paths. We assume that the number of possible supply voltages and their values are known for each computational element. We focus on solving this problem on cyclic and acyclic graphs corresponding to synchronous designs. We consider multiphase clocked sequential circuits derived using software pipelining techniques. In this paper, we present exact and heuristic methods to solve the problem. The proposed methods take the form of mathematical programming formulations and their associated solution algorithms. The exact methods are based on a mixed integer linear programming formulation of the problem. The heuristic methods are based on linear programming formulations derived from the exact problem formulation. Solution methods are analyzed experimentally in terms of their run time and effectiveness in finding designs with lower dynamic power using circuits from the ISCAS89 benchmark suite. Power reduction factors as high as 69.75% were obtained compared to designs using the highest supply voltages. One of the heuristic methods leads to solutions that are near optimal, typically within 5% from the optimal solution. Low dynamic-power designs with no or a small number of level converters, are also obtained.
Noureddine Chabini, Ismaïl Chabini, El Mostapha Aboulhamid, Yvon Savaria
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2002 A flexible floating-point format for optimizing data-paths and operators in FPGA based DSPs
abstract
Video signal processing requires complex algorithms performing many basic operations on a video stream. To perform these calculations in real-time in a FPGA, we must use innovative structures to meet speed requirements while managing complexity. As part of a project aiming at the development of a video noise reducer, we developed an optimized processing stream that required some floating-point calculations. This paper presents the rationale for developing a floating-point unit, justifies the data representation used, its implementation in a Xilinx VirtexE FPGA and reports the performance we obtained. A divider using this representation is also presented, with its implementation and performances in the same FPGA.
J. Dido, N. Géraudie, L. Loiseau, O. Payeur, Yvon Savaria, D. Poirier
FPGA5
2002 A practical approach to model long MIS interconnects in VLSI circuits
abstract
In this paper, a practical approach to model metal-insulator-semiconductor (MIS) interconnects is presented, with focus on the microstrip configuration. Starting from a one-dimensional (1-D) electromagnetic field analysis, we first extend the validity range of some closed-form expressions from 1-D to two-dimensional (2-D) and present an original RLCG-B model with five equivalent circuit parameters. These parameters, which depend on two effective widths of the physical metal strip, can be frequency dependent because of the skin effect and the dielectric losses. The original RLCG-B model is then modified and implemented with seven frequency-independent circuit parameters. These parameters are computed by analytical equations. Numerical simulations are used to validate the original and modified RLCG-B models. A formula to allow comparison of various interconnect models in the time domain is proposed. Comparisons based on this formula are presented for a single transmission line with source resistance, R/sub S/, and load capacitance, C/sub L/. Such comparisons are more meaningful in VLSI applications than comparisons of characteristics derived from swept-frequency per-unit-length parameters.
Zhong-Fang Jin, Jean-Jacques Laurin, Yvon Savaria
IEEE Trans. Very Large Scale Integr. Syst.3
2001 Determining Schedules for Reducing Power Consumption Using Multiple Supply Voltages
abstract
Dynamic power is the main source of power consumption in CMOS circuits. It depends on the square of the supply voltage. It may significantly be reduced by scaling down the supply voltage of some computational elements in the circuit, with the penalty of an increase of their execution delay. To reduce the dynamic power consumption, without degrading the performance determined assuming that the circuit operates at the highest available supply voltage, the supply voltage of computational elements off critical paths can be scaled down. Defined here as MinP/sub dyn/, the problem of minimizing the dynamic power consumption, under performance constraints, by scaling down the supply voltage of computational elements on non-critical paths is NP-hard in general. Solving MinP/sub dyn/ for multi-phase clocked sequential circuits may allow to reduce their power consumption and the required number of registers. Reducing the number of registers also allows to reduce the power consumption, the number of control signals, and the area of the circuit. In this paper, we focus on devising methods to efficiently solve MinP/sub dyn/ for designs modeled as cyclic or acyclic graphs. More precisely, once the circuit is optimized for timing constraints, then we look for schedules that allow the computational elements of the circuit to operate at the lowest possible supply voltage. We present an integer linear programming formulation for that problem, which we use to devise a polynomial time solvable method and an exact algorithm based on a branch-and-bound technique. Experimental results confirm the effectiveness of the method and power reduction factors as high as 53.84% were obtained. Also, they show that the exact algorithm produces optimal results in a small number of tries, which is due to the rules used to prune useless solutions.
Noureddine Chabini, El Mostapha Aboulhamid, Yvon Savaria
ICCD3
2001 Tools for the Characterization of Bipolar CML Testability
abstract
A methodology to characterize thoroughly the defective behavior of CML bipolar gates has been developed. This methodology produced data suitable to guide design for testability in CML circuits. Inductive Fault Analysis (IFA) is first applied to library cell layouts to characterize their sensitivity to realistic defects. The data is then processed by an automatic simulation program that can, according to a list of criteria, classify the defective behavior of all cells. This complete analysis allows a thorough characterization of defective cell behaviors of such a logic family, and helps derive specific testability rules and methods. The proposed methodology is flexible enough to be adapted easily to other logic families.
Ginette Monté, Bernard Antaki, Serge Patenaude, Yvon Savaria, Claude Thibeault, Pieter M. Trouborst
VTS4
2001 Optimal design of synchronous circuits using software pipelining techniques
abstract
We present a method to optimize clocked circuits by relocating and changing the time of activation of registers to maximize the throughput. Our method is based on a modulo scheduling algorithm for software pipelining, instead of retiming. It optimizes the circuit without the constraint on the clock phases that retiming has, which permits to always achieve the optimal clock period. The two methods have the same overall time complexity, but we avoid the computation of all pair-shortest paths, which is a heavy burden regarding both space and time. From the optimal schedule found, registers are placed in the circuit without looking at where the original registers were. The resulting circuit is a multi-phase clocked circuit, where all the clocks have the same period and the phases are automatically determined by the algorithm. Edge-triggered flip-flops are used where the combinational delays exactly match that period, whereas level-sensitive latches are used elsewhere, improving the area occupied by the circuit. Experiments on existing and newly developed benchmarks show a substantial performance improvement compared to previously published work.
François R. Boyer, El Mostapha Aboulhamid, Yvon Savaria, Michel Boyer
ACM Trans. Design Autom. Electr. Syst.3
2000 Analysis of quantization effects in a digital hardware implementation of a fuzzy ART neural network algorithm
abstract
A reformulated Adaptive Resonance Theory (ART) neural network algorithm has recently been implemented in digital hardware. Naturally, the fixed point, fixed word length data format used causes some output differences with respect to floating point computer simulation. These differences are observed when using realistic input data. The effects of input quantization and the accumulation of round off errors in the arithmetic operations making up the algorithm are analyzed. Even a small quantization or round off error can trigger a change in the clustering produced. This does not mean that the clustering is not valid. Indeed, the validity of the clustering can be comparable to that obtained by floating point computer simulation, provided the word length is sufficient. This is verified on realistic input data consisting of radar pulses received from a number of emitters.
Marc-André Cantin, Yves Blaquière, Yvon Savaria, Pierre Lavoie, Eric Granger
ISCAS3
2000 A new fully integrated CMOS phase-locked loop with low jitter and fast lock time
abstract
In this paper we describe a novel PLL circuit design. The proposed topology is based on two loops: the conventional fine loop and a new coarse loop. The fine tuning loop which includes a phase-frequency detector, a charge pump and a differential voltage controlled oscillator (unity feedback PLL) is rather slow. However the coarse tuning loop reacts faster and accelerates convergence. It also ensures a better stability, a shorter locking time, and as a result, a low jitter is obtained, as well as a lower sensitivity to power supply variations.
Youcef Fouzar, Mohamad Sawan, Yvon Savaria
ISCAS3
2000 A methodology for validating digital circuits with mutation testing
abstract
This paper proposes a systematic methodology for improving functional validation vectors developed to check digital circuits. This method exploits the mutation testing concept originally proposed for software validation. Mutation injects specific functional transformations in circuit descriptions expressed in languages like VHDL or Verilog. These programs, called mutant, are syntactically correct but functionally incorrect. Knowing how these vectors detect functional faults improves the confidence in the design and provide information on the coverage of validation vectors. The paper identifies limits of previous work on mutation testing applied to hardware and proposes method that are better suited to the task.
Patrice Vado, Yvon Savaria, Yannick Zoccarato, Chantal Robach
ISCAS2
1999 Design For Testability Method for CML Digital Circuits
abstract
This paper presents a new Design for Testability (DFT) technique for Current-Mode Logic (CML) circuits. This new technique, with little overhead, using built-in detectors, monitors all gate output swings and flags all abnormal voltage excursions. These detectors cover classes of faults that cannot be tested by stuck-at testing methods only. Circuit simulations have shown that abnormal gate output excursions caused by the presence of a defect are common with CML. We also show that this technique works well below "at-speed" frequencies. Finally, variants of the built-in detectors with reduced area overhead are proposed.
Bernard Antaki, Yvon Savaria, Nanhan Xiong, Saman Adham
DATE2
1999 Design of a JTAG Based Run Time Reconfigurable System
abstract
In the past few years, the concept of run-time reconfigurable (RTR) systems has received a great deal of attention. RTR is the ability for a system to change its hardware configuration while computing, to address changing bottlenecks. This paper proposes a practical and effective means of implementing RTR systems with widely available and proven FPGA technology that keeps most of the practical benefits of RTR at the system level.
Cynthia Cousineau, François Laperle, Yvon Savaria
FCCM3
1999 Generalization, discrimination, and multiple categorization using adaptive resonance theory
abstract
The internal competition between categories in the adaptive resonance theory (ART) neural model can be biased by replacing the original choice function by one that contains an attentional tuning parameter under external control. For the same input but different values of the attentional tuning parameter, the network can learn and recall different categories with different degrees of generality, thus permitting the coexistence of both general and specific categorizations of the same set of data. Any number of these categorizations can be learned within one and the same network by virtue of generalization and discrimination properties. A simple model in which the attentional tuning parameter and the vigilance parameter of ART are linked together is described. The self-stabilization property is shown to be preserved for an arbitrary sequence of analog inputs, and for arbitrary orderings of arbitrarily chosen vigilance levels.
Pierre Lavoie, Jean-François Crespo, Yvon Savaria
IEEE Trans. Neural Networks3
1999 Reconfigurable pipelined 2-D convolvers for fast digital signal processing
abstract
In order to make software applications simpler to write and easier to maintain, a software digital signal-processing library that performs essential signal- and image-processing functions is an important part of every digital signal processor (DSP) developer's toolset. In general, such a library provides high-level interface and mechanisms, therefore, developers only need to know how to use algorithms, not the details of how they work. Complex signal transformations then become function calls, e.g., C-callable functions. Considering the two-dimensional (2-D) convolver function as an example of great significance for DSP's, this paper proposes to replace this software function by an emulation on a field-programmable gate array (FPGA) initially configured by software programming. Therefore, the exploration of the 2-D convolver's design space will provide guidelines for the development of a library of DSP-oriented hardware configurations intended to significantly speed up the performance of general DSP processors. Based on the specific convolver, and considering operators supported in the library as hardware accelerators, a series of tradeoffs for efficiently exploiting the bandwidth between the general-purpose DSP and accelerators are proposed. In terms of implementation, this paper explores the performance and architectural tradeoffs involved in the design of an FPGA-based 2-D convolution coprocessor for the TMS320C40 DSP microprocessor available from Texas Instruments Incorporated. However, the proposed concept is not limited to a particular processor.
B. Bosi, Guy Bois, Yvon Savaria
IEEE Trans. Very Large Scale Integr. Syst.3
1998 A Comparative Analysis of Fuzzy ART Neural Network Implementations: The Advantages of Reconfigurable Computing
abstract
The paper analyzes the performance differences found between software and hardware/sofware implementations of a reformulated fuzzy ART neural network algorithm. This reformulated algorithm is a solution for a real time radar signal clustering problem. The software implementations run on a 50 MHz TMS320C40 DSP, and the hardware/sofware implementation runs on the same DSP for its software part, whereas the FPGA based application specific hardware accelerator is realized on MiroTech's X-CIM TIM40 module. This investigation of FPGA based acceleration gave excellent results for our application: acceleration factors up to 68.9 have been reached.
Pascal Poiré, Marc-André Cantin, Hervé Daniel, Yves Blaquière, Yvon Savaria
FCCM5
1998 Design of Clock Distribution Networks in Presence of Process Variations
abstract
Tolerance to process-induced skew remains one of the major concerns in the design of large-area and highspeed clock distribution networks. Indeed, despite the availability of some efficient exact-zero skew algorithms that can be applied during circuit design, the clock skew remains an important performance limiting factor after chip manufacturing, and is of increasing concern for sub-micron technologies. This tutorial reviews the importance of the problem, its sources, as well as typical examples of existing solutions. Solutions range from design rule strategies to built-in self-compensation methods.
Mohamed Nekili, Yvon Savaria, Guy Bois
Great Lakes Symposium on VLSI2
1998 Optimal design of synchronous circuits using software pipelining techniques
abstract
In this paper, we present a method to optimize clocked circuits by relocating and changing the time of activation of registers to maximize throughput. Our method is based on software pipelining instead of retiming. The two methods have the same overall complexity, but unlike previously published retiming methods, the time consuming step of searching an adequate clock period is avoided, since the optimal clock period is always a solution. The resulting circuit is a multi-phase clocked circuit, where all the clocks have the same period. Edge-triggered flip-flops are used where the combinational delays exactly match that period, whereas level-sensitive latches are used elsewhere improving the area occupied by the circuit. Experiments on existing and newly developed benchmarks show substantial performance improvement compared to previously published work.
François R. Boyer, El Mostapha Aboulhamid, Yvon Savaria, Imed E. Bennour
ICCD3
1998 Parallel ultra large scale engine SIMD architecture for real-time digital signal processing applications
abstract
The instruction set and architecture of a SIMD processor optimized for real-time digital signal processing applications is presented. A novel structure allows computation and data I/O to be performed in parallel, provides interprocessor communications and enables the cascade of multiple chips without glue-logic to provide systems with large numbers of processors. Powerful processing elements and instruction set architecture allow complex real-time linear and non-linear algorithms to be coded. Dynamically reconfigurable systems can be also be built. A demonstrator chip was manufactured by ChipExpress. Benchmark comparisons and the methodology used to implement image processing applications on the PULSE, are also presented.
Paul Marriott, Ivan C. Kraljic, Yvon Savaria
ICCD3
1998 A comparison of self-organizing neural networks for fast clustering of radar pulses
Eric Granger, Yvon Savaria, Pierre Lavoie, Marc-André Cantin
Signal Process.2
1997 Pipelined H-trees for high-speed clocking of large integrated systems in presence of process variations
abstract
This paper addresses the problem of clocking large high-speed digital systems, as well as deterministic skew modeling, a related problem. In order to provide a reliable skew model, and to avoid the frequency limitation, we propose a novel approach that distributes the clock with an H-tree, whose branches are composed of minimum-sized inverters rather than metal. With such a structure, we obtain the highest clocking rate achievable with a given technology. Indeed, clock rates around 1 GHz are possible with a 1.2 /spl mu/m CMOS technology. From the skew modeling standpoint, we derive an analytic expression of the skew between two leaves of the H-tree, which we consider to be the difference in root-to-leaf delay pairs. The skew upper bound obtained has an order of complexity which, with respect to the H-tree size D, is the same as the one that may be derived from the Fisher and Kung model for both side-to-side and neighbor-to-neighbor communications, i.e., a /spl Omega/(D/sup 2/), whereas, the Steiglitz and Kugelmass probabilistic model predicts /spl Theta/(D/spl times//spl radic/LogD). In an H-tree implemented with metallic lines, the leaf-to-leaf skew is obviously bounded by the delay between the root and the leaves. However, with the logic based H-tree proposed here, we arrive at a nonobvious result, which states that the leaf-to-leaf skew grows faster than the root-to-leaf delay in presence of a uniform transistor time constant gradient. This paper also proposes generalizations of the skew model to (1) the case of chips in a wafer subject to a smooth, but nonuniform gradient and (2) the case of H-tree configurations mixing logic and interconnections; in this respect, this paper covers the H-tree configurations based on the combination of logic and interconnections.
Mohamed Nekili, Guy Bois, Yvon Savaria
IEEE Trans. Very Large Scale Integr. Syst.3
1996 Design and performance of CMOS TSPC cells for high speed pseudo random testing
abstract
In this paper the problem of testing high speed CMOS circuits is considered. A test methodology based on a Built-in Self-Test scheme adapted to TSPC circuits is proposed. We show through HSpice simulations on netlists extracted from layout that this scheme can operate at more than 580 MHz. Moreover, an efficient solution to the problem of redesigning an easily testable functionally equivalent logic block to eliminate hard to test and untestable faults in TSPC circuits is introduced.
Mohamed Soufi, Steve Rochon, Yvon Savaria, Bozena Kaminska
VTS3
1996 Reconstruction method for jitter tolerant data acquisition system
Adel Belhaouane, Yvon Savaria, Bozena Kaminska, Daniel Massicotte
J. Electron. Test.2
1996 Timing analysis speed-up using a hierarchical and a multimode approach
abstract
In this paper, we examine the impact of using the hierarchy of the design and multiple delay models defined at different abstraction levels to speed up the timing performance evaluation of VLSI circuits. The algorithms implemented in the Dynamic and Hierarchical Timing Analysis (DHTA) tool are described. DHTA rapidly identifies the critical portions of the circuit at high hierarchical levels with rough delay models. These portions are then successively studied at more detailed levels for maximal accuracy. The effects on processing time of exploiting the design hierarchy and using several delay models are characterized. The implementation of DHTA demonstrates experimentally the benefits of using a mixed-mode approach for timing analysis. We show that considering all available hierarchical levels may degrade the computing time and heuristics are proposed to select the hierarchical levels which generally lead to a speed-up.
Yves Blaquière, Michel R. Dagenais, Yvon Savaria
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1995 A new method for testing mixed analog and digital circuits
abstract
In this paper a new method is proposed for observing analog test points inside integrated circuits that enables the simultaneous observation of a large number of points. The method permits the removal of the analog multiplexer from the signal path and a reduction of the load introduced at the observed test points. A charge coupled device analog shift register is used to sample input voltage and shift out a charge that is proportional to the input voltage.
Janusz Rzeszut, Bozena Kaminska, Yvon Savaria
Asian Test Symposium3
1995 On Using Partial Reset for Pseudo-Random Testing
abstract
An inexpensive Design For Testability (DFT) technique, called Partial Reset (PR), was recently proposed to ease automatic test pattern generation of sequential circuits. In the present paper, we propose a new PR-like method. With this method, flip-flops to be reset are selected in order to ease pseudo-random testing of sequential circuits. Moreover, the control of these reset are condensed on a reduced set of input lines. Therefore, only a small number of additional primary inputs are required. This technique has been evaluated on a large subset of the 1989 ISCAS sequential benchmark circuits and promising results were obtained.
Mohamed Soufi, Yvon Savaria, Bozena Kaminska
ISCAS2
1995 On the design of at-speed testable VLSI circuits
abstract
In this paper, a new design-for-testability technique for sequential circuits is presented. This technique may be considered as an alternative to full scan. The fault coverages obtained with this technique are comparable to those produced by full scan techniques. However, the present method improves full scan in several ways. The application test time of a device is reduced to that of applying parallel vectors at the operational speed. This characteristic of applying test vectors at the operational speed has a positive impact on the test quality. Indeed, a stuck-at test, applied at the operational speed of the circuit, identifies more defective chips than the same test applied at a lower speed. Furthermore, the timing and the area overheads, which are often considered to be serious disadvantages of DFT techniques, are in this case acceptable. With this method, all FFs are replaced with XFF gates. The XFF gate is similar to a T flip-flop without feedback. However, in some cases, when observability problems are still important, a probe observation point is inserted at the pseudo-primary inputs (PPIs).
Mohamed Soufi, Yvon Savaria, Bozena Kaminska
VTS2
1995 Producing Reliable Initialization and Test of Sequential Circuits with Pseudorandom Vectors
abstract
In this paper, the initialization of sequential circuits using pseudorandom input patterns is addressed. An extended Markov chain model that covers the initialization phase is proposed. This model supports the theoretical framework used to demonstrate that sequential circuits can be initialized with pseudorandom vectors. This leads to a uniform BIST approach in which initialization and testing are performed together with a single pseudorandom generator.>
Mohamed Soufi, Yvon Savaria, F. Darlay, Bozena Kaminska
IEEE Trans. Computers2
1995 Equivalence Proofs of Some Yield Modeling Methods for Defect-Tolerant Integrated Circuits
abstract
In this paper, two equivalence proofs of yield modeling methods for defect-tolerant integrated circuits (ICs) are presented. These proofs are generalizations of those found in Koren and Stapper (1989); one of the proofs presented in this paper is valid for any defect-tolerant IC, while the other one is valid for defect-tolerant ICs with two levels of hierarchy.>
Claude Thibeault, Yvon Savaria, Jean-Louis Houle
IEEE Trans. Computers2
1995 Bounds on the performance of partial selection networks
abstract
The evaluation of the performance of partial selection networks which select a set of M elements from a set of N inputs is addressed. The partial selection problem occurs when dealing with non-exhaustive multi-path breadth-first searches, like in the M algorithm or the bidirectional algorithm. These algorithms are used in the decoding of convolutional codes. The paper presents a set of bounds to evaluate the quality of regular, Delta class, networks of depth 1gN and width N/2, with respect to their selection capabilities. The results from the bounds are compared to Monte Carlo simulations of the selection capabilities of the Banyan and Alekseyev networks. Finally, the performance degradation associated with the use of these networks on the performance of a bidirectional decoder is presented. In particular, the authors show that even with imperfect selection, the bidirectional decoder can outperform a Viterbi decoder of comparable complexity.>
Jean Belzile, Yvon Savaria, David Haccoun, Martin Chalifoux
IEEE Trans. Commun.2
1994 Fast Convergence with Low Precision Weights in ART1 Networks
abstract
A new learning law, the Direct Coding Rule, is proposed for bottom-up long term memory learning in Adaptive Resonance Theory (ART) networks. This law requires less computational precision than the traditional Weber Law Rule and modifies the search dynamics of the network to accelerate convergence. Following a brief mathematical analysis of the new learning law, an ART1 network based on this law is applied to a passive radar detection problem. The simulation results allow comparison of the new law to the Weber Law Rule, with and without weight quantization, from the speed and cost viewpoints.>
Jean-François Crespo, Pierre Lavoie, Yvon Savaria
ISCAS3
1994 A Comparative Study of Single-Phase Clocked Latches Using Estimation Criteria
abstract
The advantage of using single-phase clocked circuits in VLSI system design is well known. This class of circuits has the advantage of simple clock distribution, low area for clock routing, reduced clock skew, and high speed. However, it is difficult to compare the characteristics and performance of these circuits, because there are no clear evaluation criteria. Indeed, such criteria are related to the application context and may therefore be misleading when taken out of context. In this paper, we will present a set of criteria which will permit designers to choose the most appropriate circuit in a particular case and therefore, obtain useful results. The proposed criteria will reduce the simulation time and help designers to reach better solutions in less design time.>
Sameh Ghannoum, Dmitri Chtchvyrkov, Yvon Savaria
ISCAS3
1994 Pseudo-Random Vector Compaction for Sequential Testability
abstract
In this paper, a pseudo-random vector compaction technique for sequential circuits is presented. This technique is based on sequential testability measures and the iterative model. The optimum number of circuit duplications is deduced from the testability analysis. The pseudo random vector compaction consists of conserving the vectors that detect faults and the n-1 previous vectors, where n is the optimum number of circuit duplications. The results indicate that fault coverage produced by 200000 pseudo-random vectors is exactly reproduced by a small set of vectors which do not exceed 1000 vectors for almost all the benchmark circuits.>
Naim Ben-Hamida, Bozena Kaminska, Yvon Savaria
ISCAS3
1994 A Fast Low-Power Driver for Long Interconnections in VLSI Systems
abstract
Propagation delays of signals on long interconnections and power consumption are very important limitations to the computing capacity of VLSI chips. Thus, this paper presents a low-power circuit for the fast propagation of signals on long interconnections in VLSI systems. Using electrical simulations with HSPICE guided by a heuristic, we present two designs that perform better in terms of speed and area than the conventional method. With both designs, power dissipation is reduced by 64%.>
Mohamed Nekili, Yvon Savaria, Guy Bois
ISCAS2
1994 A Fast CMOS Voltage-Controlled Ring Oscillator
abstract
A voltage-controlled ring oscillator based on variable capacitive loading is described. A prototype, fabricated in a 1.2 /spl mu/m CMOS-technology, has been tested. The results showed that it is capable to operate at a maximum frequency of 560 MHz. This design operates at 92% of the speed achievable by the same ring oscillator, but where the speed control circuits have been removed. An important feature of this oscillator is its tunability range, which goes from 277 to 560 MHz, or a factor of 2.0 between f/sub max/ and f/sub min/. This paper presents the simulations performed to optimize the performance of the circuit.>
Yvon Savaria, Dmitri Chtchvyrkov, John F. Currie
ISCAS1
1994 A Fast Method to Evaluate the Optimum Number of Spares in Defect-Tolerant Integrated Circuits
abstract
We present a method to accelerate the search for the number of spares to be included in defect tolerant integrated circuits. Our method is obtained by bringing two modifications to a conventional evaluation method. The main motivations behind the development of this method are: the possibilities offered by the implementation of defect tolerance, the existence of many yield models, which may predict different results in terms of optimum number of spares, and the fact that some models are very compute intensive. The modeling methods leading to several usual yield models are briefly presented. We also present results showing that our method is valid for a wide range of parameters. However, this method can be applied to all yield models considered and it can significantly reduce the time spent in the search for the best possible reconfiguration strategies.>
Claude Thibeault, Yvon Savaria, Jean-Louis Houle
IEEE Trans. Computers2
1994 A multiprocessor architecture for multiple path stack sequential decoders
abstract
The Zigangirov-Jelinek (stack) algorithm allows decoding convolutional codes with a small computational effort compared to the optimum Viterbi algorithm. However, it suffers from a variability of that computational effort that is highly undesirable. The paper describes an architecture that implements a multiple-path-like stack algorithm for reducing this variability. This architecture is organized as a linear structure comprising special processors for extending tree nodes, called extenders, and priority stacks for storing nodes in sorted metric order. The architecture is shown to have a good potential for reducing the computational variability without adding much overhead to the system. Simulations have shown that this architecture effectively reduces computational variability as the number of processors increases, even for a relatively large number of extenders. Simulations run for up to 16 extenders have also shown that using 4 to 16 extenders is a good choice. The architecture is also shown to reduce computational variability like the multiple path algorithm does, while having a better time performance.>
Normand Bélanger, David Haccoun, Yvon Savaria
IEEE Trans. Commun.3
1994 A systolic architecture for fast stack sequential decoders
abstract
The troublesome operation of reordering the stack in stack sequential decoders is alleviated by storing the nodes in a systolic priority queue that delivers the true top node in a short and constant amount of time. A new systolic priority queue is described that allows each decoding step, including retrieval, reordering and storage of the nodes, to take place in a single clock period. A complete decoder architecture designed around this queue is compared to a conventional stack-bucket architecture from both speed and cost points of view. The proposed decoder architecture appears to be faster, affordable, and compatible with convolutional codes having long memory and high coding rate.>
Pierre Lavoie, David Haccoun, Yvon Savaria
IEEE Trans. Commun.3
1994 Pipelining communications in large VLSI/ULSI systems
abstract
A simple and very effective solution to the delay incurred while propagating data through long interconnection wires is presented. Such delays can be found in large VLSI/ULSI or wafer scale systems. The basic idea of the technique relies on the fragmentation of the wires and in reconnecting them with a special device called repeater in order to form a bidirectional pipeline. A method for determining the optimum configuration of the pipeline is presented. It is shown that, even in presence of an appreciable skew in synchronous systems, the technique improves the transmission speed by 150% for 32-byte messages, when a 10 cm 8-bit bus implemented in a 1.2 /spl mu/m CMOS technology is used. The improvement increases for longer messages and for larger skews. It is also shown that the actual transmission time is close (to within a factor of 2) to the theoretical limit that could be achieved with a zero-length wire. A method based on repeaters operating at a multiple of the basic system clock frequency is also proposed. It is shown that this technique may speedup data transfer by an order of magnitude. The extension of the technique to asynchronous self-timed repeaters is also discussed. Finally, a VLSI implementation of the synchronous reconnection device is described.>
Daniel Audet, Yvon Savaria, N. Arel
IEEE Trans. Very Large Scale Integr. Syst.2
1993 Initiability: A Measure of Sequential Testability
Naim Ben-Hamida, Bozena Kaminska, Yvon Savaria
ISCAS3
1993 A High Speed Parallel Structure for the Basic Wavelet Transform Algorithm
Hakim Khali, Jean-Louis Houle, Yvon Savaria
ISCAS3
1993 Parallel Regeneration of Interconnections in VLSI & ULSI Circuits
Mohamed Nekili, Yvon Savaria
ISCAS2
1992 Test quality of hierarchical defect-tolerant integrated circuits
Claude Thibeault, Yvon Savaria, Jean-Louis Houle
J. Electron. Test.2
1992 Performance improvements to VLSI parallel systems, using dynamic concatenation of processing resources
Daniel Audet, Yvon Savaria, Jean-Louis Houle
Parallel Comput.2
1991 New VLSI architectures for fast soft-decision threshold decoders
abstract
New VLSI architectures for fast convolutional threshold decoders that process soft-quantized channel symbols are presented. The new architectures feature pipelining and parallelism and make it possible to fabricate decoders for data rates up to hundreds of Mbits per second. With these architectures, the data rate is shown to be independent of the memory of the code, implying that fast AAPP (approximate a posteriori probability) decoders can be built for long powerful codes. Furthermore, the architectures are convenient to use with low and high coding rates. Using a typical example it is shown that a soft-decision threshold decoder can provide a substantial coding gain while being less costly to implement than the hard-decision threshold decoder.>
Pierre Lavoie, David Haccoun, Yvon Savaria
IEEE Trans. Commun.3
1989 Design-for-Testability Using Test Design Yield and Decision Theory
abstract
A framework for prediction and estimation of the test yield and test cost of VLSI systems during the design-for-testability stage is given. As an important extension, the authors present a technique for evaluating the set of possible solutions and selecting the most effective one. This technique is based on the evaluation of test-related performance measures and on decision theory. As a result, a new level of design and test integration is obtained. Experimental results have confirmed the applicability and effectiveness of the method. It is shown that it is possible to derive, in a very straightforward manner, maximum and minimum expected values of the design yields of a strategy (design scheme).>
Bozena Kaminska, Yvon Savaria
ITC2
1989 A Pragmatic Approach to the Design of Self-Testing Circuits
abstract
A tool for interactively finding hard-to-test nodes and assessing the impact of test points on random pattern testability is described. With this tool, a set of test points which bring the fault coverage to almost 100% was found for the hardest ISCAS benchmark. It is shown that observation test points only are not sufficient, and thus controllable test points are required. The added test points yield test sets shorter than those necessary with weighted random test sets. On the basis of these findings, a pragmatic approach to built-in self-test (BIST) is proposed. New CMOS BILBO-like test latches that are self-tested are also proposed. A modified implementation for the force-observe (F-O) test points is proposed, and an evaluation of the relative overhead associated with both the circular scan path and the added test points is presented.>
Yvon Savaria, Bruno Laguë, Bozena Kaminska
ITC1
1988 New architectures for fast convolutional encoders and threshold decoders
abstract
Several new architectures for high-speed convolution encoders and threshold decoders are developed. In particular, it is shown that new architectures featuring both parallelism and pipelining are promising from a speed point of view. These architectures are practical for a wide range of coding rates and constant lengths. Two integrated circuits featuring these architectures have been designed and fabricated in a CMOS 3- mu m technology. The two circuits have been tested and can be used to build convolutional encoders and definite threshold decoders operating at data rates above 100 Mb/s. It is shown that with these architectures, encoders and threshold decoders could easily be designed to operate at data rates above 1 Gb/s.>
David Haccoun, Pierre Lavoie, Yvon Savaria
IEEE J. Sel. Areas Commun.3
1986 A Theory for the Design of Soft-Error-Tolerant VLSI Circuits
abstract
Soft errors caused by ionizing radiation will be a limiting factor in the reliability of VLSI circuits with submicron-feature sizes. A new approach to the design of soft-error-tolerant digital integrated circuits'is presented. It is based on the filtering of transients at register inputs, and it incurs a lower area overhead than known techniques. The method, called soft-error filtering (SEF), is derived on the basis of the analogy between a noise-sensitive finite-state machine and a noisy communication channel. The necessary characteristics of the register are examined and a design is presented for the associated filter. It is shown that SEF can be used to reduce the associated error rate to insignificant levels.
Yvon Savaria, Jeremiah F. Hayes, Nicholas C. Rumin, Vinod K. Agarwal
IEEE J. Sel. Areas Commun.1
1986 Soft-error filtering: A solution to the reliability problem of future VLSI digital circuits
abstract
As the semiconductor industry continues to scale down the feature sizes in VLSI digital circuits, soft errors will eventually limit the reliability of these circuits. An important source of these errors will be the products of radioactive decay. It is proposed to combat these transient errors by a new technique called soft-error filtering (SEF). This is based on filtering the input to every latch in the VLSI circuit, thereby preventing these transients, generated by alpha particle hits in the combinational section, from being latched in the corresponding registers. Several approaches to the problem of designing filtering latches are compared. This comparison demonstrates the superiority of a double-filter realization. The design for a CMOS implementation of the double-filter latch is presented. Not only is the design simple and efficient, but it can be expected to be tolerant to process variations. A comparison of SEF with conventional techniques for dealing with soft errors shows the former to be generally much more attractive, from the point of view of both area and time overhead.
Yvon Savaria, Nicholas C. Rumin, Jeremiah F. Hayes, Vinold K. Agarwal
Proc. IEEE1
1984 A Design for Machines with Built-In Tolerance to Soft Errors
Yvon Savaria, Vinod K. Agarwal, Nicholas C. Rumin, Jeremiah F. Hayes
ITC1