J. M. Pierre Langlois

dblp:93/6531 · also Pierre Langlois 0001 · DBLP profile ↗
← Back
37ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-1721-2520ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Artificial intelligence and machine learning · 4Computer networks · 4Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 DyRecMul: Fast and Low-Cost Approximate Multiplier for FPGAs using Dynamic Reconfiguration
abstract
Multipliers are widely-used arithmetic operators in digital signal processing and machine learning ( ML ) circuits. Due to their relatively high complexity, they can have high latency and be a significant source of power consumption. One strategy to alleviate these limitations is to use approximate computing. This article thus introduces an original FPGA-based approximate multiplier specifically optimized for ML computations. It utilizes dynamically reconfigurable lookup table (LUT) primitives in AMD-Xilinx technology to realize the core part of the computations. The article provides an in-depth analysis of the hardware architecture, implementation outcomes, and accuracy evaluations of the multiplier proposed in INT8 precision. The article also facilitates the generalization of the proposed approximate multiplier idea to other datatypes, providing analysis and estimations for hardware cost and accuracy as a function of multiplier parameters. Implementation results on an AMD-Xilinx Kintex Ultrascale+ FPGA demonstrate remarkable savings of 64% and 67% in LUT utilization for signed multiplication and multiply-and-accumulation configurations, respectively when compared to the standard Xilinx multiplier core. Accuracy measurements on four popular deep learning (DL) benchmarks indicate a minimal average accuracy decrease of less than 0.29% during post-training deployment, with the maximum reduction staying less than 0.33%. The source code of this work is available on GitHub.
Shervin Vakili, Mobin Vaziri, Amirhossein Zarei, J. M. Pierre Langlois
ACM Trans. Reconfigurable Technol. Syst.4
2023 Symbolic Analysis for Data Plane Programs Specialization
abstract
Programmable network data planes have extended the capabilities of packet processing in network devices by allowing custom processing pipelines and agnostic packet processing. While a variety of applications can be implemented on current programmable data planes, there are significant constraints due to hardware limitations. One way to meet these constraints is by optimizing data plane programs. Program optimization can be achieved by specializing code that leverages architectural specificity or by compilation passes. In the case of programmable data planes, to respond to the varying requirements of a large set of applications, data plane programs can target different architectures. This leads to difficulties when developers want to reuse the code. One solution to that is to use compiler optimization techniques. We propose performing data plane program specialization to reduce the generated program size. To this end, we propose to specialize in programs written in P4, a Domain Specific Language (DSL) designed for specifying network data planes. The proposed method takes advantage of key aspects of the P4 language to perform a symbolic analysis on a P4 program and then partially evaluate the program to specialize it. The approach we propose is independent of the target architecture. We evaluate the specialization technique by implementing a packet deparser on an FPGA. The results demonstrate that program specialization can reduce the resource usage by a factor of 2 for various packet deparsers.
Thomas Luinaud, J. M. Pierre Langlois, Yvon Savaria
ACM Trans. Archit. Code Optim.2
2021 Design Principles for Packet Deparsers on FPGAs
abstract
The P4 language has drastically changed the networking field as it allows to quickly describe and implement new networking applications. Although a large variety of applications can be described with the P4 language, current programmable switch architectures impose significant constraints on P4 programs. To address this shortcoming, FPGAs have been explored as potential targets for P4 applications. P4 applications are described using three abstractions: a packet parser, match-action tables, and a packet deparser, which reassembles the output packet with the result of the match-action tables. While implementations of packet parsers and match-action tables on FPGAs have been widely covered in the literature, no general design principles have been presented for the packet deparser. Indeed, implementing a high-speed and efficient deparser on FPGAs remains an open issue because it requires a large amount of interconnections and the architecture must be tailored to a P4 program. As a result, in several works where a P4 application is implemented on FPGAs, the deparser consumes a significant proportion of chip resources. Hence, in this paper, we address this issue by presenting design principles for efficient and high-speed deparsers on FPGAs. As an artifact, we introduce a tool that generates an efficient vendor-agnostic deparser architecture from a P4 program.Our design has been validated and simulated with a cocotb-based framework.The resulting architecture is implemented on Xilinx Ultrascale+ FPGAs and supports a throughput of more than 200 Gbps while reducing resource usage by almost 10x compared to other solutions.
Thomas Luinaud, Jeferson Santiago da Silva, J. M. Pierre Langlois, Yvon Savaria
FPGA3
2021 CARLA: A Convolution Accelerator With a Reconfigurable and Low-Energy Architecture
abstract
Convolutional Neural Networks (CNNs) have proven to be extremely accurate for image recognition, even outperforming human recognition capability. When deployed on battery-powered mobile devices, efficient computer architectures are required to enable fast and energy-efficient computation of costly convolution operations. Despite recent advances in hardware accelerator design for CNNs, two major problems have not yet been addressed effectively, particularly when the convolution layers have highly diverse structures: (1) minimizing energy-hungry off-chip DRAM data movements; (2) maximizing the utilization factor of processing resources to perform convolutions. This work thus proposes an energy-efficient architecture equipped with several optimized dataflows to support the structural diversity of modern CNNs. The proposed approach is evaluated on convolutional layers of VGGNet-16 and ResNet-50. Results show that the architecture achieves a Processing Element (PE) utilization factor of 98% for the majority of 3×3 and 1×1 convolutional layers, while limiting latency to 396.9 ms and 92.7 ms when performing convolutional layers of VGGNet-16 and ResNet-50, respectively. In addition, the proposed architecture benefits from the structured sparsity in ResNet-50 to reduce the latency to 42.5 ms when half of the channels are pruned.
Shervin Vakili, J. M. Pierre Langlois
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 Unleashing the Power of FPGAs as Programmable Switches
abstract
The P4 language and the PISA architecture have revolutionized the field of networking. Thanks to P4 and PISA, new networking applications and protocols can be rapidly evaluated on high performance switches. While P4 allows the expression of a wide range of packet processing algorithms, current programmable switch architecture limit the overall processing flexibility. To address this shortcoming recent work have proposed to implement PISA on FPGAs. However, little effort has been devoted to analyze whether FPGAs are good candidates to implement PISA. In this work, we take a step back and evaluate the micro-architecture efficiency of various PISA blocks. Using a theoretical analysis and experiments, we demonstrate that current FPGA architecture drastically limit the performance of a few PISA blocks. Thus, we explore two avenues to alleviate these shortcomings. First, we identify some network applications that are well tailored to current FPGAs. Second, to support a wider range of networking applications, we propose modifications to the FPGA architecture which can also be of interest outside the networking field.
Thomas Luinaud, Thibaut Stimpfling, Jeferson Santiago da Silva, Yvon Savaria, J. M. Pierre Langlois
FPGA5
2020 Bridging the Gap: FPGAs as Programmable Switches
abstract
The emergence of P4, a domain specific language, coupled to PISA, a domain specific architecture, is revolutionizing the networking field. P4 allows to describe how packets are processed by a programmable data plane, spanning ASICs and CPUs, implementing PISA. Because the processing flexibility can be limited on ASICs, while the CPUs performance for networking tasks lag behind, recent works have proposed to implement PISA on FPGAs. However, little effort has been dedicated to analyze whether FPGAs are good candidates to implement PISA. In this work, we take a step back and evaluate the micro-architecture efficiency of various PISA blocks. We demonstrate, supported by a theoretical and experimental analysis, that the performance of a few PISA blocks is severely limited by the current FPGA architectures. Specifically, we show that match tables and programmable packet schedulers represent the main performance bottlenecks for FPGA-based programmable switches. Thus, we explore two avenues to alleviate these shortcomings. First, we identify network applications well tailored to current FPGAs. Second, to support a wider range of networking applications, we propose modifications to the FPGA architectures which can also be of interest out of the networking field.
Thomas Luinaud, Thibaut Stimpfling, Jeferson Santiago da Silva, Yvon Savaria, J. M. Pierre Langlois
HPSR5
2019 Module-per-Object: A Human-Driven Methodology for C++-Based High-Level Synthesis Design
abstract
High-Level Synthesis (HLS) brings FPGAs to audiences previously unfamiliar to hardware design. However, achieving the highest Quality-of-Results (QoR) with HLS is still unattainable for most programmers. This requires detailed knowledge of FPGA architecture and hardware design in order to produce FPGA-friendly codes. Moreover, these codes are normally in conflict with best coding practices, which favor code reuse, modularity, and conciseness. To overcome these limitations, we propose Module-per-Object (MpO), a human-driven HLS design methodology intended for both hardware designers and software developers with limited FPGA expertise. MpO exploits modern C++ to raise the abstraction level while improving QoR, code readability and modularity. To guide HLS designers, we present the five characteristics of MpO classes. Each characteristic exploits the power of HLS-supported modern C++ features to build C++-based hardware modules. These characteristics lead to high-quality software descriptions and efficient hardware generation. We also present a use case of MpO, where we use C++ as the intermediate language for FPGA-targeted code generation from P4, a packet processing domain specific language. The MpO methodology is evaluated using three design experiments: a packet parser, a flow-based traffic manager, and a digital up-converter. Based on experiments, we show that MpO can be comparable to handwritten VHDL code while keeping a high abstraction level, humanreadable coding style and modularity. Compared to traditional C-based HLS design, MpO leads to more efficient circuit generation, both in terms of performance and resource utilization. Also, the MpO approach notably improves software quality, augmenting parameterization while eliminating the incidence of code duplication.
Jeferson Santiago da Silva, François R. Boyer, J. M. Pierre Langlois
FCCM3
2019 SHIP: A Scalable High-Performance IPv6 Lookup Algorithm That Exploits Prefix Characteristics
abstract
Due to the emergence of new network applications, current IP lookup engines must support high bandwidth, low lookup latency, and the ongoing growth of IPv6 networks. However, the existing solutions are not designed to address jointly these three requirements. This paper introduces SHIP, an IPv6 lookup algorithm that exploits prefix characteristics to build a data structure designed to meet future application requirements. Based on the prefix length distribution and prefix density, prefixes are first clustered into groups sharing similar characteristics and then encoded in hybrid trie-trees. The resulting memory-efficient and scalable data structure can be stored in low-latency memories and allows the traversal process to be parallelized and pipelined in order to support high packet bandwidth in hardware. In addition, SHIP supports incremental updates. Evaluated on real and synthetic IPv6 prefix tables, SHIP has a logarithmic scaling factor in terms of the number of memory accesses and a linear memory consumption scaling. Compared with other well-known approaches, SHIP reduces the required amount of memory per prefix by 87%. When implemented on a state-of-the-art field-programmable gate array (FPGA), the proposed architecture can support processing 588 million packets per second.
Thibaut Stimpfling, Normand Bélanger, J. M. Pierre Langlois, Yvon Savaria
IEEE/ACM Trans. Netw.3
2018 A Low-Latency Memory-Efficient IPv6 Lookup Engine Implemented on FPGA Using High-Level Synthesis
abstract
The emergence of 5G networks and real-time applications across networks has a strong impact on the performance requirements of IP lookup engines. These engines must support not only high-bandwidth but also low-latency lookup operations. This paper presents the hardware architecture of a low-latency IPv6 lookup engine capable of supporting the bandwidth of current Ethernet links. The engine implements the SHIP lookup algorithm, which exploits prefix characteristics to build a compact and scalable data structure. The proposed hardware architecture leverages the characteristics of the data structure to support low-latency lookup operations, while making efficient use of memory. The architecture is described in C++, synthesized with a highlevel synthesis tool, then implemented on a Virtex-7 FPGA. Compared to the proposed IPv6 lookup architecture, other wellknown approaches use at least 87% more memory per prefix, while increasing the lookup latency by a factor of 2.3×.
Thibaut Stimpfling, J. M. Pierre Langlois, Normand Bélanger, Yvon Savaria
CCGrid2
2018 P4-Compatible High-Level Synthesis of Low Latency 100 Gb/s Streaming Packet Parsers in FPGAs
abstract
Packet parsing is a key step in SDN-aware devices. Packet parsers in SDN networks need to be both reconfigurable and fast, to support the evolving network protocols and the increasing multi-gigabit data rates. The combination of packet processing languages with FPGAs seems to be the perfect match for these requirements. In this work, we develop an open-source FPGA-based configurable architecture for arbitrary packet parsing to be used in SDN networks. We generate low latency and high-speed streaming packet parsers directly from a packet processing program. Our architecture is pipelined and entirely modeled using templated \textttC++ classes. The pipeline layout is derived from a parser graph that corresponds to a P4 code after a series of graph transformation rounds. The RTL code is generated from the \textttC++ description using Xilinx Vivado HLS and synthesized with Xilinx Vivado. Our architecture achieves a \SI100 \giga\bit/\second data rate in a Xilinx Virtex-7 FPGA while reducing the latency by 45% and the LUT usage by 40% compared to the state-of-the-art.
Jeferson Santiago da Silva, François R. Boyer, J. M. Pierre Langlois
FPGA3
2018 Custom Low Power Processor for Polar Decoding
abstract
Cloud Radio Access Network is foreseen as one of the key features of the future 5G mobile communication standard. In this context, all the baseband processing is intended to be performed on CPUs in order to keep a high level of flexibility. The challenge is then to propose efficient software implementations of baseband processing algorithms that guarantee a sufficient throughput, while limiting the energy consumption. In this paper, as an alternative to general purpose processors, we propose an implementation of an Application Specific Instruction set Processor customized for the Successive Cancellation decoding of polar codes. The resulting software decoder achieves throughputs similar to state-of-the-art ARM processor implementations, while reducing the energy consumption by a factor 10.
Mathieu Léonardon, Camille Leroux, David Binet, J. M. Pierre Langlois, Christophe Jégo, Yvon Savaria
ISCAS4
2018 Enhanced Bloom filter utilisation scheme for string matching using a splitting approach
abstract
Bloom filters (BFs) are widely utilised to speed up string matching in crucial network applications such as real‐time intrusion detection and spam filters. This study introduces a new approach to improve the efficiency of BFs for string matching functions. The approach splits each target string into two substrings and considers the second substring for programming the BF. The objective is to minimise the false positive rate by maximising the common hash signatures from the second substring. Results show that compared to the traditional means of using BFs, the proposed approach reduces the false positive rate by averages of 76 and 88% for 32 and 64 Kb BFs, respectively. Moreover, a complete string matching architecture has been developed in hardware based on the proposed approach. Results demonstrate the advantages of this new architecture compared to similar previous works.
Shervin Vakili, J. M. Pierre Langlois, Yvon Savaria, Naraig Manjikian
IET Commun.2
2018 Explicit Ringing Removal in Image Deblurring
abstract
In this paper, we present a simple yet effective image deblurring method to produce ringing-free deblurred images. Our work is inspired by the observation that large-scale deblurring ringing artifacts are measurable through a multi-resolution pyramid of low-pass filtering of the blurred-deblurred image pair. We propose to model such a quantification as a convex cost function and minimize it directly in the deblurring process in order to reduce ringing regardless of its cause. An efficient primal-dual algorithm is proposed as a solution to this optimization problem. Since the regularization is more biased toward ringing patterns, the details of the reconstructed image are prevented from over-smoothing. An inevitable source of ringing is sensor saturation which can be detected costlessly contrary to most other sources of ringing. However, dealing with the saturation effect in deblurring introduces a non-linear operator in optimization problem. In this paper, we also introduce a linear approximation as a solution to handling saturation in the proposed deblurring method. As a result of these steps, we significantly enhance the quality of the deblurred images. Experimental results and quantitative evaluations demonstrate that the proposed method performs favorably against state-of-the-art image deblurring methods.
Ali Mosleh 0002, Yasser Elmi Sola, Farzad Zargari, Emmanuel Onzon, J. M. Pierre Langlois
IEEE Trans. Image Process.5
2017 A Configurable FPGA Implementation of the Tanh Function Using DCT Interpolation
abstract
Efficient implementation of non-linear activation functions is essential to the implementation of deep learning models on FPGAs. We introduce such an implementation based on the Discrete Cosine Transform Interpolation Filter (DCTIF). The proposed interpolation architecture combines simple arithmetic operations on the stored samples of the hyperbolic tangent function and on input data. It achieves almost 3× better precision than previous works while using a similar amount computational resources and a small amount of memory. Various combinations of DCTIF parameters can be chosen to trade off the accuracy and the overall circuit complexity of the tanh function. In one case, the proposed architecture approximates the hyperbolic tangent activation function with 0.004 maximum error while requiring only 1.45 kbits BRAM memory and 21 LUTs of a Virtex-7 FPGA.
Ahmed M. Abdelsalam, J. M. Pierre Langlois, Farida Cheriet
FCCM2
2017 Accurate and Efficient Hyperbolic Tangent Activation Function on FPGA using the DCT Interpolation Filter (Abstract Only)
Ahmed M. Abdelsalam, J. M. Pierre Langlois, Farida Cheriet
FPGA2
2017 An FPGA Overlay Architecture for Cost Effective Regular Expression Search (Abstract Only)
Thomas Luinaud, Yvon Savaria, J. M. Pierre Langlois
FPGA3
2017 An FPGA Coarse Grained Intermediate Fabric for Regular Expression Search
abstract
Deep Packet Inspection systems such as Snort and Bro express complex rules with regular expressions. In Snort, the search of a regular expression is performed with a Non-deterministic Finite Automaton (NFA). Traversing an NFA sequentially with a CPU is not deterministic in time, and it can be very time consuming. The sequential traversal of an NFA with a CPU is not deterministic in time consequently it can be time consuming. A fully parallel NFA implemented in hardware can search all rules, but most of the time only a small part is active. Furthermore, a string filter determines the traversal of an NFA. This paper proposes an FPGA Intermediate Fabric that can efficiently search regular expressions. The architecture is configured for a specific NFA based on a partial match of a rule found by the string filter. It can thus support all rules from a set such as Snort, while significantly reduce compute resources and power con-sumption compared to a fully parallel implementation. Multiple parameters can be selected to find the best tradeoff between resource consumption and the number and types of supported expressions. This architecture was implemented on a Xilinx R XC7VX1140 Virtex-7. The reported implementation, can sustain up to 512 regular expressions, while requiring 2% of the slices and 16% of the BRAM resources, for a throughput of 200 million characters per second.
Thomas Luinaud, Yvon Savaria, J. M. Pierre Langlois
ACM Great Lakes Symposium on VLSI3
2017 Scalable memory-less architecture for string matching with FPGAs
abstract
String matching hardware engines generally utilize Ternary Content Addressable Memories (TCAMs). Although TCAM-based solutions are fast, they are expensive and power hungry. This paper proposes a high-performance memory-less architecture for string matching called Split-Bucket. It offers a performance comparable to TCAM-based solutions. Moreover, it is reconfigurable and scalable to the size of the target string set and the width of the string. The architecture is characterized using the Longest Prefix Match problem for IP address lookup and is implemented on a Virtex-7 FPGA. For a real-world routing table with 524 k IPv4 prefixes, the Split-Bucket architecture achieves a throughput of 103.4 M packets per second and consumes 23% and 22% of the Look Up Tables and Flip-Flops of a Xilinx XC7V2000T chip, respectively.
Ideh Sarbishei, Shervin Vakili, J. M. Pierre Langlois, Yvon Savaria
ISCAS3
2016 Node configuration for the Aho-Corasick algorithm in Intrusion Detection Systems
abstract
In this paper, we analyze the performance and cost trade-off from selecting two representations of nodes when implementing the Aho-Corasick algorithm. This algorithm can be used for pattern matching in network-based intrusion detection systems such as Snort. Our analysis uses the Snort 2.9.7 rules set, which contains almost 26k patterns. Our methodology consists of code profiling and analysis, followed by the selection of a parameter to maximize a metric that combines clock cycles count and memory usage. The parameter determines which of two types of nodes is selected for each trie node. We show that it is possible to select the parameter to optimize the metric, which results in an improvement by up to 12× compared with the single node-type case.
Alexsandre B. Lacroix, J. M. Pierre Langlois, François R. Boyer, Antoine Gosselin, Guy Bois
ANCS2
2016 Memory-Efficient String Matching for Intrusion Detection Systems using a High-Precision Pattern Grouping Algorithm
abstract
The increasing complexity of cyber-attacks necessitates the design of more efficient hardware architectures for real-time Intrusion Detection Systems (IDSs). String matching is the main performance-demanding component of an IDS. An effective technique to design high-performance string matching engines is to partition the target set of strings into multiple subgroups and to use a parallel string matching hardware unit for each subgroup. This paper introduces a novel pattern grouping algorithm for heterogeneous bit-split string matching architectures. The proposed algorithm presents a reliable method to estimate the correlation between strings. The correlation factors are then used to find a preferred group for each string in a seed growing approach. Experimental results demonstrate that the proposed algorithm achieves an average of 41% reduction in memory consumption compared to the best existing approach found in the literature, while offering orders of magnitude faster execution time compared to an exhaustive search.
Shervin Vakili, J. M. Pierre Langlois, Bochra Boughzala, Yvon Savaria
ANCS2
2016 Red Lesion Detection Using Dynamic Shape Features for Diabetic Retinopathy Screening
abstract
The development of an automatic telemedicine system for computer-aided screening and grading of diabetic retinopathy depends on reliable detection of retinal lesions in fundus images. In this paper, a novel method for automatic detection of both microaneurysms and hemorrhages in color fundus images is described and validated. The main contribution is a new set of shape features, called Dynamic Shape Features, that do not require precise segmentation of the regions to be classified. These features represent the evolution of the shape during image flooding and allow to discriminate between lesions and vessel segments. The method is validated per-lesion and per-image using six databases, four of which are publicly available. It proves to be robust with respect to variability in image resolution, quality and acquisition system. On the Retinopathy Online Challenge's database, the method achieves a FROC score of 0.420 which ranks it fourth. On the Messidor database, when detecting images with diabetic retinopathy, the proposed method achieves an area under the ROC curve of 0.899, comparable to the score of human experts, and it outperforms state-of-the-art approaches.
Lama Séoud, Thomas Hurtut, Jihed Chelbi, Farida Cheriet, J. M. Pierre Langlois
IEEE Trans. Medical Imaging5
2015 Camera intrinsic blur kernel estimation: A reliable framework
abstract
This paper presents a reliable non-blind method to measure intrinsic lens blur. We first introduce an accurate camera-scene alignment framework that avoids erroneous homography estimation and camera tone curve estimation. This alignment is used to generate a sharp correspondence of a target pattern captured by the camera. Second, we introduce a Point Spread Function (PSF) estimation approach where information about the frequency spectrum of the target image is taken into account. As a result of these steps and the ability to use multiple target images in this framework, we achieve a PSF estimation method robust against noise and suitable for mobile devices. Experimental results show that the proposed method results in PSFs with more than 10 dB higher accuracy in noisy conditions compared with the PSFs generated using state-of-the-art techniques.
Ali Mosleh 0002, Paul Green 0002, Emmanuel Onzon, Isabelle Bégin, J. M. Pierre Langlois
CVPR5
2014 Image Deconvolution Ringing Artifact Detection and Removal via PSF Frequency Analysis
Ali Mosleh 0002, J. M. Pierre Langlois, Paul Green 0002
ECCV (4)2
2014 A computationally efficient importance sampling tracking algorithm
Rana Farah, Qifeng Gan, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria
Mach. Vis. Appl.3
2013 Finite-precision error modeling using affine arithmetic
abstract
This paper introduces a new approach for finite-precision error modeling based on affine arithmetic. The paper demonstrates that there is a common hazard in affine arithmetic-based error modeling methods described in the literature. The hazard is linked to early substitution of the signal terms that emerge in operations such as multiplication and division. The paper proposes postponed substitution combined with function maximization to address this problem. The paper also proposes a modification in the error propagation process to enhance the error modeling accuracy. An existing word length optimization method is reproduced to evaluate the efficiency of this modification. The results demonstrate that the proposed modification can improve the hardware area results by up to 7.0% at the expense of negligible complexity overhead.
Shervin Vakili, J. M. Pierre Langlois, Guy Bois
ICASSP2
2013 Enhanced Precision Analysis for Accuracy-Aware Bit-Width Optimization Using Affine Arithmetic
abstract
Bit-width allocation has a crucial impact on hardware efficiency and accuracy of fixed-point arithmetic circuits. This paper introduces a new accuracy-guaranteed word-length optimization approach for feed-forward fixed-point designs. This method uses affine arithmetic, which is a well-known analytical technique, for both range and precision analyses. This paper introduces an acceleration technique and two new semianalytical algorithms for precision analysis. While the first algorithm follows a progressive search strategy, the second one uses a tree-shaped search method for fractional width optimization. The algorithms offer two different time-complexity/cost efficiency tradeoffs. The first algorithm has polynomial complexity and achieves comparable results with existing heuristic approaches. The second algorithm has exponential complexity, but it achieves near-optimal results compared to the exhaustive search method. A commonly used set of case studies is used to evaluate the efficiency of the proposed techniques and algorithms in terms of optimization time and hardware cost. The first and second algorithms achieve 10.9% and 13.1% improvements in area, respectively, over uniform fractional width allocation. The proposed acceleration technique reduces the complexity of the fractional width selection problem by an average of 20.3%.
Shervin Vakili, J. M. Pierre Langlois, Guy Bois
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 Catching a Rat by Its Edglets
abstract
Computer vision is a noninvasive method for monitoring laboratory animals. In this article, we propose a robust tracking method that is capable of extracting a rodent from a frame under uncontrolled normal laboratory conditions. The method consists of two steps. First, a sliding window combines three features to coarsely track the animal. Then, it uses the edglets of the rodent to adjust the tracked region to the animal's boundary. The method achieves an average tracking error that is smaller than a representative state-of-the-art method.
Rana Farah, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau
IEEE Trans. Image Process.2
2012 Body temperature estimation of a moving subject from thermographic images
Guillaume-Alexandre Bilodeau, Atousa Torabi, Maxime Levesque, Charles Ouellet, J. M. Pierre Langlois, Pablo Lema, Lionel Carmant
Mach. Vis. Appl.5
2012 Real-Time Computation of Local Neighborhood Functions in Application-Specific Instruction-Set Processors
abstract
This paper presents a systematic approach to the design of application-specific instruction-set processors for high speed computation of local neighborhood functions and intra-field deinterlacing. The intended application is real-time processing of high definition video. The approach aims at an efficient utilization of the available memory bandwidth by fully exploiting the data parallelism inherent to the target algorithm class. An appropriate choice of custom instructions and application-specific registers is used together with a very long instruction word architecture in order to mimic a pipelined systolic array. This leads to a processing speed close to the limit imposed by memory bandwidth constraints. For three intra-field deinterlacing algorithms and 2-D convolution with four kernel sizes, the design approach yields speedup factors between 36 and 1330, Area-Time (AT) product improvements between 12× and 243×, and energy consumption reduction factors between 13 and 262.
P. Aubertin, J. M. Pierre Langlois, Yvon Savaria
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Comparative analysis of contrast enhancement algorithms in surveillance imaging
abstract
Image contrast enhancement methods play a key role in many image processing and vision applications. For surveillance applications, real-time contrast improvement over the whole image is required when videos are taken in poor lighting conditions. It is also necessary to highlight details in shadowed regions without introducing artifacts. In this paper, several state-of-the-art contrast enhancement methods are compared. Image quality is evaluated by means of objective metrics such as intensity contrast and brightness error, and by subjective assessment. Execution time is also measured. Experimental results show that a technique based on histogram modification presents a better trade-off considering both aspects.
Diana Carolina Gil, Rana Farah, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria
ISCAS3
2011 Combining ISA extensions and subsetting for improved ASIP performance and cost
abstract
This paper presents a fine-grained configurable processor model used to generate image processing Application Specific Instruction Set Processors (ASIPs). A methodology to develop a minimal instruction set ASIP with the processor model is also proposed. The methodology is based on using specialized instructions in conjunction with Instruction Set Architecture (ISA) subsetting to reduce hardware costs and improve execution time. The performance of an FPGA implementation of the proposed processor model is measured for a two-dimensional Gaussian filter and results are compared to a popular commercial soft core processor. With ISA subsetting and specialized instructions, the proposed processor uses up to 45% fewer slices while achieving a 1.57× speedup.
Simon Rajotte, Diana Carolina Gil, J. M. Pierre Langlois
ISCAS3
2008 Application Specific Instruction set processor specialized for block motion estimation
abstract
This paper presents a novel application specific instruction set processor specialized for block motion estimation. The proposed architecture includes an efficient register file system in terms of data reuse and parallel processing. Performances and area costs are presented for different levels of parallelism and register file dimensions. Various FPGA implementations of the architecture are further studied in order to present the most important factors affecting performance and hardware resource utilization. The proposed instruction extension block architecture enables acceleration by 3 orders of magnitude for full-search block matching algorithms.
Marc-André Daigneault, J. M. Pierre Langlois, Jean-Pierre David
ICCD2
2008 Acceleration of a 3D target tracking algorithm using an application specific instruction set processor
abstract
In todaypsilas high-tech world, intelligent video-surveillance is becoming a part of everyday life. In addition to minimizing the need for constant monitoring by an operator, it can automatically perform tasks such as accident detection or estimation of vehicle speed. A particularly useful algorithm for video surveillance is three-dimensional target tracking but, since it is both quite computationally expensive and requires the use of two cameras, it is seldom used. In this paper, we concentrate on accelerating an implementation of 3D tracking using a multiprocessor ASIP architecture based on the Tensilica Xtensa processor. Our experiments show that a speedup factor of 22 can be achieved using an extensible platform expressly optimized for this application as opposed to using a general-purpose processor.
Sebastien Fontaine, Sylvain Goyette, J. M. Pierre Langlois, Guy Bois
ICCD3
2008 Efficient FPGA implementation of complex multipliers using the logarithmic number system
abstract
In many real-time DSP applications, high performance is a prime target. However, achieving this may be done at the expense of area, power dissipation and accuracy. Attempts have been made to use alternative number systems to optimize the realization of arithmetic blocks, maintaining high performance without incurring prohibitive area and power increases. This paper presents the FPGA implementation of complex multipliers based on the logarithmic number system. Synthesis results show that a design with a 10-stage pipeline can achieve a maximum clock rate of 224 MHz and 140 MHz for 16-bit and 32-bit designs, respectively. Both designs use the lowest amount of hardware in terms of gate equivalents as compared to a complex multiplier built with regular FPGA features. In particular, the proposed architecture uses 67% and 35% fewer gates to implement a 32-bit and 16-bit complex multiplier, respectively, when compared to a design realized with embedded multipliers. Simulation results based on selected test vectors show that the greatest relative error of the logarithmic-based 16-bit complex multiplier is 2.14%
Man Yan Kong, J. M. Pierre Langlois, Dhamin Al-Khalili
ISCAS2
2007 FPGA-Based Efficient Design Approach for Large-Size Two's Complement Squarers
abstract
This paper presents an optimized design approach of two's complement large-size squarers using embedded multipliers in FPGAs. The realization is based on Baugh-Wooley's algorithm, which partitions the multiplication into unsigned and signed sections. To achieve efficient implementation, a set of optimized schemes for the realization of multi-level additions of the partial products is proposed. Our approach has been evaluated through the implementation of squarers for operands with sizes ranging from 20 to 128 bits. The designs are synthesized and implemented on Xilinx' Spartan-3 with ISE 8.1 design platform and compared with the standard implementation, and with Xilinx' IP Core. The results indicate that our approach offers substantial LUT savings by up to 52% with an average delay reduction of 13%. The usage of the number of embedded multipliers is reduced by 38% compared with the standard schemes.
Shuli Gao, Noureddine Chabini, Dhamin Al-Khalili, J. M. Pierre Langlois
ASAP4
2006 An Optimized Design Approach for Squaring Large Integers Using Embedded Hardwired Multipliers
abstract
This paper presents an efficient design methodology and a systematic approach for the implementation of squaring functions with large integers, using small-size embedded multipliers. A general architecture of the squarer and a set of equations are derived to aid in the realization. The inputs of the squarer are split into several segments leading to an efficient utilization of the small-size embedded multipliers and reduced number of required addition operations. Various benchmarks were tested for different segments ranging from 2 to 5 targeting Xilinx Spartan-3 FPGA. The synthesis was performed with the aid of the Xilinx ISE 7.1 XST tool. Our approach was compared with the traditional technique using the same tool. The results illustrate that our design approach is very efficient in terms of both timing and area saving. The combinational delay is reduced by an average of 15.8%, and the area saving is about 50 % in terms of number of slices and number of 4-input LUTs. Also, the number of required embedded multipliers is reduced by an average of 32.3% compared to the traditional technique.
Shuli Gao, Noureddine Chabini, Dhamin Al-Khalili, J. M. Pierre Langlois
AICCSA4
2002 A low power direct digital frequency synthesizer with 60 dBc spectral purity
abstract
We present a low-power sine-output Direct Digital Frequency Synthesizer (DDFS) realized in 0.18 μm CMOS that achieves 60 dBc spectral purity from DC to the Nyquist frequency. No ROM or multipliers are used, but an external DAC is required if an analog output is desired. Power consumption is 10 mW for a 100 MHz clock, which is significantly less than figures reported previously. System complexity is greatly reduced by using an efficient linear interpolation scheme to approximate a sinusoid function. This has resulted in silicon area utilization of 0.011 mm2. The design would be suitable as an IP core in a low power digital RF transceiver ASIC.
J. M. Pierre Langlois, Dhamin Al-Khalili
ACM Great Lakes Symposium on VLSI1