Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Pascal Urard

dblp:45/5880 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 1 since 2021Computer networks · 3Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Electronic design automation · 90% Integrated circuit design · 10%
Theoretical computer science
1 paper
Coding theory · 100%
Computer networks
1 paper
Physical-layer communications · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › system-level design
electronic system level design
0.122006
Building a standard ESL design and verification methodology: is it just a dream? · DAC 2006
ESL: building the bridge between systems to silicon · DAC 2005
Coding theory › error-correcting codes
LDPC codes
0.112010
Low-complexity decoding for non-binary LDPC codes in high order fields · IEEE Trans. Commun. 2010
Coding theory › error-correcting codes › LDPC codes
non-binary LDPC decoding
0.112010
Low-complexity decoding for non-binary LDPC codes in high order fields · IEEE Trans. Commun. 2010
Electronic design automation › hardware verification and test › formal verification
equivalence checking
0.112008
Leveraging sequential equivalence checking to enable system-level to RTL flows · DAC 2008
Electronic design automation
hardware verification and test
0.112008
Leveraging sequential equivalence checking to enable system-level to RTL flows · DAC 2008
Electronic design automation › hardware verification and test › formal verification
sequential equivalence checking
0.112008
Leveraging sequential equivalence checking to enable system-level to RTL flows · DAC 2008
Integrated circuit design
digital circuit design
0.112005
A 135Mbps DVB-S2 compliant codec based on 64800-bit LDPC and BCH codes (ISSCC paper 24.3) · DAC 2005
Electronic design automation
system-level design
0.112005
ESL: building the bridge between systems to silicon · DAC 2005
Electronic design automation
high-level synthesis
0.012008
Leveraging sequential equivalence checking to enable system-level to RTL flows · DAC 2008
Electronic design automation › hardware verification and test
hardware verification
0.012006
Building a standard ESL design and verification methodology: is it just a dream? · DAC 2006
Physical-layer communications › channel coding › error control coding › block codes
BCH codes
0.012005
A 135Mbps DVB-S2 compliant codec based on 64800-bit LDPC and BCH codes (ISSCC paper 24.3) · DAC 2005
Physical-layer communications
channel coding
0.012005
A 135Mbps DVB-S2 compliant codec based on 64800-bit LDPC and BCH codes (ISSCC paper 24.3) · DAC 2005
Physical-layer communications › channel coding › error control coding › block codes
LDPC codes
0.012005
A 135Mbps DVB-S2 compliant codec based on 64800-bit LDPC and BCH codes (ISSCC paper 24.3) · DAC 2005

Methods — techniques the papers use, named apart from their topics

extended min-sum decoding · 0.1density evolution · 0.1low-leakage CMOS design · 0.1sequential equivalence checking · 0.1
YearPublicationVenuePosition
2024 GemIMC: A Configurable HW Architecture for Technology Agnostic IMC Based NN Inference
abstract
This paper presents GemIMC, a High Level Synthesis (HLS) based configurable digital unit architecture for accelerating Neural Networks (NN) at the edge using In Memory Computing (IMC). The proposed architecture is capable of supporting any type of memory technology, be it CMOS or resitive, with digital or analog storage capabilities. GemIMC aims to facilitate design space exploration among different IMC-related parameters and provide a top architecture for prototyping. By taking as input the IMC-tile parameters such as storage type (analog/digital), size, latency, power and process variability, GemIMC provides a an estimation of the full system's latency, area, and power consumption, taking into account the full digital control. The Tensorfiow-to-GemIMC environment flow allows for direct evaluation of the ineference accuracy for any type of NN, considering the variability associated with analog computation (when needed). The proposed work provides a flexible and efficient solution for IMC-based NN inference on the edge, along with a methodology for writing the IMC tile model to ensure compatibility with the top architecture.
Emilien Taly, Roberto Guizzetti, Pascal Urard, Elena I. Vatajelu
VLSI-SoC3
2023 Quantization Modes for Neural Network Inference: ASIC Implementation Trade-offs
abstract
As deep neural networks migrate close to the sensors, accuracy cannot be the single target anymore: inference tasks must also be highly energy efficient. For embedded devices, the power budget for one inference is typically in the range of a few tens of µW to single-digit mW. We have three levers of action for that: computational workload, number of values to memorize-be they network parameters or intermediate activation results-, and implementation strategy. Given the fact that Application Specific Integrated Circuits are two orders of magnitude more power-efficient than processors for a given technology node, the latter issue is solved using ad-hoc hardware implementations. For the two first former issues, we detail and compare in this work different existing quantization approaches, since reducing the number of bits of the weights and activations reduces computation complexity and storage needs. In addition, we also propose two new modes specifically aiming at low silicon footprint and power optimized hardware implementations that still provide an accuracy in par with existing works. We report the area/power and accuracy trade-offs all theses quantization modes provide when targeting low to ultra-low power devices. The evaluation is done using STMicroelectronics 40nm technology. It shows that the best results vary depending on the dataset and network architecture, which calls for application and quantization aware network architecture search.
Nathan Bain, Roberto Guizzetti, Emilien Taly, Ali Oudrhiri, Bruno Paille, Pascal Urard, Frédéric Pétrot
IJCNN6
2023 Performance Modeling and Estimation of a Configurable Output Stationary Neural Network Accelerator
abstract
Neural network accelerators are designed to process Neural Networks (NN) optimizing three Key Performance Indicators (KPIs): latency, power, and chip area. This work is based on the study of Gemini, an industrial prototype near memory computing inference accelerator designed using a high-level synthesis technique. Gemini is an output stationary configurable accelerator that achieves its performance based on two structural parameters. The measurement of the KPIs requires simulations that are time-consuming and resource-intensive. This paper presents a high-level practical estimator that can instantly predict the KPIs depending on the NN and the Gemini configuration. The latency is accurately derived using an analytical model based on the architecture, the operators scheduling and the NN characteristics. The power and the chip area are computed analytically and the models are calibrated using simulations. Finally, we show how to use the estimator to derive Pareto optima for choosing the best Gemini configurations for a VGG-like NN.
Ali Oudrhiri, Emilien Taly, Nathan Bain, Alix Munier Kordon, Roberto Guizzetti, Pascal Urard
SBAC-PAD6
2015 GreenNet: An Energy-Harvesting IP-Enabled Wireless Sensor Network
abstract
This paper presents GreenNet, an energy efficient and fully operational protocol stack for IP-enabled wireless sensor networks based on the IEEE 802.15.4 beacon-enabled mode. The stack runs on a hardware platform with photovoltaic cell energy harvesting developed by STMicroelectronics (STM) that can operate autonomously for long periods of time. GreenNet integrates several standard mechanisms and enhances existing protocols, which results in an operational platform with the performance beyond the current state of the art. In particular, it includes the IEEE 802.15.4 beacon-enabled medium access control (MAC) integrated with lightweight IP routing for achieving very low duty cycles. It offers an advanced discovery scheme that accelerates the process of joining the network and proposes an adaptation scheme for adjusting the duty cycle of harvested nodes to the available energy for increased performance. Finally, it supports security at two levels: a basic standard secure operation at the link layer and advanced scalable data payload security. This paper describes all techniques and mechanisms for saving energy and operating at very low duty cycles. It also provides an evaluation of the performance and energy consumption of GreenNet.
Liviu-Octavian Varga, Gabriele Romaniello, Malisa Vucinic, Michel Favre, Andrei Banciu, Roberto Guizzetti, Christophe Planat, Pascal Urard, Martin Heusse, Franck Rousseau, Olivier Alphand, Etienne Dublé, Andrzej Duda
IEEE Internet Things J.8
2010 A new approach for minimizing buffer capacities with throughput constraint for embedded system design
abstract
The design of streaming applications (e.g. multimedia or network packet processing) must consider several optimizations such as the minimization of the whole surface of the memory needed on a Chip. The problem tackled in this paper is the minimization of the whole surface of the memory needed to reach a minimum fixed throughput. The application is modelled using a Marked Timed Weighted Event Graphs (in short MTWEG), which is a subclass of Petri nets. Transitions correspond to specific treatments and places model buffers for data transfers. It is assumed that transitions are periodically fired with a fixed throughput. The problem is first mathematically modelled using an Integer Linear Program. We then study for a unique buffer the optimum throughput according to its capacity. A polynomial simple algorithm that minimizes the overall surface of memory for a fixed throughput is derived when there is no circuit in the initial MTWEG, which corresponds to a wide class of applications. We prove in this case that the capacities of every buffer may be optimized independently. For general MTWEG, the problem is NP-Hard and an original polynomial 2-approximation algorithm is presented. For practical applications, the solution computed is very close to the optimum.
Mohamed Benazouz, Olivier Marchetti, Alix Munier Kordon, Pascal Urard
AICCSA4
2010 Static Address Generation Easing: a design methodology for parallel interleaver architectures
abstract
For high throughput applications, turbo-like iterative decoders are implemented with parallel architectures. However, to be efficient parallel architectures require to avoid collision accesses i.e. concurrent read/write accesses should not target the same memory block. This consideration applies to the two main classes of turbo-like codes which are Low Density Parity Check (LDPC) and Turbo-Codes. In this paper we propose a methodology which finds a collision-free mapping of the variables in the memory banks and which optimizes the resulting interleaving architecture. Finally, we show through a pedagogical example the interest of our approach compared to state-of-the-art techniques.
Cyrille Chavet, Philippe Coussy, Pascal Urard, Eric Martin 0001
ICASSP3
2010 Low-complexity decoding for non-binary LDPC codes in high order fields
abstract
In this paper, we propose a new implementation of the Extended Min-Sum (EMS) decoder for non-binary LDPC codes. A particularity of the new algorithm is that it takes into accounts the memory problem of the non-binary LDPC decoders, together with a significant complexity reduction per decoding iteration. The key feature of our decoder is to truncate the vector messages of the decoder to a limited number nmof values in order to reduce the memory requirements. Using the truncated messages, we propose an efficient implementation of the EMS decoder which reduces the order of complexity to ¿(nmlog2nm). This complexity starts to be reasonable enough to compete with binary decoders. The performance of the low complexity algorithm with proper compensation is quite good with respect to the important complexity reduction, which is shown both with a simulated density evolution approach and actual simulations.
Adrian Voicila, David Declercq, François Verdier, Marc P. C. Fossorier, Pascal Urard
IEEE Trans. Commun.5
2008 Leveraging sequential equivalence checking to enable system-level to RTL flows
abstract
It has long been the practice to create models in C or C++ for architectural studies, software prototyping and RTL verification in the design of Systems-on-Chip (SoC). It is often the case that by the end of a design project, multiple C models exist for different uses. Since a lot of time is invested in ensuring the functional correctness of these models via their use in system-level simulations, they often become "golden" functional reference models. Design teams are moving towards leveraging these system-level models to reduce the time needed for design and verification of RTL. On the design side, the use of high-level synthesis tools to synthesize RTL from C/C++ models is gaining ground for certain classes of blocks within a design. On the verification front, temporal differences at interfaces and in internal states between system-level models and RTL prevent the use of combinational equivalence checkers. This paper focuses on the use of sequential equivalence checking to verify functional equivalence between system-level models and RTL and describes the challenges and vale of using it in system-level to RTL flows.
Pascal Urard, Asma Maalej, Roberto Guizzetti, Nitin Chawla
DAC1
2008 Split non-binary LDPC codes
abstract
In this paper, we propose and study a new family of error-correcting codes. These achieve excellent error performance under an iterative decoding over the binary-input noisy channel and solves the memory space requirements problem of the non-binary LDPC decoders. We named this class of codes, Split non-binary LDPC codes. The main particularity of this new family of codes is that the variable and the check nodes are not defined over the same finite field GF(2p), like in the case of classical non-binary LDPC codes. The class of Split non-binary LDPC codes is obviously larger than that of existing types of codes, which gives more degrees of freedom to find good codes when the existing codes show their limits. We provide two examples of interesting split NB-LDPC codes.
Adrian Voicila, David Declercq, François Verdier, Marc P. C. Fossorier, Pascal Urard
ISIT5
2007 A design methodology for space-time adapter
abstract
This paper presents a solution to efficiently explore the design space of communication adapters. In most digital signal processing (DSP) applications, the overall architecture of the system is significantly affected by communication architecture, so the designers need specifically optimized adapters. By explicitly modeling these communications within an effective graph-theoretic model and analysis framework, we automatically generate an optimized architecture, named Space-Time AdapteR (STAR). Our design flow inputs a C description of Input/Output data scheduling, and user requirements (throughput, latency, parallelism&), and formalizes communication constraints through a Resource Constraints Graph (RCG). The RCG properties enable an efficient architecture space exploration in order to synthesize a STAR component. The proposed approach has been tested to design an industrial data mixing block example: an Ultra-Wideband interleaver.
Cyrille Chavet, Philippe Coussy, Pascal Urard, Eric Martin 0001
ACM Great Lakes Symposium on VLSI3
2007 Low-Complexity, Low-Memory EMS Algorithm for Non-Binary LDPC Codes
abstract
In this paper, we propose a new implementation of the EMS decoder for non binary LDPC codes presented in (D. Declencq and M. Fossorier, 2007). A particularity of the new algorithm is that it takes into accounts the memory problem of the non binary LDPC decoders, together with a significant complexity reduction per decoding iteration. The key feature of our decoder is to truncate the vector messages of the decoder to a limited number nm of values in order to reduce the memory requirements. Using the truncated messages, we propose an efficient implementation of the EMS decoder which reduces the order of complexity to O(nmlog2nm), which starts to be reasonable enough to compete with binary decoders. The performance of the low complexity algorithm with proper compensation are quite good with respect to the important complexity reduction, which is shown both with a simulated density evolution approach and actual FER simulations.
Adrian Voicila, David Declercq, François Verdier, Marc P. C. Fossorier, Pascal Urard
ICC5
2007 A design flow dedicated to multi-mode architectures for DSP applications
abstract
This paper addresses the design of multi-mode architectures for digital signal processing applications. We present a dedicated design flow and its associated high-level synthesis tool, named GAUT. Given a unified description of a set of time-wise mutually exclusive tasks and their associated throughput constraints, a single RTL hardware architecture optimized in area is generated. In order to reduce the register, steering logic (multiplexers) and controller (decoding logic) complexities, we propose a joint-scheduling algorithm which maximizes the similarities between control steps and specific binding approaches for both functional units and storage elements which maximize the similarities between the datapaths. We show through a set of test cases that our approach offers significant area saving relative to the state-of-the-art.
Cyrille Chavet, Caaliph Andriamisaina, Philippe Coussy, Emmanuel Casseau, Emmanuel Juin, Pascal Urard, Eric Martin 0001
ICCAD6
2007 A Methodology for Efficient Space-Time Adapter Design Space Exploration: A Case Study of an Ultra Wide Band Interleaver
abstract
This paper presents a solution to efficiently explore the design space of communication adapters. In most digital signal processing (DSP) applications, the overall architecture of the system is significantly affected by communication architecture, so the designers need specifically optimized adapters. By explicitly modeling these communications within an effective graph-theoretic model and analysis framework, we automatically generate an optimized architecture, named Space-Time AdapteR (STAR). Our design flow inputs a C description of Input/Output data scheduling, and user requirements (throughput, latency, parallelism...), and formalizes communication constraints through a Resource Constraints Graph (RCG). The RCG properties enable an efficient architecture space exploration in order to synthesize a STAR component. The proposed approach has been tested to design an industrial data mixing block example: an Ultra-Wideband interleaver.
Cyrille Chavet, Philippe Coussy, Pascal Urard, Eric Martin 0001
ISCAS3
2006 Building a standard ESL design and verification methodology: is it just a dream?
Anoosh Hosseini, Ashish Parikh, H. T. Chin, Pascal Urard, Emil F. Girczyc, S. Bloch
DAC4
2005 IP-block-based design environment for high-throughput VLSI dedicated digital signal processing systems
abstract
The Growing requirement on the correct design of a high performance DSP system in short time force us to use IP's in many design. In this paper, we propose an efficient IP block based design environment for high throughput VLSI Systems. The flow generates SystemC Register Transfer Level (RTL) architecture, starting from a Matlab functional model described as a netlist of functional IP. The refinement process inserts automatically control structures to treat delays induced by the use of RTL IPs. It also inserts a control structure to coordinate the execution of parallel clocked IP. The delays may be managed by registers or by counters included in the control structure. The experimentations show that the approach can produce efficient RTL architecture and allow a huge save of time.
Nacer-Eddine Zergainoh, Katalin Popovici, Ahmed Amine Jerraya, Pascal Urard
ASP-DAC4
2005 ESL: building the bridge between systems to silicon
abstract
Electronic System-Level design has arrived - but can ESL provide the bridge from systems to silicon? Comprised of real world designers, this DAC ESL panel will examine and debate what works, what doesn't, and what the gaps are in the methodology and tool offerings. Panelists from a variety of industry segments, including Military/aerospace, storage area networks (SAN), wireless communications and consumer electronics, will share their experiences, lessons learned and further needs.Does ESL bridge the gap between systems to silicon? Hear from designers about their real world experience with ESL. What worked according to expectations? What didn't? What are the gaps in the methodology and tool offerings that need to be filled, and why?This panel of ESL design methodology users will give us a reality check that will enable potential users to make an adoption decision, and enable ESL design tool suppliers to evaluate their product strategies against big picture requirements.
Francine Bacchini, David Maliniak, Terry Doherty, Peter McShane, Suhas A. Pai, Sriram Sundararajan, Soo-Kwan Eo, Pascal Urard
DAC8
2005 A 135Mbps DVB-S2 compliant codec based on 64800-bit LDPC and BCH codes (ISSCC paper 24.3)
abstract
A DVB-S2 compliant codec is implemented in both 130nm-8M and 90nm-7M low-leakage CMOS technologies. The system includes encoders and decoders for both Low-Density Parity Check (LDPC) codes and serially concatenated BCH codes. All requirements of the DVB-S2 standard are supported including code rates between 1/4 and 9/10, block sizes of either 16,200 bits or 64,800 bits, and four digital modulation options. The 130nm core design occupies 49.6mm2 and operates at 200MHz, while the 90nm core design occupies 15.8mm2 and operates at 300MHz.
Pascal Urard, L. Paumier, P. Georgelin, T. Michel, V. Lebars, E. Yeo
DAC1