EDBT 2026 Demo / reviewers in the wild / expert
Jean-Pierre David
dblp:81/170
· DBLP profile ↗
39ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-7707-0483ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorArtificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Prototyping Framework for P4-Programmable Traffic ManagersabstractInternational audience Karl La Grassa, André Béliveau, Mathieu Léonardon, Jean-Pierre David, Matthieu Arzel, Yvon Savaria |
RSP | 4 |
| 2025 | Enabling Rank-Based P4 Programmable Schedulers: Requirements, Implementation, and Evaluation on BMv2 SwitchesabstractSoftware-defined networking (SDN) has revolutionized network infrastructure, offering programmability to meet evolving network demands. However, the fixed-function nature of the packet scheduler in current network equipment impedes the exploration of scheduling policies within a programmable network environment. This paper proposes a novel methodology to implement rank-based programmable schedulers in programmable BMv2 switches expressed with the network-specific programming language (P4). A proposed custom networking environment facilitates the study and evaluation of various scheduling policies. This environment is used to implement 20 different scheduling and shaping policies to identify the required language constructs and components needed to express these policies with the P4 language efficiently. Our experiments reveal that specific scheduling policies do not seamlessly align with a previously proposed architecture for rank-based scheduling policies. Thus, we propose rank-based versions for five previously reported scheduling policies, making them efficiently implementable in any rank-based schedulers and programmable network equipment. The reported results confirm that the rank-based versions of these scheduling policies accurately replicate the behavior and performance of the original policies, with a maximum error of 0.5% in the resulting flow completion times (FCTs). Mostafa Elbediwy, Bill Pontikakis, Jean-Pierre David, Yvon Savaria |
IEEE Trans. Netw. | 3 |
| 2024 | Temporal Logic Explanations for Dynamic Decision Systems Using Anchors and Monte Carlo Tree Search (Abstract Reprint)abstractFor many automated perception and decision tasks, state-of-the-art performance may be obtained by algorithms that are too complex for their behavior to be completely understandable or predictable by human users, e.g., because they employ large machine learning models. To integrate these algorithms into safety-critical decision and control systems, it is particularly important to develop methods that can promote trust into their decisions and help explore their failure modes. In this article, we combine the anchors methodology with Monte Carlo Tree Search to provide local model-agnostic explanations for the behaviors of a given black-box model making decisions by processing time-varying input signals. Our approach searches for descriptive explanations for these decisions in the form of properties of the input signals, expressed in Signal Temporal Logic, which are highly likely to reproduce the observed behavior. To illustrate the methodology, we apply it in simulations to the analysis of a hybrid (continuous-discrete) control system and a collision avoidance system for unmanned aircraft (ACAS Xu) implemented by a neural network. Tzu-Yi Chiu, Jerome Le Ny, Jean-Pierre David |
AAAI | 3 |
| 2024 | DR-PIFO: A Dynamic Ranking Packet Scheduler Using a Push-In-First-Out QueueabstractSoftware-defined Networking (SDN) introduced the decoupling of control and data forwarding planes. Despite advances in the programmability of SDNs, there remains a strong need for a fully programmable packet scheduler in the data plane. In this context, the ability to adapt to various traffic patterns and the expressiveness of schedulers are of paramount importance. This paper introduces the Dynamic Ranking Push-In-First-Out (DR-PIFO), as an algorithmic model that can be used to develop programmable packet schedulers based on PIFO queues. The DR-PIFO is a highly expressive model, capable of expressing a wide range of work-conserving, non-work-conserving, and hierarchical scheduling algorithms. Additionally, its dynamic ranking capabilities allow for real-time updates to the packet’s priority within the scheduler. The proposed solution also performs error detection in the departure order of packets, which is essential to avoid starvation in strict priority scheduling. These features are crucial when implementing popular scheduling algorithms such as the pFabric. The DR-PIFO is evaluated through its algorithmic properties and by implementing two distinct case studies. Its performance is further evaluated by incorporating it as an external module, written in a high-level language, and integrating it with software switches implemented using the P4 language. The results illustrate the superior expressiveness of DR-PIFO over state-of-the-art models such as PIFO and PIEO and confirm that it is an algorithm-agnostic model. Thus, DR-PIFO represents a promising solution for implementing more fully programmable packet schedulers in SDNs, with the potential to improve performance and adaptability. Mostafa Elbediwy, Bill Pontikakis, Alireza Ghaffari, Jean-Pierre David, Yvon Savaria |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2023 | BARVINN: Arbitrary Precision DNN Accelerator Controlled by a RISC-V CPUabstractWe present a DNN accelerator that allows inference at arbitrary precision with dedicated processing elements that are configurable at the bit level. Our DNN accelerator has 8 Processing Elements controlled by a RISC-V controller with a combined 8.2 TMACs of computational power when implemented with the recent Alveo U250 FPGA platform. We develop a code generator tool that ingests CNN models in ONNX format and generates an executable command stream for the RISC-V controller. We demonstrate the scalable throughput of our accelerator by running different DNN kernels and models when different quantization levels are selected. Compared to other low precision accelerators, our accelerator provides run time programmability without hardware reconfiguration and can accelerate DNNs with multiple quantization levels, regardless of the target FPGA size. BARVINN is an open source project and it is available at https://github.com/hossein1387/BARVINN. MohammadHossein AskariHemmat, Sean Wagner, Olexa Bilaniuk, Yassine Hariri, Yvon Savaria, Jean-Pierre David |
ASP-DAC | 6 |
| 2023 | Quark: An Integer RISC-V Vector Processor for Sub-Byte Quantized DNN InferenceabstractIn this paper, we present Quark, an integer RISC-V vector processor specifically tailored for sub-byte DNN inference. Quark is implemented in GlobalFoundries' 22FDX FD-SOI technology. It is designed on top of Ara, an open-source 64-bit RISC-V vector processor. To accommodate sub-byte DNN inference, Quark extends Ara by adding specialized vector instructions to perform sub-byte quantized operations. We also remove the floating-point unit from Quarks' lanes and use the CVA6 RISC-V scalar core for the re-scaling operations that are required in quantized neural network inference. This makes each lane of Quark 2 times smaller and 1.9 times more power efficient compared to the ones of Ara. In this paper we show that Quark can run quantized models at sub-byte precision. Notably we show that for 1-bit and 2-bit quantized models, Quark can accelerate computation of Conv2d over various ranges of inputs and kernel sizes. MohammadHossein AskariHemmat, Théo Dupuis, Yoan Fournier, Nizar El Zarif, Matheus A. Cavalcante, Matteo Perotti, Frank K. Gürkaynak, Luca Benini, François Leduc-Primeau, Yvon Savaria, Jean-Pierre David |
ISCAS | 11 |
| 2023 | Temporal logic explanations for dynamic decision systems using anchors and Monte Carlo Tree Search
Tzu-Yi Chiu, Jerome Le Ny, Jean-Pierre David |
Artif. Intell. | 3 |
| 2022 | An FPGA-based HW/SW Co-Verification Environment for Programmable Network DevicesabstractBugs in network devices translate to financial losses for the service providers and degrade the quality of experience for the users. Simulation tools cannot guarantee complete fault coverage as bugs can manifest at any time in live hardware. To mitigate these issues, we propose a novel hardware/software (HW/SW) co-verification tool that targets programmable dataplane network devices. The system integrates cycle-accurate software simulation with a hardware implementation. For the software simulation, open-source tools such as CocoTB and GHDL were used. The Design Under Test (DUT) and our test interfaces are embedded in programmable hardware. Data from the software can be inserted and then extracted in real-time from the input/output (I/O) ports of the DUT. To achieve this functionality the hardware design uses data insertion and extraction blocks which also support assertions. For the hardware implementation, reported experiments have been conducted on a NetFPGA-SUME platform. When a packet flows through the NetFPGA and triggers an assertion, the data present in the DUT at that time can be captured, and sent back to the simulator for further analysis and replay. Each of our design block consumes less than 1% of the available resources on the FPGA. Mengyue Su, Jean-Pierre David, Yvon Savaria, Bill Pontikakis, Thomas Luinaud |
ISCAS | 2 |
| 2021 | RISC-V Barrel Processor for Deep Neural Network AccelerationabstractThis paper presents a barrel RISC-V processor designed to control a deep neural network accelerator. Our design has a 5-stage pipeline data path with 8 hardware threads (harts). Each thread is executed under a strict round robin scheduler and is responsible for providing data and control signals to a neural network processing element (PE). Each PE is capable of arbitrary precision GEneral Matrix Vector (GEMV) operations. The execution of each thread is independent of other threads and any communication between threads are sent through shared memory via software. To reduce the area required for implementation, our processor is an implementation of the RV32I plus a set of custom CSRs for controlling the PEs. Our design passes all riscv_test written in assembly and compiled with RISC-V gcc. Our 8-hart barrel processor runs at 250 MHz with CPI of 1 and consumes 0.372W. To demonstrate the capabilities of our design, we computed a GEMV operation with an input matrix size of 8 by 128 and a weight matrix size of 128 by 128 with two-bit precision in only 16 clock cycles. MohammadHossein AskariHemmat, Olexa Bilaniuk, Sean Wagner, Yvon Savaria, Jean-Pierre David |
ISCAS | 5 |
| 2021 | Guest Editorial Special Issue on the IEEE International NEWCAS Conference 2020abstractThis Special Issue is a selection of the best articles presented at the 18th IEEE International NEWCAS Conference (NEWCAS) 2020 that was held in Montreal, Canada, on June 16–19, 2020. As an Interregional flagship conference of the IEEE Circuits and Systems Society (CASS), this conference covers a wide spectrum of topics, research, and practice in the fields of circuits and systems, and offers an international forum for exchanging ideas and results. Jean-Pierre David, Manuel J. Barragan Asian |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | RISC-V Barrel Processor for Accelerator ControlabstractHardware accelerators are important in the post-Moore’s law era of computing. To maximize performance of such accelerators, most of the logic resources should be allocated to their execution circuits, while control mechanisms should be kept small yet flexible. In this paper, we propose a barrel processor design based on the RISC-V instruction set architecture (ISA) [1]. To the best of our knowledge, this is the first implementation of a barrel RISC-V processor made public. The purpose of this processor is to concurrently control and coordinate a set of accelerator processing elements. MohammadHossein AskariHemmat, Olexa Bilaniuk, Sean Wagner, Yvon Savaria, Jean-Pierre David |
FCCM | 5 |
| 2019 | Binary Speech Features for Keyword Spotting Tasks
Alexandre Riviello, Jean-Pierre David |
INTERSPEECH | 2 |
| 2019 | Bit-Slicing FPGA Accelerator for Quantized Neural NetworksabstractDeep Neural Networks (DNNs) become the state-of-the-art in several domains such as computer vision or speech recognition. However, using DNNs for embedded applications is still strongly limited because of their complexity and the energy required to process large data sets. In this paper, we present the architecture of an accelerator for quantized neural networks and its implementation on a Nallatech 385-A7 board with an Altera Stratix V GX A7 FPGA. The accelerator's design centers around the matrix-vector product as the key primitive, and exploits bit-slicing to extract maximum performance using low-precision arithmetic. Olexa Bilaniuk, Sean Wagner, Yvon Savaria, Jean-Pierre David |
ISCAS | 4 |
| 2018 | Ultra-low latency communication channels for FPGA-based HPC cluster
Roberto Sanchez Correa, Jean-Pierre David |
Integr. | 2 |
| 2018 | Automated Synthesis of Streaming Transfer Level Hardware DesignsabstractAs modern field-programmable gate arrays (FPGA) enable high computing performance and efficiency, their programming with low-level hardware description languages is time-consuming and remains a major obstacle to their adoption. High-level synthesis compilers are able to produce register-transfer-level (RTL) designs from C/C++ algorithmic descriptions, but despite allowing significant design-time improvements, these tools are not always able to generate hardware designs that compare to handmade RTL designs. In this article, we consider synthesis from an intermediate-level (IL) language that allows the description of algorithmic state machines handling connections between streaming sources and sinks. However, the interconnection of streaming sources and sinks can lead to cyclic combinational relations, resulting in undesirable behaviors or un-synthesizable designs. We propose a functional-level methodology to automate the resolution of such cyclic relations into acyclic combinational functions. The proposed IL synthesis methodology has been applied to the design of pipelined floating-point cores. The results obtained show how the proposed IL methodology can simplify the description of pipelined architectures while enabling performances that are close to those achievable through an RTL design methodology. Marc-André Daigneault, Jean-Pierre David |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2017 | A Cache-Coherent Heterogeneous Architecture for Low Latency Real Time ApplicationsabstractThis paper proposes a generic hardware architecture for runtime acceleration of heterogeneous high performance computing (HPC) clusters. This runtime accelerator performs real time resource allocation and management of HPC systems with low latency on multiple time scales. One of the target applications is to perform the signal processing in wireless communication systems such as LTE and 5G over the cloud. A core part of this work is to develop and characterize algorithms that can distribute workloads to server blades in a balanced manner with the aim of maximizing processor utilization in computing clusters. Resources are also managed to guarantee bandwidth for data transfer between computing nodes and reserved cache memories to enable deterministic task execution. This paper shows how a workload distributed among several server blades can be scheduled at a finer time scale than what a normal software implementation would allow, in order to minimize the makespan required to complete execution of sets of tasks. A case study is conducted on the implementation of a resource allocator for the proposed platform. A 760-time acceleration factor of the resource allocation process has been achieved compared to a pure software implementation, while enabling data transfers at the nanosecond scale. It stands as a proof of concept that confirms the viability of CPU-FPGA platforms for wireless standards virtualization. Michel Gemieux, Yvon Savaria, Jean-Pierre David, Guchuan Zhu |
ISORC | 3 |
| 2017 | Low latency and division free Gauss-Jordan solver in floating point arithmetic
Jean-Pierre David |
J. Parallel Distributed Comput. | 1 |
| 2015 | BinaryConnect: Training Deep Neural Networks with binary weights during propagationsabstractDeep Neural Networks (DNN) have achieved state-of-the-art results in a wide range of tasks, with the best results obtained with large training sets and large models. In the past, GPUs enabled these breakthroughs because of their greater computational speed. In the future, faster computation at both training and test time is likely to be crucial for further progress and for consumer applications on low-power devices. As a result, there is much interest in research and development of dedicated hardware for Deep Learning (DL). Binary weights, i.e., weights which are constrained to only two possible values (e.g. -1 or 1), would bring great benefits to specialized DL hardware by replacing many multiply-accumulate operations by simple accumulations, as multipliers are the most space and power-hungry components of the digital implementation of neural networks. We introduce BinaryConnect, a method which consists in training a DNN with binary weights during the forward and backward propagations, while retaining precision of the stored weights in which gradients are accumulated. Like other dropout schemes, we show that BinaryConnect acts as regularizer and we obtain near state-of-the-art results with BinaryConnect on the permutation-invariant MNIST, CIFAR-10 and SVHN. Matthieu Courbariaux, Yoshua Bengio, Jean-Pierre David |
NIPS | 3 |
| 2013 | High-Level Description and Synthesis of Floating-Point Accumulators on FPGAabstractDecades of research in the field of high level hardware description now result in tools that are able to automatically transform C/C++ constructs into highly optimized parallel and pipelined architectures. Such approaches work fine when the control flow is a priory known since the computation results in a large dataflow graph that can be mapped into the available operators. Nevertheless, some applications have a control flow that is highly dependant on the data. This paper focuses on the hardware implementation of such applications and presents a high level synthesis methodology applied to a Hardware Description Language (HDL) in which assignments correspond to self-synchronized connections between predefined data streaming sources and sinks. A data transfer occurs over an established connection when both source and sink are ready, according to their synchronization interfaces. Founded on a high-level communicating FSM programming model, the language allows the user to describe and dynamically modify streaming architectures exploiting spatial and temporal parallelism. Our compiler attempts to maximize the number of transfers at each clock cycle and automatically fixes the potential combinatorial loops induced by the dynamic connection of dependant sources and sinks. The methodology is applied to the synthesis of a pipelined floating point accumulator using the Delayed-Buffering (DB) reduction method. The results we obtain are similar to state-of-the-art dedicated architectures but require much less design time and expertise. Marc-André Daigneault, Jean-Pierre David |
FCCM | 2 |
| 2013 | Hardware description and synthesis of control-intensive reconfigurable dataflow architectures (abstract only)abstractField-Programmable-Gate-Arrays are used increasingly to speed up applications in various fields of science. But as modern digital designs integrate hundreds of interconnected processing and memory units, the need for a higher level of abstraction to handle their descriptions is indisputable. This paper presents a beyond-RTL concurrent hardware description language that combines both Finite-State Machine (FSM) and constraint programming paradigms. At the featured level of abstraction, the user describes dynamic connections between data sources and sinks that may not always be ready to send or receive data tokens. The high-level description methodology enables a comprehensible description of behaviors such as data transfer synchronization, exclusivity, priority and constrained scheduling by the means of logical-implication rules constraining the data transfers authorizations. Dynamically connecting resources with potential combinatorial dependencies may lead to instability or deadlock. Such situations are automatically detected and fixed by the proposed compiler that generates a dedicated control-circuit optimizing the number of transfers that can be authorized at each clock cycle. The proposed design automation methodology is applied to the problem of deeply-pipelined vector reduction. A pipelined floating point accumulator and a matrix multiplication circuits are described with a few lines of code and automatically compiled into an FPGA. Results show that the synthesis results are comparable to those obtained with hand-written RTL but with much lower effort and time. Marc-André Daigneault, Jean-Pierre David |
FPGA | 2 |
| 2013 | Self-Alignment Schemes for the Implementation of Addition-Related Floating-Point OperatorsabstractAdvances in semiconductor technology brings to the market incredibly dense devices, capable of handling tens to hundreds floating-point operators on a single chip; so do the latest field programmable gate arrays (FPGAs). In order to alleviate the complexity of resorting to these devices in computationally intensive applications, this article proposes hardware schemes for the realization of addition-related floating-point operators based on the self-alignment technique (SAT). The article demonstrates that the schemes guarantee an accuracy as if summation was computed accurately in the precision of operator’s internal mantissa, then faithfully rounded to working precision. To achieve such performance, the article adopts the redundant high radix carry-save (HRCS) format for the rapid addition of wide mantissas. Implementation results show that combining the SAT and the HRCS format allows the implementation of complex operators with reduced area and latency, more so when a fused-path approach is adopted. The article also proposes a new hardware operator for performing endomorphic HRCS additions and presents a new technique for speeding up the conversion from the redundant HRCS to a conventional binary format. Tarek Ould Bachir, Jean-Pierre David |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2012 | Raising the abstraction level of HDL for control-dominant applicationsabstractAs the complexity of modern digital systems continues to increase exponentially, the need for beyondRTL design methodologies is growing as well. In this paper, we propose a high-level hardware description language that allows the user to dynamically modify and constrain the connections between data token sources and sinks. Actual transfers occur when both sources and sinks are ready to proceed, according to different predefined synchronization protocols. At this level of abstraction, both FSM programming and constraint programming paradigms are combined to enhance the user's ability to describe and exploit fine-grain parallelismin control-intensive hardware designs. The proposed hardware description methodology is applied to the description of two hardware implementations of the QuickSort algorithm, using pipelined memory and comparator components. Marc-André Daigneault, Jean-Pierre David |
FPL | 2 |
| 2012 | Two-level configuration for FPGA: A new design methodology based on a computing fabricabstractLarge FPGAs require more and more time and expertise to efficiently target custom applications. This paper presents a new methodology based on two configuration levels. At the lowest level, the architecture is fully synthesized, placed and routed by experts to implement a 2-D mesh architecture of configurable algorithmic token machines. At the highest level, the users can program those machines to implement custom processing and routing. The architecture is data driven. The operations are triggered by the arrival of operands, leading to a large and functional pipeline spread over the whole FPGA. This methodology enables the fast implementation of data processing algorithms by people who are not experts in FPGA design, while achieving higher performances than a pure software solution. Two simple examples (FIR and FFT) illustrate the proposed methodology and demonstrate how it is possible to benefit from the expertise encapsulated at low level by just configuring the high level. Another advantage of the proposed methodology is the opportunity to dynamically reconfigure the fabric very quickly to best match the computation requirements at run time. Mathieu Allard, Patrick Grogan, Yvon Savaria, Jean-Pierre David |
ISCAS | 4 |
| 2011 | Logarithmic-Time FPGA Bitstream Analysis: A Step Towards JIT Hardware CompilationabstractJust-In-Time (JIT) compilation is frequently used in software engineering to accelerate program execution. Parts of the code are translated to machine code at runtime to speedup their execution by exploiting local and dynamic information of the computation. Modern FPGAs manufactured by Xilinx allow partial and dynamic configuration. Such features make them eligible platforms for JIT hardware compilation. Nevertheless, this has not been achieved until now because the mapping between a bitstream and the programmable points inside these FPGAs is not documented. In this article, we propose a methodology to retrieve the relevant information in logarithmic time per bit by methodically using the tools distributed by Xilinx. We give a practical case study which details the analysis of a Virtex-II Pro FPGA bitstream. The mapping of CLBs, BRAMs, and multipliers has been fully determined. Thanks to this information, we have been able to prototype tools in the fields of reverse mapping FPGA bitstreams, low-level simulation, and custom place-and-route. Finally preliminary results demonstrate that a processor embedded in an FPGA can compile, place, and route arithmetic and logic expressions inside the FPGA within a few milliseconds. Etienne Bergeron, Louis-David Perron, Marc Feeley, Jean-Pierre David |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2010 | Performing Floating-Point Accumulation on a Modern FPGA in Single and Double PrecisionabstractIn this paper, we discuss the feasibility of a floating-point accumulator (FPACC) on modern high-end FPGA devices. We explore different implementation scenarios and propose new FPACC architectures for both single and double precision floating-point addends. The proposed strategies can be easily adapted to the implement a multiply-accumulator (FPMAC), with one or two rounding stages, in both single and double precision as well. All the aforementioned designs are characterized by high operating frequencies (ranging from 130 to 300 MHz) and moderate occupation area (from 300 to 800 slices) when implemented on the VC5VSX50T FPGA, an entry level Virtex 5 from Xilinx. Tarek Ould Bachir, Jean-Pierre David |
FCCM | 2 |
| 2010 | Towards 5ps resolution TDC on a dynamically reconfigurable FPGA (abstract only)abstractThis paper presents the implementation of a high resolution time-to-digital converter (TDC) on a dynamically reconfigurable FPGA. The TDC architecture is based on the Vernier method using two ring oscillators with slightly different frequencies. The proposed oscillators can be calibrated with picoseconds resolution by taking advantage of partial reconfiguration, and moreover recalibrated over time. The results obtained on a Xilinx Virtex-II Pro FPGA show that the proposed TDC implementation can achieve unprecedented resolutions (on FPGA) as low as 5ps and precisions up to 25ps. Marc-André Daigneault, Jean-Pierre David |
FPGA | 2 |
| 2008 | Hardware JIT Compilation for Off-the-Shelf Dynamically Reconfigurable FPGAs
Etienne Bergeron, Marc Feeley, Jean-Pierre David |
CC | 3 |
| 2008 | Setting up On-Line Learning Experiments: The LearningLab PlatformabstractTo carry out experiments into classrooms, in order to test hypothesis or new learning tools needs to perform recurrent complex and time-consuming tasks. We propose means for setting up Web based experiments by distinguishing the experimentation perspective and the learning perspective.This paper presents the experimentation platform we have developed for the kaleidoscope network of excellence (NoE) ldquoshared virtual laboratoryrdquo action. This platform, called ldquoLearningLabrdquo, gives a concrete expression of our proposal, and provides researchers with useful tools for each phase of an experiment. The usability of this platform has been put to the test with two experimentations about electricity learning. Jean-Michel Adam, Anne Lejeune, Sandra Michelet, Jean-Pierre David, Christian Martel |
ICALT | 4 |
| 2008 | Application Specific Instruction set processor specialized for block motion estimationabstractThis paper presents a novel application specific instruction set processor specialized for block motion estimation. The proposed architecture includes an efficient register file system in terms of data reuse and parallel processing. Performances and area costs are presented for different levels of parallelism and register file dimensions. Various FPGA implementations of the architecture are further studied in order to present the most important factors affecting performance and hardware resource utilization. The proposed instruction extension block architecture enables acceleration by 3 orders of magnitude for full-search block matching algorithms. Marc-André Daigneault, J. M. Pierre Langlois, Jean-Pierre David |
ICCD | 3 |
| 2007 | Hardware Complexity of Modular Multiplication and ExponentiationabstractLarge integer modular multiplication (MM) and modular exponentiation (ME) are the foundation of most public-key cryptosystems, specifically RSA, Diffie-Helleman, EIGamal, and the elliptic curve cryptosystems. Thus, MM algorithms have been studied widely and extensively. Most of the work is based on the well-known Montgomery multiplication method and its variants, which require standard multiplication operations. Despite their better complexity orders, Karatsuba and FFT algorithms seem to rarely be used for hardware implementation. In this paper, we review their hardware complexity and propose original implementations of MM and ME that become useful for 24-bit operators (Karatsuba algorithm) or 373-bit operators (FFT algorithm). Jean-Pierre David, Kassem Kalach, Nicolas Tittley |
IEEE Trans. Computers | 1 |
| 2006 | Expressing Workshop Scenario with Computer Independent ModelabstractThis contribution for the workshop aims to establish a link between the initial expression of the planet game scenario, viewed by a teacher, and its computable expression Jean-Pierre David, Anne Lejeune, Emmanuelle Villiot-Leclercq |
ICALT | 1 |
| 2006 | Modeling Collaborative Learning Activities on e-Learning PlatformsabstractThe scenarization of educational activities, especially those that are going to take place within e-learning platforms, has for a number of years represented a major challenge for groups working to favor the emergence of educational standards. In this paper we describe a meta-model, LDL (Learning Design Language), to formalize scenarized activities. We particularly highlight the correct adaptation of this meta-model for modeling various collaborative situations Christian Martel, Laurence Vignollet, Christine Ferraris, Jean-Pierre David, Anne Lejeune |
ICALT | 4 |
| 2006 | LDL: An Alternative EMLabstractThis paper describes the foundations of LDL and the associated infrastructure LDL. It explains why we have chosen to define this new EML instead of using an existing one like IMS-LD Christian Martel, Laurence Vignollet, Christine Ferraris, Jean-Pierre David, Anne Lejeune |
ICALT | 4 |
| 2006 | Comparing Educational Modeling Languages on a Case StudyabstractIn the field of learning design, IMS has standardized EML, the Educational Modeling Language proposed by OUNL. However, it remains to be seen whether it will be widely adopted. Moreover, evolutions of IMS-LD and competing models are proposed by several teams. The objective of this workshop is to confront the approaches (models, tools, methodologies) through modeling and implementation experiences of collaborative learning activities. To facilitate the comparison, a case study is proposed. The outcomes of this workshop will be discussed during the panel "Learning Design of Collaborative Learning Activities: languages, models and tools". Laurence Vignollet, Jean-Pierre David, Christine Ferraris, Christian Martel, Anne Lejeune |
ICALT | 2 |
| 2006 | Expressing Learning Scenarios with Computer Independent ModelsabstractOur research focuses on methodologies and models for helping teachers to become designers of their own educational scenarios in e-learning contexts. In this article, we propose a computer independent model based on three formalisms for describing a scenario initially expressed as an informal text. The aim of this contribution is to provide teachers with appropriate tools that later facilitates the translation of a learning scenario into a complete specification in view of its implementation on a learning platform Emmanuelle Villiot-Leclercq, Jean-Pierre David, Anne Lejeune |
ICALT | 2 |
| 2004 | An Intermediate Level HDL for System Level Design
Jean-Pierre David, Etienne Bergeron |
FDL | 1 |
| 2002 | An FPGA Implementation of the Linear Cryptanalysis
François Koeune, Gaël Rouvroy, François-Xavier Standaert, Jean-Jacques Quisquater, Jean-Pierre David, Jean-Didier Legat |
FPL | 5 |
| 2002 | A Cryptanalytic Time-Memory Tradeoff: First FPGA Implementation
Jean-Jacques Quisquater, François-Xavier Standaert, Gaël Rouvroy, Jean-Pierre David, Jean-Didier Legat |
FPL | 4 |
| 2001 | Implementation of very large dataflow graphs on a reconfigurable architecture for robotic applicationsabstractIn the context of parallel processing and reconfigurable computing, new high-density reconfigurable devices offer unexplored possibilities in various areas. Floating point operations, previously reserved to dedicated hardware, can now be mapped on reconfigurable machines. In this area, robotic algorithms require the computation of large dataflow graphs. Present DSP do not provide an efficient way of implementing parallel processing in such applications. This paper presents an original application of a data driven architecture applied to robotic. Target architecture is a multi-FPGA platform equivalent to 400.000 gates. Results show that current reconfigurable devices could deliver powerful computations (true 720 MFLOPs) compared to DSP. Jean-Pierre David, Tony Postiau, Paul Fisette, Jean-Didier Legat |
IPDPS | 1 |