VLDB 2026 Research / reviewers in the wild / expert
Daniel Ménard
dblp:14/1517
· DBLP profile ↗
64ranked-venue papers
5as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 16 since 2021Systems, architecture and hardware · 24 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Partition Tree Search Acceleration for VVC: Survey and Evaluation with VTM EvolutionabstractVVC achieves up to 50% bit-rate savings over HEVC at the cost of increased encoding complexity, largely due to the QTMTT partitioning structure. This work surveys the evolution of the VVC Test Model (VTM) and evaluates partitioning acceleration techniques considering changes in complexity and internal heuristics across VTM versions. M. E. A. Kherchouche, Franck Galpin, Thierry Dumas, Daniel Ménard |
DCC | 5 |
| 2025 | Complexity Reduction Study Based on RD Costs Approximation for VVC Intra PartitioningabstractThis paper presents a comparison study of two machine-learning techniques to accelerate the Versatile Video Coding (VVC) intra-partitioning process in QuadTree (QT) configuration: a regression model for predicting Rate-Distortion (RD) costs and a Deep Q-Network (DQN)-based Reinforcement Learning (RL) approach that models partitioning as a Markov Decision Process (MDP). Both methods are size-independent and utilize neighboring RD costs and threshold values to optimize the splits of Coding Units (CUs). M. E. A. Kherchouche, Franck Galpin, Thierry Dumas, F. Schnitzler, Daniel Ménard |
DCC | 5 |
| 2025 | Stratified sampling: fast estimation of quantization effects on DNNabstractDeep neural networks complexity has exploded in recent years. This explosion brought new challenges in terms of memory, execution time and power requirements. One way of meeting these challenges is to use finite precision. However, this solution may result in a degradation in output quality. This degradation needs to be estimated, balancing between time and confidence in the estimation.This paper proposes a parametric method to estimate the degradation caused by finite precision in data processing oriented applications such as deep learning. This method aims to reduce the estimation time and energy requirements while maintaining confidence in the results. This method takes advantage of a priori information about the inputs to select more informative inputs. A method to obtain this a priori information is also proposed. The results obtained are similar to a simple degradation estimation with a time reduction of one order of magnitude. Quentin Milot, Mickaël Dardaillon, Daniel Ménard |
DDECS | 3 |
| 2025 | Can LLMs Revolutionize the Design of Explainable and Efficient TinyML Models?abstractThis paper introduces a novel framework for designing efficient neural network architectures specifically tailored to tiny machine learning (TinyML) platforms. By leveraging large language models (LLMs) for neural architecture search (NAS), a vision transformer (ViT)-based knowledge distillation (KD) strategy, and an explainability module, the approach strikes an optimal balance between accuracy, computational efficiency, and memory usage. The LLM-guided search explores a hierarchical search space, refining candidate architectures through Pareto optimization based on accuracy, multiply-accumulate operations (MACs), and memory metrics. The best-performing architectures are further fine-tuned using logits-based KD with a pre-trained ViT-B/16 model, which enhances generalization without increasing model size. Evaluated on the CIFAR-100 dataset and deployed on an STM32H7 microcontroller (MCU), the three proposed models, LMaNet-Elite, LMaNet-Core, and QwNet-Core, achieve accuracy scores of 74.50%, 74.20% and 73.00%, respectively. All three models surpass current state-of-the-art (SOTA) models, such as MCUNet-in3/in4 (69.62% / 72.86%) and XiNet (72.27%), while maintaining a low computational cost of less than 100 million MACs and adhering to the stringent 320 KB static random-access memory (SRAM) constraint. These results demonstrate the efficiency and performance of the proposed framework for TinyML platforms, underscoring the potential of combining LLM-driven search, Pareto optimization, KD, and explainability to develop accurate, efficient, and interpretable models. This approach opens new possibilities in NAS, enabling the design of efficient architectures specifically suited for TinyML. To facilitate further research and development in this field, the proposed framework and the best-performing architectures are made publicly available at Link. Christophe El Zeinaty, Wassim Hamidouche, Glenn Herrou, Daniel Ménard, Mérouane Debbah |
IJCNN | 4 |
| 2025 | Use or Produce - Carbon Impact of a Video Streaming DeviceabstractThis paper presents a detailed carbon-centered Life Cycle Analysis (LCA) of a video streaming device designed for long-term operation. We chose a development board with a screen to have full control on its operation, for which we have detailed information on its components allowing us to accurately model emissions from the production to the use of the device. Concerning its use, we perform a dedicated measurement series to determine the actual power consumption in different video playback scenarios and develop a linear power estimation model, which we use to evaluate different usage scenarios. The resulting LCA indicates that potential savings are highest when exploiting low-power sleep modes or switching off the device during its lifetime. When operating the device in a country with low carbon intensity, production and usage show similar emissions, while in countries with medium to high carbon emissions, the usage of the device causes significantly higher emissions. Pierre Le Gargasson, Olivier Weppe, Thibaut Marty, Maxime Pelcat, Daniel Ménard, Christian Herglotz |
ISCAS | 5 |
| 2025 | Style-FG: A Style-based Framework for Film Grain Analysis and SynthesisabstractFilm grain which used to be a by-product of the chemical processing in the analog film stock is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this article, we use a deep learning-based framework for film grain analysis, generation, and synthesis. Our framework Style-FG consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community. 1 Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for. To contribute further to the sustainability necessary effort of the digital information and communication field, a light-weight version of Style-FG is also proposed, which demonstrates similar quantitative and qualitative performances, while reducing the number of network parameters by a factor of 92%. Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | 3R-INN: How to Be Climate Friendly While Consuming/Delivering Videos?
Zoubida Ameur, Claire-Hélène Demarty, Daniel Ménard, Olivier Le Meur |
ECCV (75) | 3 |
| 2024 | RD-cost Regression Speed Up Technique for VVC Intra Block PartitioningabstractThe last standard Versatile Video Codec (VVC) aims to improve the compression efficiency by saving around 50% of bitrate at the same quality compared to its predecessor High Efficiency Video Codec (HEVC). However, this comes with higher encoding complexity mainly due to a much larger number of block splits to be tested on the encoder side. This paper proposes an acceleration of the VVC partitioning based on a multi-output regression model that predicts a suitable split mode for 32 × 32 Coding Unit (CU). Experimental results show that our approach improves complexity trade-offs flexibility while adding better complexity trade-offs points compared to the original encoder. M. E. A. Kherchouche, Franck Galpin, Thierry Dumas, Daniel Ménard |
ICASSP | 4 |
| 2024 | Dicetrack: Lightweight Dice Classification on Resource-Constrained Platforms with Optimized Deep Learning ModelsabstractThis paper introduces DiceTrack, an innovative Deep Learning (DL) platform for detecting dice in board games. Deploying robust models on microcontrollers (MCUs) presents challenges due to memory and computational constraints. We focus on optimizing MobileNet for seamless ESP32 deployment and propose two novel ultra-lightweight models, Separable Convolutional Layers with Quantization Network (SCLQNet) and Binarized Neural Network (BNNet), for dice classification. SCLQNet uses separable convolutional layers with quantized weights and activations, while BNNet employs a unique Binarized architecture. Further, we create DiceVision, a custom dice classification dataset tailored for real-time digital board games. Comprehensive evaluations on ESP32 and Raspberry Pi 4 showcase the efficiency of the proposed models. SCLQNet and BNNet achieve 97.5% and 97.4% accuracy with 32KB and 22KB model sizes. Notably, SCLQNet takes only 402 ms for ESP32 inference, enabling 88% latency reduction compared to quantized MobileNet. Christophe El Zeinaty, Glenn Herrou, Wassim Hamidouche, Daniel Ménard |
ICASSP | 4 |
| 2023 | Live and low energy VVC Video Decoding powered by the OpenVVC Decoder on ARM PlatformabstractThis demonstration showcases the potential of open-source software implementation for the new versatile video coding (VVC) standard, OpenVVC. The most complex VVC tools were optimized for ARM-type architectures using data parallelism through SIMD instructions. The demonstration has been tested on the NVIDIA Jetson AGX Xavier and on a NVIDIA SHIELD Android TV which showcased real-time decoding for FHD and HD video resolutions. By combining extensive data level parallelism with frame level parallelism, OpenVVC is able to maintain a remarkably low memory and energy consumption while achieving real-time decoding of videos with high resolution. These features present a great advantage for its integration on embedded devices with low computing and memory resources. Ibrahim Farhat, Pierre-Loup Cabarat, Wassim Hamidouche, Patrice Angot, Philippe Gonon, Daniel Ménard |
ISCAS | 6 |
| 2023 | Decoding Time Prediction for Versatile Video CodingabstractEnergy consumption in video decoding is a complex interplay of codec efficiency, hardware design, software optimization, and other factors such as the processor clock frequency for software implementation. The Dynamic voltage and frequency scaling (DVFS) allows adjusting dynamically the processor clock frequency according to estimated frame requirements, enabling efficient resource management for decoding tasks. However, decoding time can vary considerably from one frame to another due to variations in frame complexity. Building upon this concept, this paper proposes a machine-learning model that estimates the decoding time of individual frames. The proposed model is built on the ExtraTrees regressor that accurately predicts the decoding time of 1080p video frames with a low relative error of 5.58% and a high R2 score of 94%. Our proposal entails the utilization of frame-related data that is readily reachable by the decoder, making it highly suitable for real-time scenarios. Hafssa Boujida, Pierre-Loup Cabarat, Daniel Ménard |
MMSP | 3 |
| 2023 | Energy Efficient VVC Decoding on Mobile PlatformabstractRecently, global demand for high-resolution videos and new multimedia applications have created the need for a new video coding standard. Hence, in July 2020 the Versatile Video Coding (VVC) standard was released providing up to 40% bit-rate saving for the same video quality compared to its predecessor High Efficiency Video Coding (HEVC). However, this bit-rate saving comes at the cost of high computational complexity, particularly for live applications and on resource-constraint embedded devices. This paper presents an power-efficient VVC decoder implementation designed for low-resource platforms. This latter exploits optimization techniques such as data level parallelism using Single Instruction Multiple Data (SIMD) instructions and functional level parallelism using frame, tile and slice-based parallelisms. The results showed that the OpenVVC decoder achieve real-time decoding of Full High Definition (FHD) resolution at 30 fps targeting a platform with 8 cores with a maximum frequency of 2.2 Ghz and High Definition (HD) real-time decoding at 30 fps for platforms using 4 cores with a maximum frequency of 1.8 Ghz. In terms of average consumed power, OpenVVC showed around 5.6 watts and 1.7 watts for the 8 and 4 cores platforms, respectively. In addition, it comes with the best trade-off between what is achievable in real-time and the power consumed in comparison to the state-of-the-art implementation. Ibrahim Farhat, Pierre-Loup Cabarat, Daniel Ménard, Wassim Hamidouche, Olivier Déforges |
MMSP | 3 |
| 2023 | Energy-Aware HDR Content End-to-End VVC EncodingabstractThe proposed demonstration combines several commercially viable state-of-the-art technologies to achieve significant energy reduction while delivering HDR video without compromising QoE. The components of the demo are: (i) The demo is based our solution on VVC, the most efficient, highest quality video codec available today. (ii) We leverage Advanced HDR by Technicolor, a solution that uniquely can deliver HDR and SDR video using a single stream. (iii) We introduce a solution for dynamically adjusting peak brightness of displays via display adaptation controlled through the network using dynamic metadata encapsulated in MPEG SEI messages. This solution has been tested and shows a capability of reducing power by 10% or more without any subjective quality loss, and far more with minimal subjective impact. (iv) While we demonstrate a streaming (unicast) scenario in our live demo, our system is also capable of working in a dynamic unicast/multicast environment based on DVB transmission standards. (v) The entire demonstration is resulting from an industrial collaboration and is based on ready to market technologies and prototypes. Olivier Le Meur, Franck Aumont, Juan-Carlos Vargas, Pierre-Loup Cabarat, Daniel Ménard, Oussama Hammami, Thomas Guionnet |
MMSP | 5 |
| 2023 | Style-based film grain analysis and synthesisabstractFilm grain which used to be a by-product of the chemical processing in the analog film stock, is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this paper, we use a deep learning-based framework for film grain analysis, generation and synthesis. Our framework consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community1. Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for. Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard, Edouard François |
MMSys | 4 |
| 2023 | Open-Source Toolkit for Live End-to-End 4K VVC Intra CodingabstractVersatile Video Coding (VVC/H.266) takes video coding to the next level by doubling the coding efficiency over its predecessors for the same subjective quality, but at the cost of immense coding complexity. Therefore, VVC calls for aggressively optimized codecs to make it feasible for live streaming media applications. This paper introduces the first public end-to-end (E2E) pipeline for live 4K30p VVC intra coding and streaming. The pipeline is made up of three open-source components: 1) uvg266 for VVC encoding; 2) uvgRTP for VVC streaming; and 3) OpenVVC for VVC decoding. The proposed setup is demonstrated with a proof-of-concept prototype that implements the encoder end on AMD ThreadRipper 2990WX and the decoder end on Nvidia Jetson AGX Orin. Our prototype is almost 34 000 times as fast as the corresponding E2E pipeline built around the VTM codec. Respectively, it achieves 3.3 times speedup without any significant coding overhead over the pipeline that utilizes the fastest possible configuration of the well-known VVenC/VVdeC codec. These results indicate that our prototype is currently the only viable open-source solution for live 4K VVC intra coding and streaming. Marko Viitanen, Joose Sainio, Alexandre Mercat, Guillaume Gautier, Jarno Vanne, Ibrahim Farhat, Pierre-Loup Cabarat, Wassim Hamidouche, Daniel Ménard |
MMSys | 9 |
| 2023 | Complexity assessment of the intra prediction in Versatile Video Coding
Naima Zouidi, Amina Kessentini, Wassim Hamidouche, Nouri Masmoudi, Daniel Ménard |
Multim. Tools Appl. | 5 |
| 2023 | Machine Learning Based Efficient QT-MTT Partitioning Scheme for VVC Intra EncodersabstractThe next-generation Versatile Video Coding (VVC) standard introduces a new Multi-Type Tree (MTT) block partitioning structure that supports Binary-Tree (BT) and Ternary-Tree (TT) splits in both vertical and horizontal directions. This new approach leads to five possible splits at each block depth. It thereby improves the coding efficiency of VVC over that of the preceding High Efficiency Video Coding (HEVC) standard, which only supports Quad-Tree (QT) partitioning with a single split per block depth. However, MTT also has brought a considerable impact on encoder computational complexity. This paper proposes a two-stage learning-based technique to tackle the complexity overhead of MTT in VVC intra encoders. In our scheme, the input block is first processed by a Convolutional Neural Network (CNN) to predict its spatial features through a vector of probabilities describing the partition at each$4\times 4$edge. Subsequently, a Decision Tree (DT) model leverages this vector of spatial features to predict the most likely splits at each block. Finally, based on this prediction, only the$N$most likely splits are processed by the Rate-Distortion (RD) process of the encoder. In order to train our CNN and DT models on a wide range of image contents, we also propose a public VVC frame partitioning dataset based on existing image dataset encoded with the VVC reference software encoder. Our solution relying on the top-3 configuration reaches 47.4% complexity reduction for a negligible bitrate increase of 0.79%. A top-2 configuration enables a higher complexity reduction of 70.4% for 2.49% bitrate loss. These results emphasize a better trade-off between VTM intra-coding efficiency and complexity reduction compared to the state-of-the-art solutions. The source code of the proposed method and the training dataset are made publicly available at GitHub. Alexandre Tissier, Wassim Hamidouche, Souhaiel Belhadj Dit Mdalsi, Jarno Vanne, Franck Galpin, Daniel Ménard |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Deep-Based Film Grain Removal and SynthesisabstractIn this paper, deep learning-based techniques for film grain removal and synthesis that can be applied in video coding are proposed. Film grain is inherent in analog film content because of the physical process of capturing images and video on film. It can also be present in digital content where it is purposely added to reflect the era of analog film and to evoke certain emotions in the viewer or enhance the perceived quality. In the context of video coding, the random nature of film grain makes it both difficult to preserve and very expensive to compress. To better preserve it while compressing the content efficiently, film grain is removed and modeled before video encoding and then restored after video decoding. In this paper, a film grain removal model based on an encoder-decoder architecture and a film grain synthesis model based on a conditional generative adversarial network (cGAN) are proposed. Both models are trained on a large dataset of pairs of clean (grain-free) and grainy images. Quantitative and qualitative evaluations of the developed solutions were conducted and showed that the proposed film grain removal model is effective in filtering film grain at different intensity levels using two configurations: 1) a non-blind configuration where the film grain level of the grainy input is known and provided as input; and 2) a blind configuration where the film grain level is unknown. As for the film grain synthesis task, the experimental results show that the proposed model is able to reproduce realistic film grain with a controllable intensity level specified as input. Zoubida Ameur, Wassim Hamidouche, Edouard François, Milos Radosavljevic, Daniel Ménard, Claire-Hélène Demarty |
IEEE Trans. Image Process. | 5 |
| 2022 | Design Space Exploration for Memory-Oriented Approximate Computing TechniquesabstractModern digital systems are processing more and more data. This increase in memory requirements must match the processing capabilities and interconnections to avoid the memory wall. Approximate computing techniques exist to alleviate these requirements but usually require a thorough and tedious analysis of the processing pipeline. This paper presents an application-agnostic Design Space Exploration (DSE) of the buffer-sizing process to reduce the memory footprint of applications while guaranteeing an output quality above a defined threshold. The proposed DSE selects the appropriate bit-width and storage type for buffers to satisfy the constraint. We show in this paper that the proposed DSE reduces the memory footprint of the SqueezeNet CNN by 58.6% with identical Top-1 prediction accuracy, and the full SKA SDP pipeline by 39.7% without degradation, while only testing for a subset of the design space. The proposed DSE is fast enough to be integrated into the design stream of applications. Hugo Miomandre, Jean-François Nezan, Daniel Ménard |
ASAP | 3 |
| 2022 | Machine Learning Based Efficient Qt-Mtt Partitioning for VVC Inter CodingabstractThe Joint Video Experts Team (JVET) have standardized the Versatile Video Coding (VVC) in 2020 targeting efficient coding of the emerging video services and formats such as 8K and immersive video streaming applications. VVC standard enhances the coding efficiency by 40% at the cost of an encoder computational complexity increase estimated to 859%(x8) compared to the previous standard High Efficiency Video Coding (HEVC). This work aims at reducing the complexity of the VVC encoder under the Random Access (RA) configuration. The proposed method takes advantage of the inter prediction in order to predict the split probabilities through a convolutional neural network. Our solution reaches 31.8% of complexity reduction for a negligible bitrate increase of 1.11% outperforming state-of-the-art methods. Alexandre Tissier, Wassim Hamidouche, Jarno Vanne, Daniel Ménard |
ICIP | 4 |
| 2022 | Efficient HW Design of Adaptive Loop Filter for 4k ASIC VVC EncoderabstractVersatile Video Coding (VVC) is the next-generation video coding standard released in July 2020. VVC introduces new coding tools enhancing the coding efficiency compared to its predecessor, High Efficiency Video Coding (HEVC). These new tools significantly impact the VVC software and hardware implementations with a complexity estimated to two times and eight times the HEVC decoder and encoder complexity, respectively. In particular, the Adaptive Loop Filter (ALF), adopted in VVC as an in-loop filter, increases both the run time complexity and memory usage. These concerns need to be carefully addressed regarding the design of a VVC hardware encoder. In this paper, we present an efficient hardware implementation of the ALF tool with its decision process in the context of a professional VVC encoder. The proposed solution can reach a real-time encoding of 4K resolution videos with 4:2:2 chroma sub-sampling at 60 frames per second targeting ASIC platforms with 28-nm technology. Ibrahim Farhat, Wassim Hamidouche, Adrien Grill, Daniel Ménard, Olivier Déforges |
PCS | 4 |
| 2021 | HEVC hardware vs software decoding: An objective energy consumption analysis and comparison
Mohammed Bey Ahmed Khernache, Yahia Benmoussa, Jalil Boukhobza, Daniel Ménard |
J. Syst. Archit. | 4 |
| 2020 | Fast Kriging-based Error Evaluation for Approximate Computing SystemsabstractApproximate computing techniques trade-off the performance of an application for its accuracy. The challenge when implementing approximate computing in an application is to efficiently evaluate the quality at the output of the application to optimize the noise budgeting of the different approximation sources. It is commonly achieved with an optimization algorithm to minimize the implementation cost of the application subject to a quality constraint. During the optimization process, numerous approximation configurations are tested, and the quality at the output of the application is measured for each configuration with simulations. The optimization process is a time-consuming task. We propose a new method for infering the accuracy or quality metric at the output of an application using kriging, a geostatistical method. Justine Bonnot, Daniel Ménard, Karol Desnos |
DATE | 2 |
| 2020 | Lightweight Hardware Implementation of VVC Transform Block for ASIC DecoderabstractVersatile Video Coding (VVC) is the next generation video coding standard expected by the end of 2020. Compared to its predecessor, VVC introduces new coding tools to make compression more efficient at the expense of higher computational complexity. This rises a need to design an efficient and optimised implementation especially for embedded platforms with limited memory and logic resources. One of the newly introduced tools in VVC is the Multiple Transform Selection (MTS). This latter involves three Discrete Cosine Transform (DCT)/Discrete Sine Transform (DST) types with larger and rectangular transform blocks. In this paper, an efficient hardware implementation of all DCT/DST transform types and sizes is proposed. The proposed design uses 32 multipliers in a pipelined architecture which targets an ASIC platform. It consists in a multi-standard architecture that supports the transform block of recent MPEG standards including AVC, HEVC and VVC. The architecture is optimized and removes unnecessary complexities found in other proposed architectures by using regular multipliers instead of multiple constant multipliers. The synthesized results show that the proposed method which sustain a constant throughput of two pixels/cycle and constant latency for all block sizes can reach an operational frequency of 600 Mhz enabling to decode in real-time 4K videos at 48 fps. Ibrahim Farhat, Wassim Hamidouche, Adrien Grill, Daniel Ménard, Olivier Déforges |
ICASSP | 4 |
| 2020 | Quality-Driven Dynamic VVC Frame Partitioning for Efficient Parallel ProcessingabstractVVC is the next generation video coding standard, offering coding capability beyond HEVC standard. The high computational complexity of the latest video coding standards requires high-level parallelism techniques, in order to achieve real-time and low latency encoding and decoding. HEVC and VVC include tile grid partitioning that allows to process simultaneously rectangular regions of a frame with independent threads. The tile grid may be further partitioned into a horizontal sub-grid of Rectangular Slices (RSs), increasing the partitioning flexibility. The dynamic Tile and Rectangular Slice (TRS) partitioning solution proposed in this paper benefits from this flexibility. The TRS partitioning is carried-out at the frame level, taking into account both spatial texture of the content and encoding times of previously encoded frames. The proposed solution searches the best partitioning configuration that minimizes the trade-off between multi-thread encoding time and encoding quality loss. Experiments prove that the proposed solution, compared to uniform TRS partitioning, significantly decreases multi-thread encoding time, with slightly better encoding quality. Thomas Amestoy, Wassim Hamidouche, Cyril Bergeron, Daniel Ménard |
ICIP | 4 |
| 2020 | CNN Oriented Complexity Reduction Of VVC Intra EncoderabstractThe Joint Video Expert Team (JVET) is currently developing the next-generation MPEG/ITU video coding standard called Versatile Video Coding (VVC) and their ultimate goal is to double the coding efficiency over the state-of-the-art HEVC standard.The latest version of the VVC reference encoder, VTM6.1, is able to improve the intra coding efficiency by 24 % over the HEVC reference encoder HM16.20, but at the expense of 27 times the encoding time. The complexity overhead of VVC primarily stems from its novel block partitioning scheme that complements Quad-Tree (QT) split with Multi-Type Tree (MTT) partitioning in order to better fit the local variations of the video signal. This work reduces the block partitioning complexity of VTM6.1 through the use of Convolutional Neural Networks (CNNs). For each 64 × 64 Coding Unit (CU), the CNN is trained to predict a probability vector that speeds up coding block partitioning in encoding. Our solution is shown to decrease the intra encoding complexity of VTM6.1 by 51.5% with a bitrate increase of only 1.45%. Alexandre Tissier, Wassim Hamidouche, Jarno Vanne, Franck Galpin, Daniel Ménard |
ICIP | 5 |
| 2020 | Software HEVC video decoder: towards an energy saving for mobile applications
Naty Ould Sidaty, Julien Heulot, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
Multim. Tools Appl. | 5 |
| 2020 | Tunable VVC Frame Partitioning Based on Lightweight Machine LearningabstractBlock partition structure is a critical module in video coding scheme to achieve significant gap of compression performance. Under the exploration of the future video coding standard, named Versatile Video Coding (VVC), a new Quad Tree Binary Tree (QTBT) block partition structure has been introduced. In addition to the QT block partitioning defined in High Efficiency Video Coding (HEVC) standard, new horizontal and vertical BT partitions are enabled, which drastically increases the encoding time compared to HEVC. In this paper, we propose a lightweight and tunable QTBT partitioning scheme based on a Machine Learning (ML) approach. The proposed solution uses Random Forest classifiers to determine for each coding block the most probable partition modes. To minimize the encoding loss induced by misclassification, risk intervals for classifier decisions are introduced in the proposed solution. By varying the size of risk intervals, tunable trade-off between encoding complexity reduction and coding loss is achieved. The proposed solution implemented in the JEM-7.0 software offers encoding complexity reductions ranging from 30average for only 0.7% to 3.0% Bjxntegaard Delta Rate (BDBR) increase in Random Access (RA) coding configuration, with very slight overhead induced by Random Forest. The proposed solution based on Random Forest classifiers is also efficient to reduce the complexity of the Multi-Type Tree (MTT) partitioning scheme under the VTM-5.0 software, with complexity reductions ranging from 25% to 61% in average for only 0.4% to 2.2% BD-BR increase. Thomas Amestoy, Alexandre Mercat, Wassim Hamidouche, Daniel Ménard, Cyril Bergeron |
IEEE Trans. Image Process. | 4 |
| 2019 | Random Forest Oriented Fast QTBT Frame PartitioningabstractBlock partition structure is a critical module in video coding scheme to achieve significant gap of compression performance. Under the exploration of future video coding standard by the Joint Video Exploration Team (JVET), named Versatile Video Coding (VVC), a new Quad Tree Binary Tree (QTBT) block partition structure has been introduced. In addition to the QT block partitioning defined by High Efficiency Video Coding (HEVC) standard, new horizontal and vertical BT partitions are enabled, which drastically increases the encoding time compared to HEVC. In this paper, we propose a fast QTBT partitioning scheme based on a Machine Learning approach. Complementary to techniques proposed in literature to reduce the complexity of HEVC Quad Tree (QT) partitioning, the propose solution uses Random Forest classifiers to determine for each block which partition modes between QT and BT is more likely to be selected. Using uncertainty zones of classifier decisions, the proposed complexity reduction technique is able to reduce in average by 30% the encoding time of JEM-v7.0 software in Random Access configuration with only 0.57% Bjøntegaard Delta Rate (BD-BR) increase. Thomas Amestoy, Alexandre Mercat, Wassim Hamidouche, Cyril Bergeron, Daniel Ménard |
ICASSP | 5 |
| 2019 | Accuracy Evaluation Based on Simulation for Finite Precision Systems Using Inferential StatisticsabstractThe conversion of an algorithm to fixed-point arithmetic is commonly achieved with a large and fixed-number of simulations. Nevertheless, when simulating a fixed and arbitrary large number of samples, no confidence information is given on the characterization, and this method is often time-inefficient. To overcome this limitation, we propose a new method for noise evaluation. The error induced by fixed-point coding is statistically characterized to compute the noise power with an adaptive and reduced number of simulations. From user-defined confidence requirements, the proposed method computes the minimal number of simulations to obtain a confidence interval of the noise power. Experiments on varied signal-processing elementary blocks show that the proposed method requires on average the simulation of only 0.04% of the simulation set required by State of the Art techniques to estimate the noise power of a 64thorder FIR filter with a relative error less than 0.01%. Justine Bonnot, Karol Desnos, Daniel Ménard |
ICASSP | 3 |
| 2019 | Convex Energy Optimization of Streaming Applications for MPSoCsabstractThe energy efficiency of modern MPSoCs is enhanced by complex hardware features such as Dynamic Voltage and Frequency Scaling (DVFS) and Dynamic Power Management (DPM). This paper introduces a new method, based on convex problem solving, that determines the most energy efficient operating point in terms of frequency and number of active cores in an MPSoC. The solution can challenge the popular approaches based on never-idle (or As-Slow-As-Possible (ASAP)) and race-to-idle (or As-Fast-As-Possible (AFAP)) principles. Experimental data are reported using a Samsung Exynos 5410 MPSoC and show a reduction in energy of up to 27 % when compared to ASAP and AFAP. Erwan Nogues, Alexandre Mercat, Florian Arrestier, Maxime Pelcat, Daniel Ménard |
ICASSP | 5 |
| 2019 | Complexity Reduction Opportunities in the Future VVC Intra EncoderabstractThe Joint Video Expert Team (JVET) is developing the next-generation video coding standard called Versatile Video Coding (VVC) and their ultimate goal is to double the coding efficiency over the current state-of-the-art standard HEVC without letting complexity get out of hand. This work addresses the complexity of the VVC reference encoder called VVC Test Model (VTM) under All Intra coding configuration. The VTM3.0 is able to improve intra coding efficiency by 21% over the latest HEVC reference encoder HM16.19. This coding gain primarily stems from three new coding tools. First, the HEVC Quad-Tree (QT) structure extension with Multi-Type Tree (MTT) partitioning. Second, the duplication of intra prediction modes from 35 to 67. And third, the Multiple Transform Selection (MTS) scheme with two new discrete cosine/sine transforms (DCT-VIII and DST-VII). However, these new tools also play an integral part in making VTM intra encoding around 20 times as complex as that of HM. The purpose of this work is to analyze these tools individually and specify theoretical upper limits for their complexity reduction. According to our evaluations, the complexity reduction opportunity of block partitioning is up to 97%, i.e., the encoding complexity would drop down to 3% for the same coding efficiency if the optimal block partitioning could be directly predicted. The respective percentages for intra mode reduction and MTS optimization are 65% and 55%. We believe these results motivate VVC codec designers to develop techniques that are able to take most out of these opportunities. Alexandre Tissier, Alexandre Mercat, Thomas Amestoy, Wassim Hamidouche, Jarno Vanne, Daniel Ménard |
MMSP | 6 |
| 2019 | Hardware-friendly DST-VII/DCT-VIII approximations for the Versatile Video Coding StandardabstractVersatile Video Coding (VVC) is the next generation video coding standard expected by the end of 2020. The new concept of Multiple-Transform Selection (MTS) has been introduced in VVC. MTS enables the VVC encoder to select the transform that minimizes the rate-distortion cost among a set of pre-defined trigonometric transforms including the well known Discrete Cosine Transform (DCT)-II, DCT-VIII and Discrete Sine Transform (DST)-VII. Unlike the DCT-II that has fast computing algorithms, the DST-VII and DCT-VIII rely on more complex matrix multiplication.This paper tackles the problem of DST-VII and DCT-VIII approximations based on the DCT-II and an adjustment stage. This latter consists in a multiplication by a band-matrix with low number of non-zero coefficients per row. The approximation problem is first modeled as a constrained integer optimization problem minimizing both error and orthogonality. The genetic algorithm is then used to solve the optimization problem and find the adjustment band-matrix that minimizes a trade-off between error and orthogonality. The proposed solution enables to preserve the coding gain achieved by the MTS and considerably reduces the complexity in terms of required number of multiplications by coefficient. Moreover, the proposed approach is hardwarefriendly and will provide a lightweight shared hardware module for DST-II, DST-VII and DCT-VIII transforms. Wassim Hamidouche, Pierrick Philippe, Camar-Eddine Mohamed, Ahmed Kammoun, Daniel Ménard, Olivier Déforges |
PCS | 5 |
| 2019 | Numerical Representation of Directed Acyclic Graphs for Efficient Dataflow Embedded Resource AllocationabstractStream processing applications running on Heterogeneous Multi-Processor Systems on Chips (HMPSoCs) require efficient resource allocation and management, both at compile-time and at runtime. To cope with modern adaptive applications whose behavior can not be exhaustively predicted at compile-time, runtime managers must be able to take resource allocation decisions on-the-fly, with a minimum overhead on application performance. Resource allocation algorithms often rely on an internal modeling of an application. Directed Acyclic Graph (DAGs) are the most commonly used models for capturing control and data dependencies between tasks. DAGs are notably often used as an intermediate representation for deploying applications modeled with a dataflow Model of Computation (MoC) on HMPSoCs. Building such intermediate representation at runtime for massively parallel applications is costly both in terms of computation and memory overhead. In this paper, an intermediate representation of DAGs for resource allocation is presented. This new representation shows improved performance for run-time analysis of dataflow graphs with less overhead in both computation time and memory footprint. The performances of the proposed representation are evaluated on a set of computer vision and machine learning applications. Florian Arrestier, Karol Desnos, Eduardo Juárez Martínez, Daniel Ménard |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | A Fast and Fuzzy Functional Simulator of Inexact Arithmetic Operators for Approximate Computing SystemsabstractInexact operators are developed to exploit the tolerance of an application to imprecisions. These operators aim at reducing system energy consumption and memory footprint. In order to integrate the appropriate inexact operators in a complex system, the Quality of Service of the approximate system must be thoroughly studied through simulation. However, when simulating on a PC or workstation, the custom bit-level structures of inexact operators are not implemented in the instruction set of the simulating architecture. Consequently, the simulation requires a costly emulation, leading to expensive bit-level simulations. This paper proposes a new "Fast and Fuzzy" functional simulation method for inexact operators whose probabilistic behavior is correlated with the Most Significant Bits of the input operands. The proposed method processes real signal data and simplifies the error model for inexact operators, accelerating the simulation of the system. The modelization accuracy of the error can be controlled by a parameter called fuzzyness degree F. Using the proposed method, the bit-accurate logic-level simulation of inexact operators is replaced by an exact operator to which a pseudo-random error variable is added. Experiments on 16-bit operators show that the proposed simulation method, when compared to a bit-accurate logic level simulation, is up to 44 times faster. Justine Bonnot, Karol Desnos, Maxime Pelcat, Daniel Ménard |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | Stochastic Modeling to Accelerate Approximate Operators SimulationabstractApproximate operators have been developed to overcome the performance limitations of the original accurate arithmetic operators. They trade off the output quality of the operator and its energy consumption, area or delay. To benefit from this trade-off, the logic structure of the original accurate operator is modified. When integrating approximate operators in a complex system, numerous simulations of the application are required to ensure the fulfillment of the application requirements, despite the induced approximations. Because the hardware implementation of approximate operators is not always available in early phases of application prototyping, long software simulation of their complex bit-level structure is used. This paper proposes a fast simulator for approximate operators built from the output values of the original accurate operator. The error due to the approximation is modeled by a stochastic process whose features are learned from the errors of the approximate operator. The proposed simulator is compared to the bit-accurate logic-level simulation and to a simulator of approximate operators built on the input values of the operator. Experiments on 10-bit operators show that the proposed method is up to 63× faster than a bit-accurate logic level simulation. Justine Bonnot, Karol Desnos, Daniel Ménard |
ISCAS | 3 |
| 2018 | Machine Learning Based Choice of Characteristics for the One-Shot Determination of the HEVC Intra Coding TreeabstractIn the last few years, the Internet of Things (IoT) has become a reality. Forthcoming applications are likely to boost mobile video demand to an unprecedented level. A large number of systems are likely to integrate the latest MPEG video standard High Efficiency Video Coding (HEVC) in the long run and will particularly require energy efficiency. In this context, constraining the computational complexity of embedded HEVC encoders is a challenging task, especially in the case of software encoders. The most energy consuming part of a software intra encoder is the determination of the coding tree partitioning, i.e. the size of pixel blocks. This determination usually requires an iterative process that leads to repeating some encoding tasks. State-of-the-art studies have focused on predicting, from “easily” computed characteristics, an efficient coding tree. They have proposed and evaluated independently many characteristics for one-shot quad-tree prediction. In this paper, we present a fair comparison of these characteristics using a Machine Learning approach and a real-time HEVC encoder. Both computational complexity and information gain are considered, showing that characteristics are far from equivalent in terms of coding tree prediction performance. Alexandre Mercat, Florian Arrestier, Maxime Pelcat, Wassim Hamidouche, Daniel Ménard |
PCS | 5 |
| 2018 | Reproducible Evaluation of System Efficiency With a Model of Architecture: From Theory to PracticeabstractCurrent trends in high performance and embedded computing include design of increasingly complex hardware architectures with high parallelism, heterogeneous processing elements, and nonuniform communication resources. In order to take hardware and software design decisions, early evaluations of the system nonfunctional properties are needed. These evaluations of system efficiency require electronic system-level information on both algorithms and architecture. Contrary to algorithm models for which a major body of work has been conducted on defining formal models of computation (MoCs), architecture models from the literature are mostly empirical models from which reproducible experimentation requires the accompanying software. In this paper, a precise definition of a model of architecture (MoA) is proposed that focuses on reproducibility and abstraction and removes the overlap previously existing between the notions of MoA and MoC. A first MoA, called the linear system-level architecture model (LSLA), is presented. To demonstrate the generic nature of the proposed new architecture modeling concepts, we show that the LSLA model can be integrated flexibly with different MoCs. LSLA is then used to model the energy consumption of a state-of-the-art multiprocessor system-on-chip (MPSoC) when running an application described using the synchronous dataflow MoC. A method to automatically learn LSLA model parameters from platform measurements is introduced. Despite the high complexity of the underlying hardware and software, a simple LSLA model is demonstrated to estimate the energy consumption of the MPSoC with a fidelity of 86%. Maxime Pelcat, Alexandre Mercat, Karol Desnos, Luca Maggiani, Yanzhou Liu 0001, Julien Heulot, Jean-François Nezan, Wassim Hamidouche, Daniel Ménard, Shuvra S. Bhattacharyya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2017 | The hidden cost of functional approximation against careful data sizing - A case studyabstractMany applications are error-resilient, allowing for the introduction of approximations in the calculations, as long as a certain accuracy target is met. Traditionally, fixed-point arithmetic is used to relax accuracy, by optimizing the bit-width. This arithmetic leads to important benefits in terms of delay, power and area. Lately, several hardware approximate operators were invented, seeking the same performance benefits. However, a fair comparison between the usage of this new class of operators and classical fixed-point arithmetic with careful truncation or rounding, has never been performed. In this paper, we first compare approximate and fixed-point arithmetic operators in terms of power, area and delay, as well as in terms of induced error, using many state-of-the-art metrics and by emphasizing the issue of data sizing. To perform this analysis, we developed a design exploration framework, APXPERF, which guarantees that all operators are compared using the same operating conditions. Moreover, operators are compared in several classical real-life applications leveraging relevant metrics. In this paper, we show that considering a large set of parameters, existing approximate adders and multipliers tend to be dominated by truncated or rounded fixed-point ones. For a given accuracy level and when considering the whole computation data-path, fixed-point operators are several orders of magnitude more accurate while spending less energy to execute the application. A conclusion of this study is that the entropy of careful sizing is always lower than approximate operators, since it require significantly less bits to be processed in the data-path and stored. Approximated data therefore always contain on average a greater amount of costly erroneous, useless information. Benjamin Barrois, Olivier Sentieys, Daniel Ménard |
DATE | 3 |
| 2017 | Exploiting computation skip to reduce energy consumption by approximate computing, an HEVC encoder case studyabstractApproximate computing paradigm provides methods to optimize algorithms with considering both computational accuracy and complexity. This paradigm can be exploited at different levels of abstraction, from technological to application levels. Approximate computing at algorithm level aims at reducing computational complexity by approximating or skipping block functions of the computation. Numerous applications in the signal and image processing domain integrate algorithms based on discrete optimization techniques. These techniques minimize a cost function by exploring the search space. In this paper, a new approach is proposed to exploit the computation-skipping approximate computing concept by using the Smart Search Space Reduction (Sssr) technique. Sssr enables early selection of the best candidate configurations to reduce the search space. An efficient SSSR technique adjusts configuration selectivity to reduce execution complexity while selecting the most suitable functions to skip. The High Efficiency Video Coding (HEVC) encoder in All Intra (AI) profile is used as a case study to illustrate the benefits of SSSR. In this application, two functions use discrete optimization to explore different solutions and select the one leading to the minimal cost in terms of bitrate/quality and computational energy: coding-tree partitioning and intra-mode prediction. By applying SSSR to this use case, energy reductions from 20% to 70% are explored through Pareto in Rate-Energy space. Alexandre Mercat, Justine Bonnot, Maxime Pelcat, Wassim Hamidouche, Daniel Ménard |
DATE | 5 |
| 2017 | Energy reduction opportunities in an HEVC real-time encoderabstractHigh Efficiency Video Coding (HEVC) is one of the latest released video standards and offers up to 40% bitrate savings when compared to the widespread H.264/AVC standard, at the cost of a substantial complexity growth. Constraining the complexity of HEVC encoding is a challenging task for embedded applications based on a software encoder. In the last few years, the Internet of Thingss (IoTs) has become a reality. Forecoming applications are likely to boost mobile video demand to an unprecedented level. In this context, designing energy-efficient HEVC real-time encoders is becoming a major challenge for software and hardware designers. In this paper, an analysis is conducted of the energy reduction opportunities offered by an HEVC encoder. The energy reduction search space is demonstrated, and the impact on energy consumption of encoding tools at various levels of granularity is measured. Alexandre Mercat, Florian Arrestier, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
ICASSP | 5 |
| 2017 | Constrain the Docile CTUs: An In-Frame complexity allocator for HEVC Intra encodersabstractHigh Efficiency Video Coding (HEVC) is one of the latest released video standards and offers up to 40% bitrate savings when compared to the widespread H.264/AVC standard, at the cost of a substantial complexity growth. Constraining the complexity of HEVC encoding is a challenging task for embedded applications based on a software encoder. The most frequent approach to solve this problem is to optimise the coding tree structure to balance compression efficiency and computational complexity. In this context, we propose and assess a method to adequately allocate the computational complexity among coding units in a frame encoded in Intra mode. By studying an open-source real-time HEVC encoder, correlations are observed between Rate-Distortion (RD)-cost and encoding complexity that motivate a new complexity allocation technique. This technique, called “Constrain the Docile CTUs” (CDC), consists of allocating less computational complexity to units with low RD-costs and using RD-costs from preceding images as predictors for the current RD-costs. Experimental results demonstrate substantial gains, up to 36% of Bjøntegaard Delta Bit Rate (BD-BR), when using CDC method instead of other allocation methods. Alexandre Mercat, Florian Arrestier, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
ICASSP | 5 |
| 2017 | Smart search space reduction for approximate computing: A low energy HEVC encoder case study
Alexandre Mercat, Justine Bonnot, Maxime Pelcat, Karol Desnos, Wassim Hamidouche, Daniel Ménard |
J. Syst. Archit. | 6 |
| 2016 | New non-uniform segmentation technique for software function evaluationabstractEmbedded applications use more and more sophisticated computations. These computations can integrate composition of elementary functions and can easily be approximated by polynomials. Indeed, polynomial approximation methods allow to find a trade-off between accuracy and computation time. Software implementation of polynomial approximation in fixed-point processors is considered in this paper. To obtain a moderate approximation error, segmentation of the interval I on which the function is computed is necessary. This paper proposes a new method to compute the values of a function on I using non-uniform segmentation and polynomial approximation. Non-uniform segmentation allows one to minimize the number of segments created and is modeled by a tree-structure. The specifications of the segmentation set the balance between memory space requirement and computation time. Besides, compared to table-based methods or the CORDIC algorithm, our approach significantly reduces, the memory size and the function evaluation time respectively. Justine Bonnot, Erwan Nogues, Daniel Ménard |
ASAP | 3 |
| 2016 | Energy Efficient Scheduling of Real Time Signal Processing Applications through Combined DVFS and DPMabstractThis paper proposes a framework to design energy efficient signal processing systems. The energy efficiency is provided by combining Dynamic Frequency and Voltage Scaling (DVFS) and Dynamic Power Management (DPM). The framework is based on Synchronous Dataflow (SDF) modeling of signal processing applications. A transformation to a single rate form is performed to expose the application parallelism. An automated scheduling is then performed, minimizing the constraint of energy efficiency and providing DVFS and DPM decisions. This framework uses an architecture model including the number of available cores, the per-actor processing load and the energy per-cycle, derived from time and power measurements of modelled applications. After introducing the proposed framework, the energy characterization of big.LITTLE SoC systems is described. A generic approach is presented to generate the energy model of a platform from power measurements as customized polynomials. Finally, the experimental results on a Samsung Exynos 5410 big.LITTLE processor show that the energy optimal execution is not obtained by Linux governors that can execute either as-fast-as-possible or as-slow-as-possible. Instead, the most energy efficient scheduling is obtained by adapting both DVFS and DPM to application needs. Erwan Nogues, Maxime Pelcat, Daniel Ménard, Alexandre Mercat |
PDP | 3 |
| 2015 | A DVFS based HEVC decoder for energy-efficient software implementation on embedded processorsabstractSoftware video decoders for mobile devices are now a reality thanks to recent advances in Systems-on-Chip (SoC). The challenge has now moved to designing energy efficient systems. In this paper, we propose a light Dynamic Voltage Frequency Scaling (DVFS)-enabled software adapted to the much varying processing load of High Efficiency Video Coding (HEVC) real-time decoding. We analyze a practical evaluation of a HEVC decoder using our proposal on a Samsung Exynos low-power SoC widely used in portable devices. Experimental results show more than 50% of power savings on a real-time decoding when compared to the same software managed by the OnDemand Linux power management. For mobile applications, the proposed method can achieve 720p video HEVC decoding at 60 frames per second consuming approximately 1.1W with pure software decoding on a general purpose processor. Erwan Nogues, Romain Berrada, Maxime Pelcat, Daniel Ménard, Erwan Raffin |
ICME | 4 |
| 2014 | Integer word-length optimization for fixed-point systemsabstractTime-to-market and implementation cost are high-priority considerations in the automation of digital hardware design. Nowadays, digital signal processing applications are implemented into fixed-point architectures due to its advantage of manipulating data with lower word-length (WL). Thus, floating-point to fixed point conversion is mandatory. However, this conversion is translated into optimizing the integer word length (IWL) and fractional word length (FWL). Optimizing the IWL can significantly reduce the cost when the application is tolerant to a low probability of overflows. In this paper, we propose a new IWL optimization algorithm that exploits selective simulation technique to reduce both the implementation cost and optimization time. The efficiency of the algorithm is illustrated through experiments, where 17 to 22 % of cost reduction with respect to interval arithmetic and acceleration factor up to 617 with respect to classical max-1 algorithm are reported. Riham Nehmeh, Daniel Ménard, Andrei Banciu, Thierry Michel, Romuald Rocher |
ICASSP | 2 |
| 2014 | Implementation of a Stereo Matching algorithm onto a Manycore Embedded SystemabstractStereo Matching techniques aim at reconstructing the disparity maps with a pair of images. The use of Stereo Matching techniques in embedded systems is very challenging due to the complexity of the state of the art algorithms. This paper proposes a real-time Stereo Matching algorithm optimised for the last generation of Manycore Embedded Systems. The features and parameters of the algorithms have been chosen to optimise the trade-off between an high quality and a low complexity. A memory analysis is performed for the algorithm's decomposition and the resulting mapping on the Manycore platform is provided. Algorithm and arithmetic optimisations have been applied to each part of the algorithm to decrease the execution time down to 160ms for a CIF resolution. Alexandre Mercat, Jean-François Nezan, Daniel Ménard |
ISCAS | 3 |
| 2014 | Accelerated Performance Evaluation of Fixed-Point Systems With Un-Smooth OperationsabstractThe problem of accuracy evaluation is one of the most time consuming tasks during the fixed-point refinement process. Analytical techniques based on perturbation theory have been proposed in order to overcome the need for long fixed-point simulation. However, these techniques are not applicable in the presence of certain operations classified as un-smooth operations. In such circumstances, fixed-point simulation should be used. In this paper, an algorithm detailing the hybrid technique which makes use of an analytical accuracy evaluation technique used to accelerate fixed-point simulation is presented. This technique is applicable to signal processing systems with both feed-forward and feedback interconnect topology between its operations. The acceleration obtained as a result of applications of the proposed technique is consistent with fixed-point simulation, while reducing the time taken for fixed-point simulation by several orders of magnitude. Karthick Parashar, Daniel Ménard, Olivier Sentieys |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | A polynomial time algorithm for solving the word-length optimization problemabstractTrading off accuracy to the system costs is popularly addressed as the word-length optimization (WLO) problem. Owing to its NP-hard nature, this problem is solved using combinatorial heuristics. In this paper, a novel approach is taken by relaxing the integer constraints on the optimization variables and obtain an alternate noise-budgeting problem. This approach uses the quantization noise power introduced into the system due to fixed-point word-lengths as optimization variables instead of using the actual integer valued fixed-point word-lengths. The noise-budgeting problem is proved to be convex in the rounding mode quantization case and can therefore be solved using analytical convex optimization solvers. An algorithm with linear time complexity is provided in order to realize the actual fixed-point word-lengths from the noise budgets obtained by solving the convex noise-budgeting problem. Karthick Parashar, Daniel Ménard, Olivier Sentieys |
ICCAD | 2 |
| 2012 | Latency-Energy Optimized MAC Protocol for Body Sensor NetworksabstractThis paper presents a self organized asynchronous medium access control (MAC) protocol for wireless body area sensor (WBASN). The protocol is optimized in terms of latency and energy under variable traffic. A body sensor network (BSN) exhibits a wide range of traffic variations based on different physiological data emanating from the monitored patient. For example, electrocardiogram data rate is multiple times more in comparison with body temperature rate. In this context, we exploit the traffic characteristics being observed at each sensor node and propose a novel technique for latency-energy optimization at the MAC layer. The protocol relies on dynamic adaptation of wake-up interval based on a traffic status register bank. The proposed technique allows the wake-up interval to converge to a steady state for variable traffic rates, which results in optimized energy consumption and reduced delay during the communication. A comparison with other energy efficient protocols is presented. The results show that our protocol outperforms the other protocols in terms of energy as well as latency under the variable traffic of WBASN. Muhammad Mahtab Alam, Olivier Berder, Daniel Ménard, Olivier Sentieys |
BSN | 3 |
| 2012 | From Scilab to High Performance Embedded Multicore Systems: The ALMA ApproachabstractThe mapping process of high performance embedded applications to today's multiprocessor system on chip devices suffers from a complex tool chain and programming process. The problem here is the expression of parallelism with a pure imperative programming language which is commonly C. This traditional approach limits the mapping, partitioning and the generation of optimized parallel code, and consequently the achievable performance and power consumption of applications from different domains. The Architecture oriented paraLlelization for high performance embedded Multicore systems using scilAb (ALMA) European project aims to bridge these hurdles through the introduction and exploitation of a Scilab-based toolchain which enables the efficient mapping of applications on multiprocessor platforms from high level of abstraction. This holistic solution of the toolchain allows the complexity of both the application and the architecture to be hidden, which leads to a better acceptance, reduced development cost, and shorter time-to-market. Driven by the technology restrictions in chip design, the end of exponential growth of clock speeds, and an unavoidable increasing request of computing performance, ALMA is a fundamental step forward in the necessary introduction of novel computing paradigms and methodologies. Jürgen Becker 0001, Timo Stripf, Oliver Wolf, Michael Hübner 0001, Steven Derrien, Daniel Ménard, Olivier Sentieys, Gerard K. Rauwerda, Kim Sunesen, Nikolaos Kavvadias, Kostas Masselos, George Goulas, Panayiotis Alefragis, Nikos S. Voros, Dimitrios Kritharidis, Nikolaos Mitas, Diana Göhringer |
DSD | 6 |
| 2011 | Exploiting reconfigurable SWP operators for multimedia applicationsabstractImplementing image processing applications in embedded systems is a difficult challenge due to the drastic constraints in terms of cost, energy consumption and real time execution. Reconfigurable architectures are good candidates to take-up this challenge and especially when the architecture is able to support different word-lengths of pixel through Sub-Word Parallelism (SWP) capabilities. Exploiting the diversity of supported data-types requires automation tools able to optimize the data word-length under an accuracy constraint. In this paper, a new approach for word-length optimization in the case of SWP operations is proposed. Compared to existing approaches the optimization time is significantly reduced without sacrificing the quality of the optimized solution. The results show the ability of our approach to exploit the SWP capabilities associated with multimedia processors. Daniel Ménard, Hai-Nam Nguyen, François Charot, Stéphane Guyetant, Jérémie Guillot, Erwan Raffin, Emmanuel Casseau |
ICASSP | 1 |
| 2010 | Analytical approach for analyzing quantization noise effects on decision operatorsabstractThe presence of decision operators has proved to be a serious impediment for a fully analytical noise power estimation technique. This paper proposes a generalized decision operator which can potentially capture the behavior of all possible types of decision operators and provides a fully analytical technique to handle them while performing quantization noise power estimation. The proposed method is applied to BPSK and 16-QAM decision operators. The total error rate and the PDF of the error signal are found to follow the simulation to a great degree of accuracy. Karthick Parashar, Romuald Rocher, Daniel Ménard, Olivier Sentieys |
ICASSP | 3 |
| 2010 | Fast performance evaluation of fixed-point systems with un-smooth operatorsabstractFixed-point refinement of signal processing systems is an essential step performed before implementation of any signal processing system. Existing analytical techniques to evaluate performance of fixed-point systems are not applicable to the errors due to quantization in the presence of un-smooth operators. Thus, it is inevitable to use simulation to evaluate performance of fixed-point systems in the presence un-smooth operators. This paper proposes a hybrid technique which can be used in place of pure simulation to accelerate the performance evaluation. The principle idea in the proposed hybrid approach is to selectively simulate parts of the system only when un-smooth errors occur but use analytical results otherwise. The acceleration thus obtained reduces the performance evaluation time which can be used to explore a wider word-length design space or speedup the optimization process. This method has been tried on a complex MIMO sphere decoding algorithm and the results obtained show several orders of magnitude improvement in terms of evaluation time. Karthick Parashar, Daniel Ménard, Romuald Rocher, Olivier Sentieys, David Novo, Francky Catthoor |
ICCAD | 2 |
| 2009 | Reconfigurable SWP Operator for Multimedia ProcessingabstractFor performance enhancement, reconfigurable processors have to overcome the overheads of reconfigurations such as the complexity of the interconnection network and reconfiguration time. In processors dealing with multimedia applications these overheads can be reduced by providing the reconfigurability inside the processing units rather than at interconnection level. Due to the low precision data nature of multimedia applications, reconfiguration at operator level also provides additional speedup through parallel execution of low precision data. In this paper a pipelined architecture of a reconfigurable coarse grain subword parallel (SWP) operator is presented for multimedia applications. This operator not only eliminates the need of reconfiguration time but also provides the reconfigurability at both data size level (different pixel data sizes) and at operation level (different multimedia oriented operations). This ensures a better utilization of the processor resources and reduces the reconfiguration overheads significantly. Shafqat Khan, Emmanuel Casseau, Daniel Ménard |
ASAP | 3 |
| 2009 | Dynamic Precision Scaling for Low Power WCDMA ReceiverabstractOne of the most important applications of digital signal processing (DSP) is wireless communication. This kind of application requires low power implementation of DSP, which generally uses fixed-point arithmetic. The fixed-point architectures should be developed to maintain the energy consumption power at a reasonable level. In this paper, an approach which adapts the fixed-point specification according to the input receiver signal-to-noise ratio (SNR) is proposed. To underline our approach interest, the rake receiver of a WCDMA receiver is examined. Results show about 25% - 40% energy savings with our dynamic precision approach. Hai-Nam Nguyen, Daniel Ménard, Olivier Sentieys |
ISCAS | 2 |
| 2006 | Fixed-point configurable hardware components for adaptive filtersabstractTo reduce the gap between the VLSI technology capability and the designer productivity, design reuse based on IP (intellectual properties) is commonly used. In terms of arithmetic accuracy, the generated architecture can generally only be configured through the input and output word-lengths. In this paper, a new kind of fixed-point arithmetic IP is presented through the LMS and delayed-LMS examples. The operator and memory word-lengths are optimized under an accuracy constraint defined by the user. To significantly reduce the optimization and design times, the architecture parameter determination is based on analytical approach Romuald Rocher, Nicolas Hervé, Daniel Ménard, Olivier Sentieys |
ISCAS | 3 |
| 2005 | Accuracy evaluation of fixed-point APA algorithm [adaptive filter applications]abstractThe implementation of adaptive filters with fixed-point arithmetic requires us to evaluate the computation quality. The accuracy can be determined by calculating the global quantization noise power in the system output. In this paper, a new model for evaluating analytically the global noise power in the APA (affine projection algorithm) is developed. The model is presented and applied to the NLMS-OCF. The accuracy of our model is analyzed by experimentation. Romuald Rocher, Daniel Ménard, Olivier Sentieys, Pascal Scalart |
ICASSP (5) | 2 |
| 2004 | Accuracy evaluation of fixed-point LMS algorithmabstractThe implementation of adaptive filters with fixed-point arithmetic requires the computation quality to be evaluated. The accuracy may be determined by calculating the global quantization noise power in the system output. A new model for evaluating analytically the global noise power in the LMS algorithm and in the NLMS algorithm is developed. Two existing models are presented, then the model is detailed and compared with the ones before. The accuracy of our model is analyzed by simulation. Romuald Rocher, Daniel Ménard, Olivier Sentieys, Pascal Scalart |
ICASSP (5) | 2 |
| 2004 | DSP Code Generation with Optimized Data Word-Length Selection
Daniel Ménard, Olivier Sentieys |
SCOPES | 1 |
| 2002 | Automatic floating-point to fixed-point conversion for DSP code generationabstractThe development of methodologies for the automatic implementation of floating-point algorithms in fixed-point architectures is required for the minimization of cost, power consumption and time to market of digital signal processing applications. In this paper, a new methodology of implementation in Digital Signal Processors (DSP) under accuracy constraint is presented. In comparison with the existing methodologies, the DSP architecture is completely taken into account for optimizing the execution time under accuracy constraint. The justification and the different stages of our methodology are presented. Daniel Ménard, Daniel Chillet, François Charot, Olivier Sentieys |
CASES | 1 |
| 2002 | Automatic Evaluation of the Accuracy of Fixed-Point AlgorithmsabstractThe minimization of cost, power consumption and time-to-market of DSP applications requires the development of methodologies for the automatic implementation of floating-point algorithms in fixed-point architectures. In this paper a new methodology for evaluating the quality of an implementation through the automatic determination of the Signal to Quantization Noise Ratio (SQNR) is under consideration. The theoretical concepts and the different phases of the methodology are explained. Then, the ability of our approach for computing the SQNR efficiently and its beneficial contribution in the process of data word-length minimization are shown through some examples. Daniel Ménard, Olivier Sentieys |
DATE | 1 |
| 2002 | A methodology for evaluating the precision of fixed-point systemsabstractThe minimization of cost, power consumption and time-to-market of DSP applications requires the development of methodologies for the automatic implementation of floating-point algorithms in fixed-point architectures. In this paper, a new methodology for evaluating the quality of an implementation through the automatic determination of the Signal to Quantization Noise Ratio (SQNR) is presented. The modelization of the system at the quantization noise level and the expression of the output noise power is detailed for linear systems. Then, the different phases of the methodology are explained and the ability of our approach for computing the SQNR efficiently is shown through examples. Daniel Ménard, Olivier Sentieys |
ICASSP | 1 |