EDBT 2026 Demo / reviewers in the wild / expert
Lionel Lacassagne
dblp:81/3164
· DBLP profile ↗
26ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 2 since 2021Systems, architecture and hardware · 11 · 5 since 2021Artificial intelligence and machine learning · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-aware scheduling strategies for partially-replicable task chains on heterogeneous processors
Yacine Idouar, Adrien Cassagne, Laércio Lima Pilla, Julien Sopena, Manuel Bouyer, Diane Orhan, Lionel Lacassagne, Dimitri Galayko, Denis Barthou, Christophe Jégo |
Parallel Comput. | 7 |
| 2024 | A New Efficient Split & Merge Algorithm for Embedded SystemsabstractThis article presents a new image segmentation algorithm based on a Split & Merge approach. By nature, the execution time of Split & Merge algorithms is data-dependent, as their halting conditions are tied to the homogeneity of each region. While previous algorithms made the Split step less sensitive to input data, the execution time of the more complex Merge step remains highly sensitive to image content. This paper tackles the sensitivity and performance problems from a system and architecture perspective. Memory reallocations due to array fusions are eliminated with the introduction of a TTA (Three Table Array) structure in the Merge step. As iterating over entries in this structure causes a loss of memory locality, we propose two new mechanisms that implement a software cache to mitigate this. An experimental study on an embedded system (Nvidia Jetson Xavier NX) has shown our Merge algorithm to be 10.6 times faster than the state-of-the-art Split & Merge algorithm for $960 \times 720$ images. Moreover, the execution time of our algorithm is also more resistant to image characteristics. Nathan Maurice, Julien Sopena, Lionel Lacassagne |
ICIP | 3 |
| 2023 | Real-Time and Approximate Iterative Optical Flow Implementation on Low-Power Embedded CPUsabstractOptical flow estimation is used in many embedded computer vision applications, and it is known to be computationally intensive. In the literature, many methods exist to estimate optical flow. Thus, the challenge is to find a method that matches the applicative constraints. In an embedded system, a trade-off between power consumption and execution time has to be made to meet both energy and framerate constraints. This work proposes methods to implement an approximate Horn & Schunck optical flow estimation that meets embedded CPUs constraints. This is achieved thanks to architectural optimizations, software optimizations and algorithm tuning. For instance, on the NVIDIA Jetson Nano, and for HD video sequences, the achieved frame latency is 12 ms for 5 Watts. To the best of our knowledge, this is the fastest optical flow implementation on embedded CPUs. Maxime Millet, Adrien Cassagne, Nicolas Rambaux, Lionel Lacassagne |
ASAP | 4 |
| 2022 | Using HLS for Designing a Parametric Optical Flow Hierarchical Algorithm in FPGAsabstractIn this work HLS is used for designing a parametric optical flow Hierarchical algorithm in FPGAs. The algorithm that is designed is the Hierarchical (pyramid) Horn and Schunck algorithm, both a multi-rate and multi-level (multi-scale) algorithm, which achieves larger motion displacement detection than the mono-scale ones. With the help of HLS, we parametrize our design in terms of the levels of the pyramid, the iteration factor and the number of pixels computed per clock. We are reusing the same resources in each level of the pyramid to keep the usage of DSPs and RAM low. We perform a design space exploration of the algorithm and we show that our fastest design achieves a throughput of 461 Mpixel/s in a 2048×2048 resolution pixel image. Ilias Bournias, Roselyne Chotin, Lionel Lacassagne |
ISCAS | 3 |
| 2022 | Max-Tree Computation on GPUsabstractIn Mathematical Morphology, the max-tree is a region-based representation that encodes the inclusion relationship of the threshold sets of an image. This tree has proved useful in numerous image processing applications. For the last decade, work has led to improving the construction time of this structure; mixing algorithmic optimizations, parallel and distributed computing. Nevertheless, there is still no algorithm that benefits from the computing power of the massively parallel architectures. In this work, we propose the first GPU algorithm to compute the max-tree. The proposed approach leads to significant speed-ups, and is up to one order of magnitude faster than the current State-of-the-Art parallel CPU algorithms. This work paves the way for a max-tree integration in image processing GPU pipelines and real-time image processing based on Mathematical Morphology. It is also a foundation for porting other image representations from Mathematical Morphology on GPUs. Nicolas Blin, Edwin Carlinet, Florian Lemaitre, Lionel Lacassagne, Thierry Géraud |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | Taming Voting Algorithms on Gpus for an Efficient Connected Component Analysis AlgorithmabstractConnected Component Analysis is vastly used as a building block for many Computer Vision algorithms from many fields like medical image processing, surveillance, or autonomous driving. It extends Connected Component Labeling by computing some features of the connected components like their bounding box or their surface. As such, Connected Component Analysis is a voting algorithm just like histogram computation or Hough transform. Voting algorithms are difficult on many-core architectures like GPUs because of the serialization of atomic memory accesses. The trend to increase the number of cores makes this issue even more critical. This paper explores multiple ways to reduce those conflicts for voting algorithms and especially for Connected Component Analysis. We show that our new algorithm is from 4 up to 10 times faster than State-of-the-Art on average on an Nvidia A100. Florian Lemaitre, Arthur M. Hennequin, Lionel Lacassagne |
ICASSP | 3 |
| 2021 | FPGA Acceleration of the Horn and Schunck Hierarchical AlgorithmabstractThis work proposes a highly tunable motion estimation architecture. We implement the Horn and Schunck algorithm with the hierarchical extension for larger motion estimations in FPGAs. Different architectures are explored dealing with interpolation, pipeline, parallelism and arithmetic format, in order to fit performance. We show in our exploration, how the different cores of our system should be used to increase the throughput. Our smallest design achieves a 30.8 Mpixel/s in a 1024×1024 resolution and the fastest 507 Mpixel/s which is one of the fastest ever achieved, as far as we know, for FPGAs. Ilias Bournias, Roselyne Chotin, Lionel Lacassagne |
ISCAS | 3 |
| 2018 | Harris corner detection on a NUMA manycore
Olfa Haggui, Claude Tadonki, Lionel Lacassagne, Fatma Sayadi, Bouraoui Ouni |
Future Gener. Comput. Syst. | 3 |
| 2017 | Cholesky factorization on SIMD multi-core architectures
Florian Lemaitre, Ben Couturier, Lionel Lacassagne |
J. Syst. Archit. | 3 |
| 2017 | Color enhanced local binary patterns in covariance matrices descriptors (ELBCM)
Michèle Gouiffès, Andrés Romero Mier y Terán, Lionel Lacassagne |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | A modular system for global and local abnormal event detection and categorization in videos
Ahmed Chamseddine Ben Abdallah, Michèle Gouiffès, Lionel Lacassagne |
Mach. Vis. Appl. | 3 |
| 2015 | Parallel light speed labeling: An efficient connected component labeling algorithm for multi-core processorsabstractThe paper introduces the parallel version of the Light Speed Labeling (LSL) and compares it with the parallel versions of the competitors. A benchmark shows that the parallel Light Speed Labeling is ×1.8 faster than all the other algorithms for random images on average. This factor reaches ×3.2 for structured random images. More importantly, we show that thanks to its run-based processing (segments), LSL is intrinsically more efficient than all pixel-based algorithms. Laurent Cabaret, Lionel Lacassagne, Daniel Etiemble |
ICIP | 2 |
| 2014 | The numerical template toolbox: A modern C++ design for scientific computing
Pierre Estérie, Joël Falcou, Mathias Gaunard, Jean-Thierry Lapresté, Lionel Lacassagne |
J. Parallel Distributed Comput. | 5 |
| 2012 | Accelerator-Based implementation of the Harris Algorithm
Claude Tadonki, Lionel Lacassagne, Elwardani Dadi, El Mostafa Daoudi |
ICISP | 2 |
| 2012 | Motion histogram quantification for human action recognition
Hedi Tabia, Michèle Gouiffès, Lionel Lacassagne |
ICPR | 3 |
| 2010 | Projection-histograms for mean-shift trackingabstractThis paper proposes an extension to the mean shift tracking by using XY projection-histograms to model the object. More than providing statistical information about the target to track, they embed information about the spatial arrangement of pixels. This approach, without any complexity increase, provides a better robustness and quality of the tracking. That is asserted by the experiments performed on several sequences showing either vehicles or pedestrians in various contexts. Michèle Gouiffès, Florence Laguzet, Lionel Lacassagne |
ICIP | 3 |
| 2010 | Color Connectedness Degree for Mean-Shift TrackingabstractThis paper proposes an extension to the mean shift tracking. We introduce the color connectedness degrees (CCD) which, more than providing statistical information about the target to track, embeds information about the amount of connectedness of the color intervals which compose the target. With a low increase of complexity, this approach provides a better robustness and quality of the tracking compared to the use of the RGB space. This is asserted by the experiments performed on several sequences showing vehicles and pedestrians in various contexts. Michèle Gouiffès, Florence Laguzet, Lionel Lacassagne |
ICPR | 3 |
| 2009 | Algorithmic Skeletons within an Embedded Domain Specific Language for the CELL ProcessorabstractEfficiently using the hardware capabilities of the Cell processor, a heterogeneous chip multiprocessor that uses several levels of parallelism to deliver high performance, and being able to reuse legacy code are real challenges for application developers. We propose to use Generative Programming and more precisely template meta-programming to design an domain specific embedded language using algorithmic skeletons to generate applications based on a high-level mapping description. The method is easy to use by developers and delivers performance close to the performance of optimized hand-written code, as shown on various benchmarks ranging from simple BLAS kernels to image processing applications. Tarik Saidani, Joël Falcou, Claude Tadonki, Lionel Lacassagne, Daniel Etiemble |
PACT | 4 |
| 2009 | Motion detection: Fast and robust algorithms for embedded systemsabstractThis article introduces a new hierarchical version of a set of motion detection algorithms called ¿¿. These new algorithms are designed to preserve as much as possible the computational efficiency of the basic ¿¿ estimation, in order to target real-time implementation for low power consumption processors and embedded systems. Lionel Lacassagne, Antoine Manzanera, Antoine Dupret |
ICIP | 1 |
| 2009 | Light Speed Labeling for RISC architecturesabstractThis article introduces a fast algorithm for Connected Component Labeling of binary images called Light Speed Labeling. It is segment-based and a line-relative labeling that was especially thought for RISC computers. An extensive benchmark on both structured and unstructured images substanciates that the algorithm, the way it is designed, is faster and more runtime predictable than Wu's algorithm claimed to be the world fastest in 2007. Lionel Lacassagne, Bertrand Y. Zavidovique |
ICIP | 1 |
| 2007 | Adaptive Multiresolution for Low Power CMOS Image SensorabstractTo be implemented on an analog CMOS image sensor, a robust algorithm based on recursive operations is presented. It allows sensor's acuity adaptation to the scene activity. The main interest of the presented motion detection with adaptive thresholding is that, in a context of embedded steady camera, such a system allows focusing on targets with high resolution while keeping background in low resolution. Drastic power consumption reduction is achieved by tremendously reducing the amount of processed data. Arnaud Verdant, Antoine Dupret, Hervé Mathias, Patrick Villard, Lionel Lacassagne |
ICIP (5) | 5 |
| 2007 | Parallelization Strategies for the Points of Interests Algorithm on the Cell Processor
Tarik Saidani, Lionel Lacassagne, Samir Bouaziz, Taj Muhammad Khan |
ISPA | 2 |
| 2004 | 16-Bit FP Sub-Word Parallelism to Facilitate Compiler Vectorization and Improve Performance of Image and Media ProcessingabstractWe consider the implementation of 16-bit floating point instructions on a Pentium 4 and a PowerPC G5 for image and media processing. By measuring the execution time of benchmarks with these new simulated instructions, we show that significant speed-up is obtained compared to 32-bit FP versions. For image processing, the speed-up both comes from doubling the number of operations per SIMD instruction and the better cache behavior with byte storage. For data stream processing with arrays of structures, the speed-up mainly comes from the wider SIMD instructions. Daniel Etiemble, Lionel Lacassagne |
ICPP | 2 |
| 2000 | Object Image Retrieval with Image Compactness VectorsabstractWe present in this paper a new global measure to characterise an image: the compactness vector. This measure considers both object shape and grey level distribution function and does not require any preliminary segmentation. It is invariant to rotation, translation, scale and luminance and is then a powerful tool for image retrieval from a query image. We present here some object retrieval examples from large database images. Catherine Achard, Jean Devars, Lionel Lacassagne |
ICPR | 3 |
| 1999 | A generic methodology for the software managing of caches in multi-processors DSP architecturesabstractThis article introduces a novel software engineering methodology designed for the real-time execution of low-level image operators running on multi-processors DSP architectures. We detail the results we gained while implementing our approach on the TMS320C80, a shared memory multi-processors architecture. Our contribution compares to other existing C80's image processing libraries in terms of genericity, flexibility, and performance improvement. More specifically, generic mechanisms allows one to address various operator's requirements as well as expanding them using a standard framework. Our approach is flexible enough to allow for the dynamic composing of concurrent and reconfigurable processing chains thanks to a modular library implementing basic operators. Processing chains work on various image sizes and with any number of processors. Above all, our methodology permits performance improvement by enhancing data locality. Frantz Lohier, Lionel Lacassagne, Patrick Garda |
ICASSP | 2 |
| 1998 | Real time execution of optimal edge detectors on RISC and DSP processorsabstractThis paper presents the real time implementations of the Canny-Deriche (1986, 1990) optimal edge detectors on RISC and DSP processors. For each type of architecture, the most leading optimization techniques are described. A comparison is then made between RISC and DSP processing speeds. Lionel Lacassagne, Frantz Lohier, Patrick Garda |
ICASSP | 1 |