EDBT 2026 Demo / reviewers in the wild / expert
Marc Duranton
dblp:16/1360
· DBLP profile ↗
15ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0003-3762-264XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Emerging computing paradigms · 37% Hardware accelerators and domain-specific architectures · 35% Embedded and real-time systems · 12% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Programming languages and type systems · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
convolutional neural network |
0.2 | 1 | 2014 | Efficient Data Encoding for Convolutional Neural Network application · ACM Trans. Archit. Code Optim. 2014 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.2 | 1 | 2014 | Efficient Data Encoding for Convolutional Neural Network application · ACM Trans. Archit. Code Optim. 2014 |
Emerging computing paradigms
neuromorphic computing |
0.1 | 1 | 2012 | Capacitance of TSVs in 3-D stacked chips a problem?: not for neuromorphic systems! · DAC 2012 |
Programming languages and type systems › programming models
stream processing |
0.1 | 1 | 2006 | N-synchronous Kahn networks: a relaxed model of synchrony for real-time systems · POPL 2006 |
Programming languages and type systems › domain-specific languages › synchronous languages
synchronous dataflow languages |
0.1 | 1 | 2006 | N-synchronous Kahn networks: a relaxed model of synchrony for real-time systems · POPL 2006 |
Embedded and real-time systems
synchronous programming |
0.1 | 1 | 2006 | N-synchronous Kahn networks: a relaxed model of synchrony for real-time systems · POPL 2006 |
Emerging computing paradigms
approximate computing |
0.1 | 1 | 2014 | Efficient Data Encoding for Convolutional Neural Network application · ACM Trans. Archit. Code Optim. 2014 |
Processor architecture and microarchitecture › multi-chip architecture
3d stacking |
0.0 | 1 | 2012 | Capacitance of TSVs in 3-D stacked chips a problem?: not for neuromorphic systems! · DAC 2012 |
Integrated circuit design › 3d integration
through-silicon via |
0.0 | 1 | 2012 | Capacitance of TSVs in 3-D stacked chips a problem?: not for neuromorphic systems! · DAC 2012 |
Methods — techniques the papers use, named apart from their topics
significant position encoding · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Few hints towards more sustainable AlabstractArtificial Intelligence (AI) is now everywhere and its domains of application grow every day. But its demand in data and in computing power is also growing at an exponential rate, faster than used to be the “Moore's law”. The largest structures, like GPT-3, have impressive results but also trigger questions about the resources required for their learning phase, in the order of magnitude of hundreds of MWh. Once the learning done, the use of Deep Learning solutions (the “inference” phase) is far less energy demanding, but the systems are often duplicated in quantities (e.g. for consumer applications) and reused multiple times, so the cumulative energy consumption is also important. It is therefore of paramount importance to improve the efficiency of AI solutions in all their lifetime. This can only be achieved by combining efforts on several domains: on the algorithmic side, on the codesign application/algorithm/hardware, on the hardware architecture and on the (silicon) technology for example. The aim of this short tutorial is to raise awareness on the energy consumption of AI and to show different tracks to improve this problem, from distributed and federated learning, to optimization of Neural Networks and their data representation (e.g. using “Spikes” for information coding), to architectures specialized for AI loads, including systems where memory and computation are near, and systems using emerging memories or 3D stacking. Marc Duranton |
DATE | 1 |
| 2014 | Advanced technologies for brain-inspired computingabstractThis paper aims at presenting how new technologies can overcome classical implementation issues of Neural Networks. Resistive memories such as Phase Change Memories and Conductive-Bridge RAM can be used for obtaining low-area synapses thanks to programmable resistance also called Memristors. Similarly, the high capacitance of Through Silicon Vias can be used to greatly improve analog neurons and reduce their area. The very same devices can also be used for improving connectivity of Neural Networks as demonstrated by an application. Finally, some perspectives are given on the usage of 3D monolithic integration for better exploiting the third dimension and thus obtaining systems closer to the brain. Fabien Clermidy, Rodolphe Héliot, Alexandre Valentian, Christian Gamrat, Olivier Bichler, Marc Duranton, Bilel Belhadj, Olivier Temam |
ASP-DAC | 6 |
| 2014 | The improbable but highly appropriate marriage of 3D stacking and neuromorphic acceleratorsabstract3D stacking is a promising technology (low latency/power/area, high bandwidth); its main shortcoming is increased power density. Simultaneously, motivated by energy constraints, architectures are evolving towards greater customization, with tasks delegated to accelerators. Due to the widespread use of machine-learning algorithms and the re-emergence of neural networks (NNs) as the preferred such algorithms, NN accelerators are receiving increased attention. They turn out to be well matched to 3D stacking: inherently 3D structures with a low power density and high across-layer bandwidth requirements. We present what is, to the best of our knowledge, the first 3D stacked NN accelerator Bilel Belhadj, Alexandre Valentian, Pascal Vivet, Marc Duranton, Liqiang He, Olivier Temam |
CASES | 4 |
| 2014 | Low complexity multi-target tracking for embedded systems
Aziz Dziri, Marc Duranton, Roland Chapuis |
FUSION | 2 |
| 2014 | Efficient Data Encoding for Convolutional Neural Network applicationabstractThis article presents an approximate data encoding scheme called Significant Position Encoding (SPE) . The encoding allows efficient implementation of the recall phase (forward propagation pass) of Convolutional Neural Networks (CNN)—a typical Feed-Forward Neural Network. This implementation uses only 7 bits data representation and achieves almost the same classification performance compared with the initial network: on MNIST handwriting recognition task, using this data encoding scheme losses only 0.03% in terms of recognition rate (99.27% vs. 99.3%). In terms of storage, we achieve a 12.5% gain compared with an 8 bits fixed-point implementation of the same CNN. Moreover, this data encoding allows efficient implementation of processing unit thanks to the simplicity of scalar product operation—the principal operation in a Feed-Forward Neural Network. Hong-Phuc Trinh, Marc Duranton, Michel Paindavoine |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Capacitance of TSVs in 3-D stacked chips a problem?: not for neuromorphic systems!abstractIn order to cope with increasingly stringent power and variability constraints, architects need to investigate alternative paradigms. Neuromorphic architectures are increasingly considered (especially spike-based neurons) because of their inherent robustness and their energy efficiency. Yet, they have two limitations: the massive parallelism among neurons is hampered by 2D planar circuits, and the most cost-effective hardware neurons are analog implementations that require large capacitors, We show that 3D stacking with Through-Silicon-Vias applied to neuromorphic architectures can solve both issues: not only by providing massive parallelism between layers, but also by turning the parasitic capacitances of TSVs into useful capacitive storage. Antoine Joubert, Marc Duranton, Bilel Belhadj, Olivier Temam, Rodolphe Héliot |
DAC | 2 |
| 2012 | Balancing Programmability and Silicon Efficiency of Heterogeneous Multicore ArchitecturesabstractMulticore architectures provide scalable performance with a lower hardware design effort than single core processors. Our article presents a design methodology and an embedded multicore architecture, focusing on reducing the software design complexity and boosting the performance density. First, we analyze characteristics of the Task-Level Parallelism in modern multimedia workloads. These characteristics are used to formulate requirements for the programming model. Then we translate the programming model requirements to an architecture specification, including a novel low-complexity implementation of cache coherence and a hardware synchronization unit. Our evaluation demonstrates that the novel coherence mechanism substantially simplifies hardware design, while reducing the performance by less than 18% relative to a complex snooping technique. Compared to a single processor core, the multicores have already proven to be more area- and energy-efficient. However, the multicore architectures in embedded systems still compete with highly efficient function-specific hardware accelerators. In this article we identify five architectural methods to boost performance density of multicores; microarchitectural downscaling, asymmetric multicore architectures, multithreading, generic accelerators, and conjoining. Then, we present a novel methodology to explore multicore design spaces, including the architectural methods improving the performance density. The methodology is based on a complex formula computing performances of heterogeneous multicore systems. Using this design space exploration methodology for HD and QuadHD H.264 video decoding, we estimate that the required areas of multicores in CMOS 45 nm are 2.5 mm 2 and 8.6 mm 2 , respectively. These results suggest that heterogeneous multicores are cost-effective for embedded applications and can provide a good programmability support. Andrei Sergeevich Terechko, Jan Hoogerbrugge, Ghiath Alkadi, Surendra Guntur, Anirban Lahiri, Marc Duranton, Clemens C. Wüst, Phillip Christie, Axel Nackaerts, Aatish Kumar |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2010 | Erbium: a deterministic, concurrent intermediate representation to map data-flow tasks to scalable, persistent streaming processesabstractTuning applications for multicore systems involve subtle concurrency concepts and target-dependent optimizations. This paper advocates for a streaming execution model, called ER, where persistent processes communicate and synchronize through a multi-consumer processing applications, we demonstrate the scalability and efficiency advantages of streaming compared to data-driven scheduling. To exploit these benefits in compilers for parallel languages, we propose an intermediate representation enabling the compilation of data-flow tasks into streaming processes. This intermediate representation also facilitates the application of classical compiler optimizations to concurrent programs. Cupertino Miranda, Antoniu Pop, Philippe Dumont, Albert Cohen 0001, Marc Duranton |
CASES | 5 |
| 2008 | A Look-Ahead Task Management Unit for Embedded Multi-Core ArchitecturesabstractEfficient utilization of multi-core architectures relies on the partitioning of applications into tasks and mapping the tasks to cores. In some applications (e.g. H.264 video decoding parallelized at macro-block level) these tasks have dependencies among each other. Task scheduling, consisting of selecting a task with satisfied dependencies and mapping it to a core, is typically a functionality delegated to the Operating System. In this paper we present a hardware Task Management Unit (TMU) that looks ahead in time to find tasks to be executed by a multi-core architecture. The look-ahead functionality is shown to reduce the task management overhead by 40-50% when executing a parallelized version of an H.264 video decoder on an architecture with up to 16 cores. In overall, the TMU-based multi-core architecture reaches a speedup of more than 14x on 16 cores running H.264 video decoding, assuming CABAC is implemented in a dedicated coprocessor. Magnus Själander, Andrei Sergeevich Terechko, Marc Duranton |
DSD | 3 |
| 2006 | The Challenges for High Performance Embedded SystemsabstractConsumer electronics devices traditionally rely on non-programmable circuits for their "streaming" part. Recent demands on flexibility moved the balance towards the use of programmable components. However, there is a major gap between the current programmable processors and the actual requirements of applications. To bridge this gap, it is necessary to use parallel architectures consisting of multiple, programmable compute blocks, specifically designed for efficient processing of data streams. Programming those architectures poses major challenges and requires appropriate tools. Managing the ever-increasing complexity of those embedded systems is certainly one of the most important challenges beside power consumption. Complexity will make systems unreliable and unpredictable. The new technology nodes (65nm, and below) will also bring their own additional challenges: the global interconnect delay that does not scale, the predominant leakage current and the increasing variability of components Marc Duranton |
DSD | 1 |
| 2006 | N-synchronous Kahn networks: a relaxed model of synchrony for real-time systemsabstractThe design of high-performance stream-processing systems is a fast growing domain, driven by markets such like high-end TV, gaming, 3D animation and medical imaging. It is also a surprisingly demanding task, with respect to the algorithmic and conceptual simplicity of streaming applications. It needs the close cooperation between numerical analysts, parallel programming experts, real-time control experts and computer architects, and incurs a very high level of quality insurance and optimization.In search for improved productivity, we propose a programming model and language dedicated to high-performance stream processing. This language builds on the synchronous programming model and on domain knowledge -- the periodic evolution of streams -- to allow correct-by-construction properties to be proven by the compiler. These properties include resource requirements and delays between input and output streams. Automating this task avoids tedious and error-prone engineering, due to the combinatorics of the composition of filters with multiple data rates and formats. Correctness of the implementation is also difficult to assess with traditional (asynchronous, simulation-based) approaches. This language is thus provided with a relaxed notion of synchronous composition, called n-synchrony: two processes are n-synchronous if they can communicate in the ordinary (0-)synchronous model with a FIFO buffer of size n.Technically, we extend a core synchronous data-flow language with a notion of periodic clocks, and design a relaxed clock calculus (a type system for clocks) to allow non strictly synchronous processes to be composed or correlated. This relaxation is associated with two sub-typing rules in the clock calculus. Delay, buffer insertion and control code for these buffers are automatically inferred from the clock types through a systematic transformation into a standard synchronous program. We formally define the semantics of the language and prove the soundness and completeness of its clock calculus and synchronization transformation. Finally, the language is compared with existing formalisms. Albert Cohen 0001, Marc Duranton, Christine Eisenbeis, Claire Pagetti, Florence Plateau, Marc Pouzet |
POPL | 2 |
| 2005 | Synchronization of periodic clocksabstractWe propose a programming model dedicated to real-time video-streaming applications for embedded media devices, including high-definition TVs. This model is built on the synchronous programming model extended with domain-specific knowledge --- periodic evolution of streams --- to allow correct-by-construction properties of the application to be proven by the compiler. These properties include buffer requirements and delays between input and output streams.Such properties are tedious to analyze by hand, due to the combinatorics of video filters, multiple data rates and formats. We show how to extend a core synchronous data-flow language with a notion of periodic clocks, and to design a relaxed clock calculus (a type system for clocks) to allow non strictly synchronous processes to be composed. This relaxation is associated with a subtyping rule in the clock calculus. Delay, buffer insertion and control code for these buffers are automatically inferred from the clock types through a systematic program transformation. Albert Cohen 0001, Marc Duranton, Christine Eisenbeis, Claire Pagetti, Florence Plateau, Marc Pouzet |
EMSOFT | 2 |
| 2003 | Understanding Video Pixel Processing Applications for Flexible ImplementationsabstractMedia processing system-on-chips (SoCs) mainly consist of audio encoding/decoding (e.g. AC-3, MP3), video encoding/decoding (e.g. H263, MPEG-2) and video pixel processing functions (e.g. de-interlacing, noise reduction). Video pixel processing functions have very high computational demands, as they require a large amount of computations on large amount of data (note that the data are pixels of completely decoded pictures). In this paper, we focus on video pixel processing functions. Usually, these functions are implemented in dedicated hardware. However, flexibility (by means of programmability or reconfigurability) is needed to introduce the latest innovative algorithms, to allow differentiation of products, and to allow bug fixing after fabricating chips. It is impossible to fulfill the computational requirements of these functions by current programmable media processors. To achieve efficient implementations for flexible solutions, we will study, in this paper, the application characteristics of some representative video pixel processing functions. The characteristics considered are granularity of operations, amount and kind of data accesses and degree of parallelism present in these functions. We observe that from computational granularity point of view many functions can be expressed in terms of kernels e.g. Median3 (i.e. median of three values), finite impulse response (FIR) filters, table lookups (LUT) etc. that are coarser grain than ALU, Mult, MAC, etc. Regarding the kind of data accesses, we categorize these functions as regular, regular with some data rearrangement and irregular data access patterns. Furthermore, the degree of parallelism present in these functions is expressed in terms of data level parallelism (DLP) and instruction/operation level parallelism (ILP). We show with an example that these properties can be exploited to make specialized programmable processors. Om Prakash Gangwal, Johan G. W. M. Janssen, Selliah Rathnam, Erwin B. Bellers, Marc Duranton |
DSD | 5 |
| 2002 | Multi-periodic Process Networks: Prototyping and Verifying Stream-Processing Systems
Albert Cohen 0001, Daniela Genius, Abdesselem Kortebi, Zbigniew Chamski, Marc Duranton, Paul Feautrier |
Euro-Par | 5 |
| 1992 | Lneuro 1.0: a piece of hardware LEGO for building neural network systemsabstractNeural network simulations on a parallel architecture are reported. The architecture is scalable and flexible enough to be useful for simulating various kinds of networks and paradigms. The computing device is based on an existing coarse-grain parallel framework (INMOS transputers), improved with finer-grain parallel abilities through VLSI chips, and is called the Lneuro 1.0 (for LEP neuromimetic) circuit. The modular architecture of the circuit makes it possible to build various kinds of boards to match the expected range of applications or to increase the power of the system by adding more hardware. The resulting machine remains reconfigurable to accommodate a specific problem to some extent. A small-scale machine has been realized using 16 Lneuros, to experimentally test the behavior of this architecture. Results are presented on an integer version of Kohonen feature maps. The speedup factor increases regularly with the number of clusters involved (to a factor of 80). Some ways to improve this family of neural network simulation machines are also investigated. Nicolas Mauduit, Marc Duranton, Jean Gobert, Jacques Ariel Sirat |
IEEE Trans. Neural Networks | 2 |