Paul Marchal

dblp:63/1767 · also Pol Marchal · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
0since 2021 · last 2013
0000-0003-2821-8119ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 1 first-authorSoftware engineering, systems software and programming languages · 9 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Integrated circuit design · 43% Processor architecture and microarchitecture · 17% Parallel and multicore computing · 14%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
3d integration
0.222010
3-D stacked die: now or future? · DAC 2010
3-D Technology Assessment: Path-Finding the Technology/Design Sweet-Spot · Proc. IEEE 2009
Processor architecture and microarchitecture › multi-chip architecture
3d stacking
0.112011
3D heterogeneous system integration: application driver for 3D technology development · DAC 2011
Integrated circuit design › 3d integration
through-silicon via
0.112010
3-D stacked die: now or future? · DAC 2010
Electronic design automation
design technology co-optimization
0.112009
3-D Technology Assessment: Path-Finding the Technology/Design Sweet-Spot · Proc. IEEE 2009
Processor architecture and microarchitecture
chip multiprocessor
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Energy-efficient computing
power-performance tradeoff
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Parallel and multicore computing
programming models
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Parallel and multicore computing › parallel programming models › hybrid programming models
shared memory and message passing
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Embedded and real-time systems › embedded software › embedded operating systems
embedded memory management
0.012004
An integrated hardware/software approach for run-time scratchpad management · DAC 2004
Embedded and real-time systems › embedded software › embedded operating systems › embedded memory management
scratchpad memory management
0.012004
An integrated hardware/software approach for run-time scratchpad management · DAC 2004
Integrated circuit design
heterogeneous integration
0.012011
3D heterogeneous system integration: application driver for 3D technology development · DAC 2011
Integrated circuit design
packaging
0.012010
3-D stacked die: now or future? · DAC 2010
Parallel and multicore computing
parallel programming models
0.012007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Runtime systems and virtual machines
runtime memory management
0.012004
An integrated hardware/software approach for run-time scratchpad management · DAC 2004
Memory systems › on-chip memory
scratchpad memory
0.012004
An integrated hardware/software approach for run-time scratchpad management · DAC 2004

Methods — techniques the papers use, named apart from their topics

path-finding methodology · 0.1hardware-software co-design · 0.1hardware-software tuning · 0.1comparative analysis · 0.1
YearPublicationVenuePosition
2013 Design issues in heterogeneous 3D/2.5D integration
abstract
Efficient processing of fine-pitched Through Silicon Vias, micro-bumps and back-side re-distribution layers enable face-to-back or face-to-face integration of heterogeneous ICs using 3D stacking and/or Silicon Interposers. While these technology features are extremely compelling, they considerably stress the existing design practices and EDA tool flows typically conceived for 2D systems. With all system, technology and implementation level options brought with these features, the design space increases to an extent where traditional 2D tools cannot be used any more for efficient exploration. Therefore, the cost-effective design of future 3D ICs products will require new planning and co-optimisation techniques and tools that are fast and accurate enough to cope with these challenges. In this paper we present design methodology and the practical EDA tool chain that covers different aspects of the design flow and is specific to efficient design of 3D-ICs. Flow features include: fast synthesis and 3D design partitioning at gate level, TSV/micro-bump array planning, 3D floor planning, placement and routing, congestion analysis, fast thermal and mechanical modeling, easy technology vs. implementation trade-off analysis, 3D device models generations and Design-for-Test (DfT). The application of the tool chain is illustrated using concrete example of a real-world design, showing not only the applicability of the tool chain, but also the benefits of heterogeneous 2.5 and 3D integration technologies.
Dragomir Milojevic, Paul Marchal, Erik Jan Marinissen, Geert Van der Plas, Diederik Verkest, Eric Beyne
ASP-DAC2
2011 3D heterogeneous system integration: application driver for 3D technology development
abstract
Three dimensional integration complements semiconductor scaling; it enables a higher integration density as well as heterogeneous technology integration. Using 3D chip stacking, it is possible to extend the number of functions per 3D chip well beyond the near-term capabilities of traditional scaling. The 3D strata may be realized using advanced CMOS technology nodes but may also exploit a wide variety of device technologies to optimize system performance.
Eric Beyne, Paul Marchal, Geert Van der Plas
DAC2
2010 3-D stacked die: now or future?
abstract
The continuation of Moore's law by conventional CMOS scaling is becoming challenging. 3D Packaging with 3D through silicon vias (TSV) interconnects is showing promise for extending scaling using mature silicon technology, providing another path towards the "More than Moore". Two years ago, the big unceasing question was "Why 3D?" Today, as we move forward with the concrete implementation of the technology, the questions are now "When 3D?" and "How 3D?" There are quite a few brave souls who have taken this disruptive interconnect technology and are investing in it today to gain benefit from it. However, for many the lingering questions remain "Are we there yet?" "Is it now or the future?"
Samta Bansal, Juan C. Rey, Myung-Soo Jang, L. C. Lu, Philippe Magarshack, Paul Marchal, Riko Radojcic
DAC7
2010 An RDL-configurable 3D memory tier to replace on-chip SRAM
abstract
In a conventional SoC designs, on-chip memories occupy more than the 50% of the total die area. 3D technology enables the distribution of logic and memories on separate stacked dies (tiers). This allows redesigning the memory tier as a configurable product to be used in multiple system designs. Previously proposed dynamic re-configurable solutions demonstrate strong dependence between read latency and dimensions of the mapped memory, leading to potential performance limitations. In this paper we propose a one-time configurable memory tier designed to minimize the performances overhead due to the commodity. Flexible configuration is enabled by smart memory macros and I/Os organization and a customizable redistribution layer routing. With respect to the dynamic re-configurability, the proposed design offers up to 40% faster access time, while saving more than 10% of energy per access. In addition production cost trade offs are analyzed.
Marco Facchini, Paul Marchal, Francky Catthoor, Wim Dehaene
DATE2
2010 3D integration: Circuit design, test, and reliability challenges
abstract
3D-Stacked ICs (3D-SIC) based on Through-Silicon Vias (TSVs) offer alleviation of the performance and interconnect density bottlenecks faced by traditional CMOS scaling. As a result there is a lot of industrial focus to make this technology available for the next generation of SoCs. However, for 3D integration to become a viable product approach, it requires that the additional processing steps necessary preserve the integrity of both front-end and back-end of devices and constituting materials. 3D processing steps such as TSV insertion and wafer thinning, have an impact on the functionality and performance of analog and digital circuits, which needs to be accounted for during the design phase. Moreover, testing 3D-SICs calls for more complex test flow trade-offs and enhanced design-for-test architectures for test access within the stack. Finally, the reliability consequences with respect to thermal and mechanical stress in dense stacks of thinned wafers need to be carefully assessed to guarantee a target product life time. In this presentation we discuss the above mentioned challenges and some of the emerging solutions.
Nikolaos Minas, Ingrid De Wolf, Erik Jan Marinissen, Michele Stucchi, Herman Oprins, Abdelkarim Mercha, Geert Van der Plas, Dimitrios Velenis, Paul Marchal
IOLTS9
2010 3D NoCs - Unifying inter & intra chip communication
abstract
Networks-on-chip have been developed in the last few years to address the scalability challenges of global on-chip communication. VLSI technology is now rapidly moving into vertical stacking to overcome fundamental communication and integration bottlenecks, however this technology is not mature yet, and significant reliability challenges must be overcome. In this paper we describe our effort in establishing a 3DNoC design flow and in designing circuits and architectural solutions for variability and reliability characterization and tolerance.
Igor Loi, Paul Marchal, Antonio Pullini, Luca Benini
ISCAS2
2009 System-level power/performance evaluation of 3D stacked DRAMs for mobile applications
abstract
Convergence of communication, consumer applications and computing within mobile systems pushes memory requirements both in terms of size, bandwidth and power consumption. The existing solution for the memory bottle-neck is to increase the amount of on-chip memory. However, this solution is becoming prohibitively expensive, allowing 3D stacked DRAM to become an interesting alternative for mobile applications. In this paper, we examine the power/performance benefits for three different 3D stacked DRAM scenarios. Our high-level memory and Through Silicon Via (TSV) models have been calibrated on state-of-the-art industrial processes. We model the integration of a logic die with TSVs on top of both an existing DRAM and a DRAM with redesigned transceivers for 3D. Finally, we take advantage of the interconnect density enabled by 3D technology to analyze an ultra-wide memory interface. Experimental results confirm that TSV-based 3D integration is a promising technology option for future mobile applications, and that its full potential can be unleashed by jointly optimizing memory architecture and interface logic.
Marco Facchini, Trevor E. Carlson, Anselme Vignon, Martin Palkovic, Francky Catthoor, Wim Dehaene, Luca Benini, Paul Marchal
DATE8
2009 A novel DRAM architecture as a low leakage alternative for SRAM caches in a 3D interconnect context
abstract
This paper presents a DRAM architecture that improves the DRAM performance/power trade-off to increase their usability on low power chip design using 3D interconnect technology. The use of a finer matrix subdivision and buffering the bitline signal at the localblock level allows to reduce both the energy per access and the access time. The obtained performances match those of a typical low power SRAM, while achieving a significant area and static power reduction compared to these memories. The 128 kb memory architecture proposed here achieves an access time of 1.3 ns for a dynamic energy of less than 0.2 pJ per bit. A localized refresh mechanism allows gaining a factor of 10 in static power consumption associated with the cell, and a factor of 2 in area, when compared with an equivalent SRAM.
Anselme Vignon, Stefan Cosemans, Wim Dehaene, Paul Marchal, Marco Facchini
DATE4
2009 3-D Technology Assessment: Path-Finding the Technology/Design Sweet-Spot
abstract
It is widely acknowledged that three-dimensional (3-D) technologies offer numerous opportunities for system design. In recent years, significant progress has been made on these 3-D technologies, and they have become probably the best hope for carrying the semiconductor industry beyond the path of Moore's law. However, a clear roadmap is missing to successfully introduce this 3-D technology onto the market. Today, a plurality of 3-D technology options exists, which requires different design and test strategies. To crystallize the many technology options in a few mainstream technologies, it is mandatory to coexplore both technology and design options. The contribution of this paper is to introduce a novel path finding methodology to untangle the many intertwined design/technology options. This holistic approach will be applied on a representative 3-D case study. Initial results demonstrate the benefits of the proposed path-finding methodology to steer the technology development and fine-tune design strategies.
Paul Marchal, Bruno Bougard, Guruprasad Katti, Michele Stucchi, Wim Dehaene, Antonis Papanikolaou, Diederik Verkest, Bart Swinnen, Eric Beyne
Proc. IEEE1
2008 HOT TOPIC - 3D Integration or How to Scale in the 21st Century
abstract
Summary form only given. 3D integration offers numerous opportunities for design, and is probably the best hope for carrying ICs along (and even beyond) the path of Moore's Law in the 21st century. However, many questions still need to be answered to take advantage of 3D. First, what will become the mainstream 3D technology? Today, many technology options are proposed, but each having different cost, design and test implications. Secondly, how to make 3D designs reliable? Many unknowns still exist related to thermal load, reliability and signal integrity challenges. Finally, what about design solutions/methods and architectural modifications for 3D integration? The objective of this special session is to create a better understanding of forthcoming 3D technologies, their implication on design and test. An attempt will be made to roadmap 3D technologies and their design implications. This will enable R&D planning by design houses, EDA vendors, foundries and academia, paving the way for a widespread acceptance of 3D technologies.
Bruno Bougard, Paul Marchal, Luca Benini, Doris Keitel-Schulz, Neal Checka
DATE2
2008 How to Live with Uncertainties: Exploiting the Performance Benefits of Self-Timed Logic In Synchronous Design
abstract
Ultra low power digital systems are key for any future wireless sensor nodes but also inside nomadic embedded systems (such as inside the digital front end of software defined radios). These systems require the highest possible energy efficiency of logic, which can only be achieved by operating in moderate inversion. Unfortunately, when operating near the threshold voltage, transistors become highly sensitive to process variations, thereby increasing leakage currents and complicating timing closure. Rather than pursuing a worst-case design approach for dealing with these uncertainties, we present a hybrid self-timed/synchronous approach. It will be demonstrated on the VEX VLIW core designed for ultra low-power operations. Experimental results of our approach demonstrate performance benefits up to 2times and significant energy savings at low throughput rates.
Giacomo Paci, Axel Nackaerts, Francky Catthoor, Luca Benini, Paul Marchal
DSD5
2007 Exploration of Low Power Adders for a SIMD Data Path
abstract
Hardware for ambient intelligence needs to achieve extremely high computational efficiency (up to 40GOPS/W). An important way for reaching this is exploiting parallelism, and more specifically data-level parallelism enabled by SIMD. Whereas a large body of research exists on the benefits of the architectural design of and compilation onto SIMD, the design of energy-optimal functional units for SIMD has received limited attention. It appears that existing SIMD functional units are designed in an area optimal, but not energy optimal way. By exploiting the difference in critical path length for the types of operations (e.g., 4times8/2times16/1times32), SIMD adders can be developed that save up to 40% of energy. In this paper, the authors present these adders, the issues of building them and quantify their benefits for different usage scenarios and operating frequencies.
Giacomo Paci, Paul Marchal, Luca Benini
ASP-DAC2
2007 Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support
abstract
In today's multiprocessor SoCs (MPSoCs), parallel programming models are needed to fully exploit hardware capabilities and to achieve the 100 Gops/W energy efficiency target required for ambient intelligence applications. However, mapping abstract programming models onto tightly power-constrained hardware architectures imposes overheads which might seriously compromise performance and energy efficiency. The objective of this work is to perform a comparative analysis of message passing versus shared memory as programming models for single-chip multiprocessor platforms. Our analysis is carried out from a hardware-software viewpoint: we carefully tune hardware architectures and software libraries for each programming model. We analyze representative application kernels from the multimedia domain, and identify application-level parameters that heavily influence performance and energy efficiency. Then, we formulate guidelines for the selection of the most appropriate programming model and its architectural support
Francesco Poletti, Antonio Poggiali, Davide Bertozzi, Luca Benini, Paul Marchal, Mirko Loghi, Massimo Poncino
IEEE Trans. Computers5
2006 Physical design implementation of segmented buses to reduce communication energy
abstract
The amount of energy consumed for interconnecting the IP-blocks is increasing significantly due to the suboptimal scaling of long wires. To limit this energy penalty, segmented buses have gained interest in the architectural community. However, the netlist topology and the physical design stage significantly influence the final communication energy cost. We present in this paper an automated way to implement a netlist consisting of hard macro blocks, which are interconnected with heavily segmented buses in an energy optimal fashion for communication. We optimize the network wires energy dissipation in two separate, but related steps: minimizing the number of segments for active communication paths at the first step (block ordering), followed by the activity aware floorplanning step to minimize the physical length of these segments. Energy gains of up to a factor of 4 are achieved compared to a standard system implementation using a shared bus. Especially, the block ordering step contributes significantly to the network energy optimization process.
Jin Guo 0001, Antonis Papanikolaou, Paul Marchal, Francky Catthoor
ASP-DAC3
2006 Exploring "temperature-aware" design in low-power MPSoCs
abstract
The power density inside high performance systems continues to rise with every process technology generation, thereby increasing the operating temperature and creating "hot spots" on the die. As a result, the performance, reliability and power consumption of the system degrade. To avoid these "hot spots", "temperature-aware" design has become a must. For low-power embedded systems though, it is not clear whether similar thermal problems occur. These systems have very different characteristics from the high performance ones: they consume hundred times less power, they are based on a multi-processor architecture with lots of embedded memory and rely on cheap packaging solutions. In this paper, we investigate the need for temperature- aware design in low-power systems-on-a-chip and provide guidlines to delimit the conditions for which temperature-aware design is needed
Giacomo Paci, Paul Marchal, Francesco Poletti, Luca Benini
DATE2
2006 Reliability issues in deep deep sub-micron technologies: time-dependent variability and its impact on embedded system design
abstract
Technology scaling has traditionally offered advantages to embedded system design in terms of reduced energy consumption and cost and increased performance. Scaling past the 45 nm technology node, however, brings a host of problems, whose impact on system-level design has not been evaluated. Random intra-die process variability, reliability and their combined impact on the system level parametric quality metrics are effects that are gaining prominence and that needs to be tackled in the next few years. Dealing with these new challenges requires a paradigm shift in the system level design phase
Antonis Papanikolaou, Miguel Corbalan, Francky Catthoor, M. Satyakiran, Paul Marchal, Ben Kaczer, C. Bruynseraede, Zsolt Tokei
VLSI-SoC6
2005 Flexible Hardware/Software Support for Message Passing on a Distributed Shared Memory Architecture
abstract
With the advent of multiprocessor systems on a chip, the interest for message passing libraries has revived. Message passing helps in mastering the design complexity of parallel systems. However to satisfy the stringent energy-budget of embedded applications, the message passing overhead should be limited. Recently, several hardware extensions have been proposed for reducing the transfer cost on a distributed memory architecture. Unfortunately, they ignore the synchronization cost between sender/receiver and/or require many dedicated hardware blocks. To overcome the above limitations, we present in this paper lightweight support for message passing. Moreover, we have made our library as flexible as possible such that we can optimally match the application with the target architecture. We demonstrate the benefits of our approach by means of representative benchmarks from the multimedia domain.
Francesco Poletti, Antonio Poggiali, Paul Marchal
DATE3
2004 Optimizing the Memory Bandwidth with Loop Morphing
José Ignacio Gómez, Paul Marchal, Sven Verdoolaege, Luis Piñuel, Francky Catthoor
ASAP2
2004 An integrated hardware/software approach for run-time scratchpad management
abstract
An ever increasing number of dynamic interactive applications are implemented on portable consumer electronics. Designers depend largely on operating systems to map these applications on the architecture. However, today's embedded operating systems abstract away the precise architectural details of the platform. As a consequence, they cannot exploit the energy efficiency of scratchpad memories. We present in this paper a novel integrated hardware/software solution to support scratchpad memories at a high abstraction level. We exploit hardware support to alleviate the transfer cost from/to the scratchpad memory and at the same time provide a high-level programming interface for run-time scratchpad management. We demonstrate the effectiveness of our approach with a case-study.
Francesco Poletti, Paul Marchal, David Atienza 0001, Luca Benini, Francky Catthoor, Jose Manuel Mendias
DAC2
2003 SDRAM-Energy-Aware Memory Allocation for Dynamic Multi-Media Applications on Multi-Processor Platforms
Paul Marchal, José Ignacio Gómez, Luis Piñuel, Davide Bruni, Luca Benini, Francky Catthoor, Henk Corporaal
DATE1
2001 Task concurrency management methodology summary
abstract
This paper summarizes a new methodology for the design of concurrent dynamic real-time embedded systems. An embedded system can be specified at a grey-box abstraction level in a combined MTG-CDFG model. The authors believe that task concurrency management can be implemented in four major steps. Firstly, the grey box model is built, including the necessary concurrency extraction. Then transformations are applied on the specified MTG-CDFG to increase the opportunities for concurrency exploration and cost minimization. Then static scheduling will be applied on the design time analyzable parts of the grey-box model, including processor assignment in the multiple processor context. Finally, a dynamic scheduler will schedule the dynamic and coarse-grain constructs at run time on the given platform while making trade-offs based on Pareto curves.
Chun Wong, Paul Marchal, Francky Catthoor, Hugo De Man, Aggeliki S. Prayati, Nathalie Cossement, Rudy Lauwereins, Diederik Verkest
DATE2
2001 Optimisation Problems for Dynamic Concurrent Task-Based Systems
abstract
One of the most critical bottlenecks in many novel multi-media applications is their very dynamic concurrent behaviour. This is especially true because of the quality-of-service (QoS) aspects of these applications. Prominent examples of this can be found in MPEG4 and JPEG2000 standards and especially the new MPEG21 standard. In order to deal with these dynamic issues where tasks and complex data types are created and deleted at run-time based on non-deterministic events, a novel system design paradigm is required. Because of the high computation and communicating requirement from these kind of applications, multi-processor SoC (system on chip) is accepted as a solution (e.g., TriMedia), which is also promised by the advance of the processing technology. It is different from traditional general-purpose computers because it is application specific and it is an embedded system, which means costs like energy consumption are of major concern. This paper focused on the new requirements in system-level synthesis. In particular a "task concurrency management" problem formulation is proposed, with special focus on the formal definition of the most crucial optimization problems. The concept of Pareto curve based exploration is crucial in these formulations.
Diederik Verkest, Chun Wong, Paul Marchal
ICCAD4