VLDB 2026 Research / reviewers in the wild / expert
Dragomir Milojevic
dblp:83/3348
· DBLP profile ↗
28ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-5915-5160ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Focus Session: Advanced CMOS and 5.5D Packaging: Perspectives and Challenges for Design, Reliability and SecurityabstractAI workloads are driving exceptional demand for performance and energy efficiency, forcing semiconductor innovation to advance along two major directions simultaneously. On the device roadmap, the transition from FinFETs to gate-all-around nanosheet FETs and, subsequently, monolithic 3D Complementary FETs (CFETs) is enabling scaling toward the 2 nm era and beyond while targeting aggressive logic density. In parallel, advanced packaging, spanning 2.5D integration on silicon interposers, true 3D stacking, and hybrid 5.5D assemblies, is becoming essential to deliver ultra-high bandwidth, low-energy die-to-die connectivity required by rapidly growing AI model sizes and the resulting memory-wall bottlenecks. This focus session discusses the opportunities and challenges of this co-evolution, with emphasis on system-technology co-optimization and the inevitable need for multiphysics analysis across electrical, thermal, mechanical, and reliability domains. We highlight how reliability and security concerns are increasingly shaping architectural and packaging choices, and we discuss the role of deep learning as a practical enabler for faster simulation and design-space exploration under rising complexity. Hussam Amrouch, Dragomir Milojevic, Giorgio Di Natale, Jérôme Toublanc |
DATE | 2 |
| 2025 | Bandwidth-Latency-Thermal Co-Optimization of Interconnect-Dominated Many-Core 3D-ICabstractThe ongoing integration of advanced functionalities in contemporary system-on-chips (SoCs) poses significant challenges related to memory bandwidth, capacity, and thermal stability. These challenges are further amplified with the advancement of artificial intelligence (AI), necessitating enhanced memory and interconnect bandwidth and latency. This article presents a comprehensive study encompassing architectural modifications of an interconnect-dominated many-core SoC targeting the significant increase of intermediate, on-chip cache memory bandwidth and access latency tuning. The proposed SoC has been implemented in 3-D using A10 nanosheet technology and early thermal analysis has been performed. Our workload simulations reveal, respectively, up to 12- and 2.5-fold acceleration in the 64-core and 16-core versions of the SoC. Such speed-up comes at 40% increase in die-area and a 60% rise in power dissipation when implemented in 2-D. In contrast, the 3-D counterpart not only minimizes the footprint but also yields 20% power savings, attributable to a 40% reduction in wirelength. The article further highlights the importance of pipeline restructuring to leverage the potential of 3-D technology for achieving lower latency and more efficient memory access. Finally, we discuss the thermal implications of various 3-D partitioning schemes in High Performance Computing (HPC) and mobile applications. Our analysis reveals that, unlike high-power density HPC cases, 3-D mobile case increases$T_{\max }$only by$2~^{\circ } $C–$3~^{\circ } $C compared to 2-D, while the HPC scenario analysis requires multiconstrained efficient partitioning for 3-D implementations. Sudipta Das, Samuel Riedel, Mohamed Naeim, Moritz Brunion, Marco Bertuletti, Luca Benini, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2024 | GNN-assisted Back-side Clock Routing Methodology for Advance TechnologiesabstractThe back-side metal layers exhibit lower parasitics compared to the front-side layers in advanced technologies, making them suitable for clock-net distribution. In this study, we explore the advantages of using back-side metal layers for clock routing, which is shared with a power delivery network. Our Graph Neural Network (GNN) based framework, effectively distributes the clock-tree between the front and back sides. We address the back-side clock nets' creation by incorporating back-side buffers. Our results demonstrate better clock and full-chip metrics represented by an increase of up to 13% in the effective frequency with equivalent power consumption, using 3 nm technology. Nesara Eranna Bethur, Pruek Vanna-Iampikul, Odysseas Zografos, Lingjun Zhu, Giuliano Sisto, Dragomir Milojevic, Alberto García Ortiz, Geert Hellings, Julien Ryckaert, Francky Catthoor, Sung Kyu Lim |
DAC | 6 |
| 2024 | 3D Partitioning with Pipeline Optimization for Low-Latency Memory Access in Many-Core SoCsabstractThis paper presents an investigation of System-on-Chip (SoC) communication latency optimization for 3D system integration and highlights the role of architectural modifications to maximize the Power, Performance, & Area (PPA) benefits. An instance of a highly configurable RISC-V SoC is implemented using ∼2nm nanosheet technology and different 3D stacking options using design flow from sign-off tools. The proposed implementation targets performance optimization for different 3D partitioning scenarios: Memory-on-Logic (MoL) & Logic-on-Logic (LoL). We target 2-die 3D Integrated Circuits (3D-IC) with high density 3D interconnect using Face-to-Face (F2F) hybrid bonding (∼1µm), and 3-die stack, as Face-to-Back (F2B) on top of F2F. Our analysis of the 16-core SoC instance shows that the proposed architectural optimizations bring a significant reduction of 4 pipeline stages in the design hierarchy at a marginal cost of 9% effective frequency loss when implemented in 3D in comparison to the baseline 2D architecture. Further, going from 2D to 3D allows more than 40% total system wire-length reduction & 10% less cell area, resulting in 20% power savings. These findings hold promise for further explorations on many-core SoC instances (256 & more) facing system interconnect challenges. Sudipta Das, Samuel Riedel, Marco Bertuletti, Luca Benini, Moritz Brunion, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic |
ISCAS | 9 |
| 2024 | Hier-3D: A Methodology for Physical Hierarchy Exploration of 3-D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. Complemented with an automated floorplanning solution, the flow allows for system-level physical and architectural exploration of 3D designs. As a result, we significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 -2.2× compared with 2D, where all metrics are improved simultaneously, including up to 20% power savings. Nesara Eranna Bethur, Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Multidie 3-D Stacking of Memory Dominated Neuromorphic ArchitecturesabstractEvent-driven neuromorphic processors for artificial intelligence (AI) inference on edge/IoT devices require largeon-chip memory capacity, for efficient execution of spiking neural networks (NNs). In this work, we evaluate 3-D stacking benefits on SENECA, a digital neuromorphic accelerator core, sweeping itson-chip memory capacity from 2 up to 32 Mb in both legacy planar and advanced nanosheet CMOS logic nodes. In a planar CMOS node (GF-22 nm), two-die memory-on-logic (MoL) partitioning enables$8\times $moreon-chip memory, and it boosts operating frequency by 7% with 26% less power than the 2-D. Moving to an advanced nanosheet technology (imec A10), multidie (up to 7 dies) MoL stacking enables a performance increase of up to 29% and power savings up to 31%. Furthermore, a core folding (CF) partitioning in A10 shows up to 16% performance improvement with 12% total power savings with respect to the 2-D implementation on the same technology. We also demonstrate no thermal overhead for multidie stacking at advanced nodes for designs exhibiting low power density. These physical design explorations lay the foundation for system technology co-optimization studies for edge devices. Leandro M. G. Rocha, Refik Bilgic, Mohamed Naeim, Sudipta Das, Herman Oprins, Amirreza Yousefzadeh, Mario Konijnenburg, Dragomir Milojevic, James Myers, Julien Ryckaert, Dwaipayan Biswas |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2022 | MemPool-3D: Boosting Performance and Efficiency of Shared-L1 Memory Many-Core Clusters with 3D IntegrationabstractThree-dimensional integrated circuits promise power, performance, and footprint gains compared to their 2D counter-parts, thanks to drastic reductions in the interconnects' length through their smaller form factor. We can leverage the potential of 3D integration by enhancing MemPool, an open-source many-core design with 256 cores and a shared pool of L1 scratchpad memory connected with a low-latency interconnect. MemPool's baseline 2D design is severely limited by routing congestion and wire propagation delay, making the design ideal for 3D integration. In architectural terms, we increase MemPool's scratchpad memory capacity beyond the sweet spot for 2D designs, improving performance in a common digital signal processing kernel. We propose a 3D MemPool design that leverages a smart partitioning of the memory resources across two layers to balance the size and utilization of the stacked dies. In this paper, we explore the architectural and the technology parameter spaces by analyzing the power, performance, area, and energy efficiency of MemPool instances in 2D and 3D with 1 MiB, 2 MiB, 4 MiB, and 8 MiB of scratchpad memory in a commercial 28 nm technology node. We observe a performance gain of 9.1% when running a matrix multiplication on MemPool-3D with 4 MiB of scratchpad memory compared to the MemPool 2D counterpart. In terms of energy efficiency, we can implement the MemPool-3D instance with 4 MiB of L1 memory on an energy budget 15 % smaller than its 2D counterpart, and 3.7 % smaller than the MemPool-2D instance with a fourth of the L1 scratchpad memory capacity. Matheus A. Cavalcante, Anthony Agnesina, Samuel Riedel, Moritz Brunion, Alberto García Ortiz, Dragomir Milojevic, Francky Catthoor, Sung Kyu Lim, Luca Benini |
DATE | 6 |
| 2022 | Hier-3D: A Hierarchical Physical Design Methodology for Face-to-Face-Bonded 3D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. We significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 to 2.2 × compared with 2D, where all metrics are improved simultaneously, including up to power savings. Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
ISLPED | 5 |
| 2022 | Evaluation of Nanosheet and Forksheet Width Modulation for Digital IC Design in the Sub-3-nm EraabstractIn this article, we provide a comprehensive evaluation of width modulation capabilities of both nanosheet (NS) and forksheet (FS) devices, going from device level to a block level implementation. The main innovation introduced by the FS consists of a dielectric wall added between the p- and nMOS transistors. Leveraging this feature, FS shows approximately the same current behavior as NS, considered a state-of-the-art reference, but reduced parasitic capacitance thanks to its fewer but wider stacked sheets. At block level, an area reduction up to 12% is observed with FS, alongside a 13% power reduction and 10% frequency increase. Following the device comparison, the potential of sheet width modulation as additional power, performance, and area (PPA) optimization technique during synthesis and place and route (PNR) is investigated. A description of the specific steps required to enable this knob in a conventional electronic design automation (EDA) framework is provided. As demonstrated by the obtained experimental results, the same frequency of the single-width implementation can be achieved using mixed libraries with lower power consumption (13% and 16% for NS and FS, respectively), leading to improved energy efficiency. Furthermore, it is shown how designs implemented using FS benefit more from this type of optimization than the ones using NS, with a 12%–15% energy reduction compared to the 8.5%–14% obtained with NS. Giuliano Sisto, Odysseas Zografos, Bilal Chehab, Naveen Kakarla, Dragomir Milojevic, Pieter Weckx, Geert Hellings, Julien Ryckaert |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2021 | Power, Performance, Area and Cost Analysis of Memory-on-Logic Face-to-Face Bonded 3D Processor DesignsabstractIn this paper, we present a power, performance, area and cost (PPAC) analysis for large-scale 3D processor designs based on wafer-to-wafer bonding. From the evaluation of our cost model, we investigate a typically disregarded opportunity in 3D that is area savings due to buffer savings and better routability, offering unexpected cost savings. We explore the viability of this factor with the feedback of a state-of-the-art 3D memory-on-logic implementation flow. We show how this affects the PPAC of full-chip GDS implementations of a large-scale manycore processor design. Experiments show that our memory-on-logic 3D implementation offers 7% silicon area savings, resulting in 53.5% footprint reduction. We also obtain a 40% power-performance-cost improvement compared with 2D counterparts Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Dragomir Milojevic, Francky Catthoor, Manu Perumkunnil Komalan, Sung Kyu Lim |
ISLPED | 5 |
| 2021 | High-Performance Logic-on-Memory Monolithic 3-D IC Designs for Arm Cortex-A ProcessorsabstractMonolithic 3-D IC (M3-D) is a promising solution to improve the performance and energy-efficiency of modern processors. But, designers are faced with challenges in design tools and methodologies, especially for power and thermal verifications. We developed a new physical design flow that optimally places and routes cache modules in one tier and logic gates in the other. Our tool also builds high-quality clock and power delivery networks targeting logic-on-memory M3-D designs. Finally, we developed a sign-off analysis tool flow to evaluate power, performance, area (PPA), thermal, and voltage-drop quality for given M3-D designs. Using our complete register transfer level (RTL)-to-Graphic Design System (GDS) tool flow, we designed commercial quality 2-D and M3-D implementation of Arm Cortex-A7 and Cortex-A53 processors in a commercial 28-nm technology. Experimental results show that our 3-D processors offer 20% (A7) and 21% (A53) performance gain, compared with their 2-D commercial counterparts. The voltage-drop degradation of our 3-D Cortex-A7 and Cortex-A53 processors is less than 3% of the supply voltage, while temperature increase is 10.71 °C and 13.04 °C, respectively. Lingjun Zhu, Lennart Bamberg, Sai Pentapati, Kyungwook Chang, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Brian Cline, Saurabh Sinha 0001, Alberto García Ortiz, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | 3D-Stacked Integrated Circuits: How Fine Should System Partitioning Be?abstract3D stacked ICs package multiple, independently manufactured dies to reduce total system wire-length, improve timing, and reduce area and power. When designing stacked 3D-ICs, arises the question of the grain at which one should consider system partitioning to optimize the gains. This work uses known MAX-CUT graph partitioning algorithms to split designs from 42k up to 800k gates, with gates clustered from 8 and up to 32768 partitions. It has been found that with 2048 clusters, i.e. 20 to 400 gates per cluster depending on the design, a partitioning of the system allows on average to cut 35% of the nets that account for 73% of the total wire-length in 3D. Quentin Delhaye, Dragomir Milojevic, Joël Goossens |
ISCAS | 2 |
| 2017 | IR-drop aware Design & technology co-optimization for N5 node with different device and cell height optionsabstractIn this paper we propose a novel Design-Technology Co-Optimization (DTCO) framework that enables PDK generation and design implementation of sub-10nm technology nodes. The framework allows to study the impact of different technology options at design level and use effective design Power, Performance and Area (PPA) to decide on right technology option. Design implementation flow is IR-drop aware, allowing integration of optimized Power Delivery Network (PDN) for different device/cell options. Using N5-like technology node assumptions (contacted poly and metallization pitch of 42 and 32nm), we generate digital PDKs for different device (finFET, 2 & 3 nanowires) and standard cell options (3, 2 or 1 fins & 7.5 or 6-Tracks cell height). Different PDKs have been used to implement and characterize a wire dominated circuit. Our study shows that the design PDN/IR-drop awareness is fundamental to complete DTCO approach for sub-10nm nodes. Using our dedicated design methodology we reach the IR-drop target of 2.5% VDD (on the lowest metal layers), while minimizing the area degradation induced by the PDN. Further, we demonstrate that such optimized PDN is mandatory to enable the 20% area gain when moving from 7.5 to 6-Tracks cell height. Finally, we show that the impact of different device options is in range of 15% Power, 2X Performance and 20% Area, further validating the need of a fully integrated DTCO. Luca Mattii, Dragomir Milojevic, Peter Debacker, Yasser Sherazi, Mladen Berekovic, Praveen Raghavan |
ICCAD | 2 |
| 2016 | Energy minimization at all layers of the data center: The ParaDIME project
Oscar Palomar, Santhosh Kumar Rethinagiri, Gulay Yalcin, J. Rubén Titos Gil, Pablo Prieto, Emma Torrella, Osman S. Unsal, Adrián Cristal, Pascal Felber, Anita Sobe, Yaroslav Hayduk, Mascha Kurpicz, Christof Fetzer, Thomas Knauth, Malte Schneegaß, Jens Struckmeier, Dragomir Milojevic |
DATE | 17 |
| 2016 | How much cost reduction justifies the adoption of monolithic 3D ICs at 7nm node?abstractIn this paper we study power, performance, and cost (PPC) tradeoffs for 2-tier, gate-level, full-chip GDS monolithic 3D ICs (M3D) built using a foundry-grade 7nm bulk FinFET technology. We first develop highly-accurate wafer and die cost models for 2D and M3D to study PPC tradeoffs. In our study, both 2D and M3D designs are optimized in terms of the number of BEOL metal layers used for routing to obtain the best possible PPC values. We develop a new CAD methodology for 2-tier gate-level M3D, named Projected 2D Flow, that allows us to accurately compare RC parasitics of equivalent nets in both 2D and M3D designs. Our experiments based on two different circuit types (BEOL-dominant vs. FEOL-dominant) confirm that M3D designs indeed offer a significant footprint saving. However, to our surprise, the PPC quality of M3D turns out to be worse than that of 2D by 34% due to the high wafer cost of M3D. Our study also reveals that M3D wafer yield should be as high as 90% of 2D wafer yield, and the M3D device manufacturing cost should be less than 33% of that of 2D to justify the adoption of M3D technology at the 7nm era. Lastly, and counter-intuitively, our study shows that FEOL-dominant circuit shows more PPC benefits from M3D technology than BEOL-dominant circuit. Bon Woong Ku, Peter Debacker, Dragomir Milojevic, Praveen Raghavan, Sung Kyu Lim |
ICCAD | 3 |
| 2016 | Physical Design Solutions to Tackle FEOL/BEOL Degradation in Gate-level Monolithic 3D ICsabstractIn this paper, we develop physical design tools and methodologies to tackle the inter-tier performance variations caused by low temperature manufacturing in 2-tier gate-level monolithic 3D ICs (M3D). First, we model the top tier front-end-of-line (FEOL) device mobility degradation and its impact on cell delay/power values. Next, we quantify the impact of tungsten interconnect and cost-driven metal layer saving in the back-end-of-line (BEOL) of the bottom tier. These device and interconnect degradation models are used in our new full-chip M3D physical design flow named Derated 2D. This flow overcomes the well-known drawback of the state-of-the-art Shrunk 2D that requires shrinking of layout objects and RC parasitics. Also, Derated 2D performs low-temperature process-aware tier partitioning to effectively keep timing-critical components in the bottom tier. Moreover, Derated 2D conducts timing-driven monolithic inter-tier via (MIV) planning to cope with the resistivity increase in tungsten BEOL. Lastly, Derated 2D offers an effective timing closure solution through a post-route optimization. Experiments based on a foundry-grade 7nm FinFET process design kit (PDK) show that Derated 2D achieves up to 36% performance improvement and 10% energy saving compared with Shrunk 2D. Using a post-route optimization, Derated 2D further improves timing under various FEOL/BEOL degradation settings at a minimum energy overhead. Bon Woong Ku, Peter Debacker, Dragomir Milojevic, Praveen Raghavan, Diederik Verkest, Aaron Thean, Sung Kyu Lim |
ISLPED | 3 |
| 2016 | Exploring Energy Reduction in Future Technology Nodes via Voltage Scaling with Application to 10nmabstractDI-fusion, le Dépôt institutionnel numérique de l'ULB, est l'outil de référencementde la production scientifique de l'ULB.L'interface de recherche DI-fusion permet de consulter les publications des chercheurs de l'ULB et les thèses qui y ont été défendues. Gulay Yalcin, Santhosh Kumar Rethinagiri, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Dragomir Milojevic |
PDP | 6 |
| 2014 | A context aware cache controller to bridge the gap between theory and practice in real-time systemsabstractNowadays, most processing platforms make use of cache memories to improve the execution speed of the tasks running on the processors. However, when a processor switches from a task to another, the caches must be reloaded with the context of the upcoming task. This is time consuming and is usually not predictable and thus affects the worst-case execution time of the task. Such unpredictability should be avoided in real-time systems in which the instant at which a result is available is as important as the result itself. In this paper, we present a hardware component named hardware context switch (HwCS) which replaces the standard L1 cache controller of a processor. It divides the cache in two interchangeable layers and enables to save or restore the content of one layer while the second is simultaneously used as a usual cache by the processor. Saving the cache content after a preemption and restoring this content before resuming the execution of the preempted task, makes the preemption overheads negligible in comparison to the task worst-case execution times. It is theoretically proven that the existing scheduling theory can be used “as is” with the HwCS by simply reducing the task deadlines, thereby bridging the gap between theory and practice. The HwCS has been implemented in an uniprocessor system as a proof of concept. The first results show a neat improvements on the processor utilisation for a small cost in silicon surface. Yannick Allard, Geoffrey Nelissen, Joël Goossens, Dragomir Milojevic |
RTCSA | 4 |
| 2013 | Design issues in heterogeneous 3D/2.5D integrationabstractEfficient processing of fine-pitched Through Silicon Vias, micro-bumps and back-side re-distribution layers enable face-to-back or face-to-face integration of heterogeneous ICs using 3D stacking and/or Silicon Interposers. While these technology features are extremely compelling, they considerably stress the existing design practices and EDA tool flows typically conceived for 2D systems. With all system, technology and implementation level options brought with these features, the design space increases to an extent where traditional 2D tools cannot be used any more for efficient exploration. Therefore, the cost-effective design of future 3D ICs products will require new planning and co-optimisation techniques and tools that are fast and accurate enough to cope with these challenges. In this paper we present design methodology and the practical EDA tool chain that covers different aspects of the design flow and is specific to efficient design of 3D-ICs. Flow features include: fast synthesis and 3D design partitioning at gate level, TSV/micro-bump array planning, 3D floor planning, placement and routing, congestion analysis, fast thermal and mechanical modeling, easy technology vs. implementation trade-off analysis, 3D device models generations and Design-for-Test (DfT). The application of the tool chain is illustrated using concrete example of a real-world design, showing not only the applicability of the tool chain, but also the benefits of heterogeneous 2.5 and 3D integration technologies. Dragomir Milojevic, Paul Marchal, Erik Jan Marinissen, Geert Van der Plas, Diederik Verkest, Eric Beyne |
ASP-DAC | 1 |
| 2013 | Panel: "will 3D-IC remain a technology of the future... even in the future?"
Marco Casale-Rossi, Patrick Leduc, Giovanni De Micheli, Patrick Blouet, Brendan Farley, Anna Fontanelli, Dragomir Milojevic |
DATE | 7 |
| 2012 | Techniques Optimizing the Number of Processors to Schedule Multi-threaded TasksabstractThese last years, we have witnessed a dramatic increase in the number of cores available in computational platforms. Concurrently, a new coding paradigm dividing tasks into smaller execution instances called threads, was developed to take advantage of the inherent parallelism of multiprocessor platforms. However, only few methods were proposed to efficiently schedule hard real-time multi-threaded tasks on multiprocessor. In this paper, we propose techniques optimizing the number of processors needed to schedule such sporadic parallel tasks with constrained deadlines. We first define an optimization problem determining, for each thread, an intermediate (artificial) deadline minimizing the number of processors needed to schedule the whole task set. The scheduling algorithm can then schedule threads as if they were independent sequential sporadic tasks. The second contribution is an efficient and nevertheless optimal algorithm that can be executed online to determine the thread's deadlines. Hence, it can be used in dynamic systems were all tasks and their characteristics are not known a priori. We finally prove that our techniques achieve a resource augmentation bound of 2 when the threads are scheduled with algorithms such as U-EDF, PD2, LLREF, DP-Wrap, etc. Geoffrey Nelissen, Vandy Berten, Joël Goossens, Dragomir Milojevic |
ECRTS | 4 |
| 2012 | U-EDF: An Unfair But Optimal Multiprocessor Scheduling Algorithm for Sporadic TasksabstractA multiprocessor scheduling algorithm named U-EDF, was presented in [1] for the scheduling of periodic tasks with implicit deadlines. It was claimed that U-EDF is optimal for periodic tasks (i.e., it can meet all deadlines of every schedulable task set) and extensive simulations showed a drastic improvement in the number of task preemptions and migrations in comparison to state-of-the-art optimal algorithms. However, there was no proof of its optimality and U-EDF was not designed to schedule sporadic tasks. In this work, we propose a generalization of U-EDF for the scheduling of sporadic tasks with implicit deadlines, and we prove its optimality. Contrarily to all other existing optimal multiprocessor scheduling algorithms for sporadic tasks, U-EDF is not based on the fairness property. Instead, it extends the main principles of EDF so that it achieves optimality while benefiting from a substantial reduction in the number of preemptions and migrations. Geoffrey Nelissen, Vandy Berten, Vincent Nélis, Joël Goossens, Dragomir Milojevic |
ECRTS | 5 |
| 2012 | Thermal characterization of cloud workloads on a power-efficient server-on-chipabstractWe propose a power-efficient many-core server-on-chip system with 3D-stacked Wide I/O DRAM targeting cloud workloads in datacenters. The integration of 3D-stacked Wide I/O DRAM on top of a logic die increases available memory bandwidth by using dense and fast Through-Silicon Vias (TSVs) instead of off-chip IOs, enabling faster data transfers at much lower energy per bit. We demonstrate a methodology that includes full-system microarchitectural modeling and rapid virtual physical prototyping with emphasis on the thermal analysis. Our findings show that while executing CPU-centric benchmarks (e.g. SPECInt and Dhrystone), the temperature in the server-on-chip (logic+DRAM) is in the range of 175-200°C at a power consumption of less than 20W, exceeding the reliable operating bounds without any cooling solutions, even with embedded cores. However, with real cloud workloads, the power density in the server-on-chip remains much below the temperatures reached by the CPU-centric workloads as a result of much lower power burnt by memory-intensive cloud workloads. We show that such a server-on-chip system is feasible with a low-cost passive heat sink eliminating the need for a high-cost active heat sink with an attached fan, creating an opportunity for overall cost and energy savings in datacenters. Dragomir Milojevic, Sachin Idgunji, Djordje Jevdjic, Emre Ozer 0001, Pejman Lotfi-Kamran, Andreas Panteli, Andreas Prodromou, Chrysostomos Nicopoulos, Damien Hardy, Babak Falsafi, Yiannakis Sazeides |
ICCD | 1 |
| 2011 | An analytical compact model for estimation of stress in multiple Through-Silicon Via configurationsabstractWe present a compact model that provides a quick estimation of the stress and mobility patterns around arbitrary configurations of Through-Silicon Via's (TSVs). No separate TCAD simulations are required for these configurations. It estimates nFET and pFET mobility for industry-standard as well as for (100)/substrate orientations. As the model provides mobility info in less than 0.1 millisecond/transistor/TSV, it is possible to be used in combination with layouting tools and circuit simulators to optimise layouts of circuits for digital and analog applications. The model has been integrated into the 3D PathFinding flow, for steering 3D IO placement during stack definition. Geert Eneman, J. Cho, V. Moroz, Dragomir Milojevic, M. Choi, Kristin De Meyer, Abdelkarim Mercha, Eric Beyne, Thomas Hoffmann 0001, Geert Van der Plas |
DATE | 4 |
| 2011 | Reducing Preemptions and Migrations in Real-Time Multiprocessor Scheduling Algorithms by Releasing the FairnessabstractAbstract-Over the past two decades, numerous optimal scheduling algorithms for real-time systems on multiprocessor platforms have been proposed for the Liu & Layland task model. However, recent studies showed that even if optimal algorithms can theoretically schedule any feasible task set, suboptimal algorithms usually perform better when executed on real computation platforms. This can be explained by the runtime overheads that such optimal algorithms induce. We have observed that all current optimal online multiprocessor real-time scheduling algorithms are (completely or partially) based on the notion of fairness. The respect of this fairness can be the cause of numerous preemptions and migrations. We therefore propose a new algorithm -named U-EDF- which releases the property of fairness and instead use an EDF-like scheduling policy. The simulation results are really encouraging since they show that, in average, U-EDF produces less than one preemption and one migration per job released during the schedule. Furthermore, we strongly believe in the optimality of our algorithm since all tested task sets were correctly scheduled under U-EDF. Geoffrey Nelissen, Vandy Berten, Joël Goossens, Dragomir Milojevic |
RTCSA (1) | 4 |
| 2008 | Power-Aware Real-Time Scheduling upon Dual CPU Type Multiprocessor Platforms
Joël Goossens, Dragomir Milojevic, Vincent Nélis |
OPODIS | 2 |
| 2008 | Concepts and Implementation of Spatial Division Multiplexing for Guaranteed Throughput in Networks-on-ChipabstractTo ensure low power consumption while maintaining flexibility and performance, future Systems-on-Chip (SoC) will combine several types of processor cores and data memory units of widely different sizes. To interconnect the IPs of these heterogeneous platforms, Networks-on-Chip (NoC) have been proposed as an efficient and scalable alternative to shared buses. NoCs can provide throughput and latency guarantees by establishing virtual circuits between source and destination. State-of-the-art NoCs currently exploit Time-Division Multiplexing (TDM) to share network resources among virtual circuits, but this typically results in high network area and energy overhead with long circuit set-up time. We propose an alternative solution based on Spatial Division Multiplexing (SDM). This paper describes our design of an SDM-based network, discusses design alternatives for network implementation and shows why SDM can be better adapted to NoCs than TDM in a specific context. Our case study clearly illustrates the advantages of our technique over TDM in terms of energy consumption, area overhead, and flexibility. A comparison is also performed with a State-of-the-Art industrial reference NoC: Arteris. Anthony Leroy, Dragomir Milojevic, Diederik Verkest, Frédéric Robert, Francky Catthoor |
IEEE Trans. Computers | 2 |
| 2005 | Implementation of Ranking Filters on General Purpose and Reconfigurable Architecture Based on High Density FPGA DevicesabstractIn this paper we present the implementation of ranking filters on two different computer architectures. First we consider a general purpose computer based on Intel Pentium 4 microprocessor and we show that when SSE2 extension is used the throughput vary from 20 to 200 MB/sec, depending on the neighborhood size and rank value. Secondly, we consider a reconfigurable architecture using hundreds of processing elements running bit-serial algorithms and implemented in a high density FPGA device. The throughput of such system varies from 900 to 2600MB/sec which is 13 to 40 times faster than the considered general purpose architecture. Dragomir Milojevic |
FPL | 1 |