Giovanni Mariani

dblp:62/4914 · DBLP profile ↗
← Back
27ranked-venue papers
14as first author
3since 2021 · last 2022
0000-0001-7611-5187ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 14 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Performance modeling and evaluation · 40% Electronic design automation · 28% Memory systems · 15%
Artificial intelligence
2 papers
Efficient and distributed learning · 95% Image recognition and object detection · 5%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.922021
Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces · ICCV 2021
TAPAS: Train-Less Accuracy Predictor for Architecture Search · AAAI 2019
Electronic design automation
design space exploration
0.842018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
DeSpErate++: An Enhanced Design Space Exploration Framework Using Predictive Simulation Scheduling · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
OSCAR: An Optimization Methodology Exploiting Spatial Correlation in Multicore Design Spaces · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
hardware-aware neural architecture search
0.512021
Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces · ICCV 2021
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.512021
Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces · ICCV 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces · ICCV 2021
Performance modeling and evaluation › surrogate modeling
machine-learning-based performance modeling
0.412019
NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning · DAC 2019
Memory systems › processing-in-memory
near-memory processing
0.412019
NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning · DAC 2019
Performance modeling and evaluation
performance prediction
0.412019
NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning · DAC 2019
Processor architecture and microarchitecture
multicore design
0.312018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
Performance modeling and evaluation
processor performance modeling
0.312018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
Compilers and program optimization › autotuning
compiler autotuning
0.212016
COBAYN: Compiler Autotuning Framework Using Bayesian Networks · ACM Trans. Archit. Code Optim. 2016
Computer vision › Image recognition and object detection
image classification
0.112019
TAPAS: Train-Less Accuracy Predictor for Architecture Search · AAAI 2019
Memory systems
processing-in-memory
0.112019
NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning · DAC 2019
Electronic design automation
multi-objective optimization
0.112010
A correlation-based design space exploration methodology for multi-processor systems-on-chip · DAC 2010
Embedded and real-time systems › embedded hardware platform › MPSoC
multiprocessor system-on-chip design
0.112010
A correlation-based design space exploration methodology for multi-processor systems-on-chip · DAC 2010
High-performance computing › supercomputing
exascale systems
0.112018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
Performance modeling and evaluation
workload characterization
0.112016
COBAYN: Compiler Autotuning Framework Using Bayesian Networks · ACM Trans. Archit. Code Optim. 2016
Performance modeling and evaluation › simulation › parallel and distributed simulation
parallel simulation
0.112015
DeSpErate++: An Enhanced Design Space Exploration Framework Using Predictive Simulation Scheduling · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Performance modeling and evaluation
simulation
0.012012
OSCAR: An Optimization Methodology Exploiting Spatial Correlation in Multicore Design Spaces · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012

Methods — techniques the papers use, named apart from their topics

static analysis · 0.5knowledge distillation · 0.5iterative compilation · 0.5evolutionary search · 0.5bayesian network · 0.5accuracy prediction · 0.5ensemble learning · 0.4deep neural network · 0.4architecture search · 0.4hardware performance counters · 0.3analytic modeling · 0.3simulation scheduling · 0.2approximate analytic prediction models · 0.2multi-objective optimization · 0.1design of experiments · 0.1
YearPublicationVenuePosition
2022 RANK - Robotic Ankle: Design and testing on irregular terrains
abstract
Despite the large amount of available exoskeletons, their use in daily life is still limited due to the absence of testing in real-life environments. Thus, the present work aims to test on a series of uneven terrains a wearable ankle exoskeleton, named RANK, designed for walking assistance and drop-foot prevention. RANK consists of a 3D-printed brace attached to the user and a piezoresistive insole, to be incorporated into the user's shoe. Real-time analysis of the insole's sensor outputs enables the control system to provide torque assistance to the ankle joint through a four-bar linkage mechanism. Two healthy male subjects were enrollee, asking them to walk on three different terrain conditions (flat, soft, and irregular) with and without exoskeleton. Human kinematics was gathered via inertial measurements units (IMUs). The effects of ankle exoskeleton on lower limb joint angles were assessed in terms of range of motion (ROM), whereas statistical parametric map method was applied to compare joint angle curves. As expected, a reduction of the ankle ROM approximatively of 10° was found in all terrain conditions between the trails performed with and without exoskeleton. No effects induced on the hip and knee joint were observed. Moreover, no significant differences were found over the almost totality of the gait cycle regardless the terrain conditions. Results demonstrate the capability of the exoskeleton to work properly regardless the type of walking surface.
Juri Taborri, Ilaria Mileti, Giovanni Mariani, Luca Mattioli, Lorenzo Liguori, Stefano Salvatori, Eduardo Palermo, Fabrizio Patane, Stefano Rossi
IROS3
2021 Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces
abstract
Current state-of-the-art Neural Architecture Search (NAS) methods neither efficiently scale to multiple hardware platforms, nor handle diverse architectural search-spaces. To remedy this, we present DONNA (Distilling Optimal Neural Network Architectures), a novel pipeline for rapid, scalable and diverse NAS, that scales to many user scenarios. DONNA consists of three phases. First, an accuracy predictor is built using blockwise knowledge distillation from a reference model. This predictor enables searching across diverse networks with varying macro-architectural parameters such as layer types and attention mechanisms, as well as across micro-architectural parameters such as block repeats and expansion rates. Second, a rapid evolutionary search finds a set of pareto-optimal architectures for any scenario using the accuracy predictor and on-device measurements. Third, optimal models are quickly fine-tuned to training-from-scratch accuracy. DONNA is up to 100× faster than MNasNet in finding state-of-the-art architectures on-device. Classifying ImageNet, DONNA architectures are 20% faster than EfficientNet-B0 and Mo-bileNetV2 on a Nvidia V100 GPU and 10% faster with 0.5% higher accuracy than MobileNetV2-1.4x on a Samsung S20 smartphone. In addition to NAS, DONNA is used for search-space extension and exploration, as well as hardware-aware model compression.
Bert Moons, Parham Noorzad, Andrii Skliar, Giovanni Mariani, Dushyant Mehta, Chris Lott, Tijmen Blankevoort
ICCV4
2021 Efficient image dataset classification difficulty estimation for predicting deep-learning accuracy
abstract
Abstract In the deep-learning community, new algorithms are published at a very fast pace. Therefore, solving an image classification problem for new datasets becomes a challenging task, as it requires to re-evaluate published algorithms and their different configurations in order to find a close to optimal classifier. To facilitate this process, before biasing our decision toward a class of neural networks or running an expensive search over the network space, we propose to estimate the classification difficulty of the dataset. Our method computes a single number that characterizes the dataset difficulty $$97\times $$ 97 × faster than training state-of-the-art networks. The proposed method can be used in combination with network topology and hyper-parameter search optimizers to efficiently drive the search toward promising neural network configurations.
Florian Scheidegger, Roxana Istrate, Giovanni Mariani, Luca Benini, Costas Bekas, Cristiano Malossi
Vis. Comput.3
2020 Evolutionary Algorithm with Non-parametric Surrogate Model for Tensor Program optimization
abstract
The efficiency of tensor operators is key to implement fast deep learning models. However, identifying the fastest implementation of a tensor operator for a target hardware is challenging. A wide range of different configurations have to be considered, and the evaluation of a configuration is time consuming as it requires compilation and execution of the operator. A common approach to address these issues is to boost traditional optimization algorithms with a surrogate modet, i.e., a machine learning model that approximates the objective function and is cheap to query compared to the target hardware. However, as the surrogate model grows in complexity, so does the time needed to train and maintain it. In this work, we propose to use an evolutionary optimizer and augment it with a non-parametric surrogate model (a weighted k-Nearest-Neighbor regression). We evaluate our approach on the convolution layers of a ResNetl8, and show a convergence speedup of up to 1.4×; when compared to baseline operator tuners.
Ioannis Gatopoulos, Romain Lepert, Auke J. Wiggers, Giovanni Mariani, Jakub M. Tomczak
CEC4
2019 TAPAS: Train-Less Accuracy Predictor for Architecture Search
abstract
In recent years an increasing number of researchers and practitioners have been suggesting algorithms for large-scale neural network architecture search: genetic algorithms, reinforcement learning, learning curve extrapolation, and accuracy predictors. None of them, however, demonstrated highperformance without training new experiments in the presence of unseen datasets. We propose a new deep neural network accuracy predictor, that estimates in fractions of a second classification performance for unseen input datasets, without training. In contrast to previously proposed approaches, our prediction is not only calibrated on the topological network information, but also on the characterization of the dataset-difficulty which allows us to re-tune the prediction without any training. Our predictor achieves a performance which exceeds 100 networks per second on a single GPU, thus creating the opportunity to perform large-scale architecture search within a few minutes. We present results of two searches performed in 400 seconds on a single GPU. Our best discovered networks reach 93.67% accuracy for CIFAR-10 and 81.01% for CIFAR-100, verified by training. These networks are performance competitive with other automatically discovered state-of-the-art networks however we only needed a small fraction of the time to solution and computational resources.
Roxana Istrate, Florian Scheidegger, Giovanni Mariani, Dimitrios S. Nikolopoulos, Costas Bekas, Cristiano Malossi
AAAI3
2019 NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning
abstract
The cost of moving data between the memory/storage units and the compute units is a major contributor to the execution time and energy consumption of modern workloads in computing systems. A promising paradigm to alleviate this data movement bottleneck is near-memory computing (NMC), which consists of placing compute units close to the memory/storage units. There is substantial research effort that proposes NMC architectures and identifies workloads that can benefit from NMC. System architects typically use simulation techniques to evaluate the performance and energy consumption of their designs. However, simulation is extremely slow, imposing long times for design space exploration. In order to enable fast early-stage design space exploration of NMC architectures, we need high-level performance and energy models.
Gagandeep Singh 0002, Juan Gómez-Luna, Giovanni Mariani, Geraldo F. Oliveira, Stefano Corda, Sander Stuijk, Onur Mutlu, Henk Corporaal
DAC3
2019 FloatX: A C++ Library for Customized Floating-Point Arithmetic
abstract
We present FloatX (Float eXtended), a C ++ framework to investigate the effect of leveraging customized floating-point formats in numerical applications. FloatX formats are based on binary IEEE 754 with smaller significand and exponent bit counts specified by the user. Among other properties, FloatX facilitates an incremental transformation of the code, relies on hardware-supported floating-point types as back-end to preserve efficiency, and incurs no storage overhead. The article discusses in detail the design principles, programming interface, and datatype casting rules behind FloatX. Furthermore, it demonstrates FloatX’s usage and benefits via several case studies from well-known numerical dense linear algebra libraries, such as BLAS and LAPACK; the Ginkgo library for sparse linear systems; and two neural network applications related with image processing and text recognition.
Goran Flegar, Florian Scheidegger, Vedran Novakovic, Giovanni Mariani, Andrés Tomás, Cristiano Malossi, Enrique S. Quintana-Ortí
ACM Trans. Math. Softw.4
2018 Predicting cloud performance for HPC applications before deployment
Giovanni Mariani, Andreea Anghel, Rik Jongerius, Gero Dittmann
Future Gener. Comput. Syst.1
2018 Analytic Multi-Core Processor Model for Fast Design-Space Exploration
abstract
Simulators help computer architects optimize system designs. The limited performance of simulators even of moderate size and detail makes the approach infeasible for design-space exploration of future exascale systems. Analytic models, in contrast, offer very fast turn-around times. In this paper we propose an analytic multi-core processor-performance model that takes as inputs a) a parametric microarchitecture-independent characterization of the target workload, and b) a hardware configuration of the core and the memory hierarchy. The processor-performance model considers instruction-level parallelism (ILP) per type, models single instruction, multiple data (SIMD) features, and considers cache and memory-bandwidth contention between cores. We validate our model by comparing its performance estimates with measurements from hardware performance counters on Intel Xeon and ARM Cortex-A15 systems. We estimate multi-core contention with a maximum error of 11.4 percent. The average single-thread error increases from 25 percent for a state-of-the-art simulator to 59 percent for our model, but the correlation is still 0.8, a high relative accuracy, while we achieve a speedup of several orders of magnitude. With a much higher capacity than simulators and more reliable insights than back-of-the-envelope calculations it makes automated design-space exploration of exascale systems possible, which we show using a real-world case study from radio astronomy.
Rik Jongerius, Andreea Anghel, Gero Dittmann, Giovanni Mariani, Erik Vermij, Henk Corporaal
IEEE Trans. Computers4
2017 Predicting Cloud Performance for HPC Applications: a User-oriented Approach
abstract
Cloud computing enables end users to execute high-performance computing applications by renting the required computing power. This pay-for-use approach enables small enterprises and startups to run HPC-related businesses with a significant saving in capital investment and a short time to market. When deploying an application in the cloud, the users may a) fail to understand the interactions of the application with the software layers implementing the cloud system, b) be unaware of some hardware details of the cloud system, and c) fail to understand how sharing part of the cloud system with other users might degrade application performance. These misunderstandings may lead the users to select suboptimal cloud configurations in terms of cost or performance. To aid the users in selecting the optimal cloud configuration for their applications, we suggest that the cloud provider generate a prediction model for the provided system. We propose applying machine-learning techniques to generate this prediction model. First, the cloud provider profiles a set of training applications by means of a hardware-independent profiler and then executes these applications on a set of training cloud configurations to collect actual performance values. The prediction model is trained to learn the dependencies of actual performance data on the application profile and cloud configuration parameters. The advantage of using a hardware-independent profiler is that the cloud users and the cloud provider can analyze applications on different machines and interface with the same prediction model. We validate the proposed methodology for a cloud system implemented with OpenStack. We apply the prediction model to the NAS parallel benchmarks. The resulting relative error is below 15% and the Pareto optimal cloud configurations finally found when maximizing application speed and minimizing execution cost on the prediction model are also at most 15% away from the actual optimal solutions.
Giovanni Mariani, Andreea Anghel, Rik Jongerius, Gero Dittmann
CCGrid1
2017 MeSAP: A fast analytic power model for DRAM memories
abstract
The design of an energy-efficient memory subsystem is one of the key issues that system architects face today. To achieve this goal, architects usually rely on system simulators and trace-based DRAM power models. However, their long execution time makes the approach infeasible for the design-space exploration of next-generation exascale computing systems. Analytic models, in contrast, are orders of magnitude faster. In this paper, we propose a new analytic memory-scheduler-agnostic power model for DRAM, henceforth referred to as MeSAP. Similarly to state-of-the-art trace-based approaches, our analytic model achieves an average error of 20%, while being an order of magnitude faster. Furthermore, we integrate MeSAP into an analytic performance model of general-purpose processors and show its applicability to the design of a computing system targeting scientific image processing applications.
Sandeep Poddar, Rik Jongerius, Leandro Fiorin, Giovanni Mariani, Gero Dittmann, Andreea Anghel, Henk Corporaal
DATE4
2017 Classification of thread profiles for scaling application behavior
Giovanni Mariani, Andreea Anghel, Rik Jongerius, Gero Dittmann
Parallel Comput.1
2016 COBAYN: Compiler Autotuning Framework Using Bayesian Networks
abstract
The variety of today’s architectures forces programmers to spend a great deal of time porting and tuning application codes across different platforms. Compilers themselves need additional tuning, which has considerable complexity as the standard optimization levels, usually designed for the average case and the specific target architecture, often fail to bring the best results. This article proposes COBAYN : Compiler autotuning framework using BAYesian Networks, an approach for a compiler autotuning methodology using machine learning to speed up application performance and to reduce the cost of the compiler optimization phases. The proposed framework is based on the application characterization done dynamically by using independent microarchitecture features and Bayesian networks. The article also presents an evaluation based on using static analysis and hybrid feature collection approaches. In addition, the article compares Bayesian networks with respect to several state-of-the-art machine-learning models. Experiments were carried out on an ARM embedded platform and GCC compiler by considering two benchmark suites with 39 applications. The set of compiler configurations, selected by the model (less than 7% of the search space), demonstrated an application performance speedup of up to 4.6 × on Polybench (1.85 × on average) and 3.1 × on cBench (1.54 × on average) with respect to standard optimization levels. Moreover, the comparison of the proposed technique with (i) random iterative compilation, (ii) machine learning--based iterative compilation, and (iii) noniterative predictive modeling techniques shows, on average, 1.2 × , 1.37 × , and 1.48 × speedup, respectively. Finally, the proposed method demonstrates 4 × and 3 × speedup, respectively, on cBench and Polybench in terms of exploration efficiency given the same quality of the solutions generated by the random iterative compilation model.
Amir H. Ashouri, Giovanni Mariani, Gianluca Palermo, Eunjung Park, John Cavazos, Cristina Silvano
ACM Trans. Archit. Code Optim.2
2015 Analytic processor model for fast design-space exploration
abstract
In this paper, we propose an analytic model that takes as inputs a) a parametric microarchitecture-independent characterization of the target workload, and b) a hardware configuration of the core and the memory hierarchy, and returns as output an estimation of processor-core performance. To validate our technique, we compare our performance estimates with measurements on an Intel® Xeon® system. The average error increases from 21% for a state-of-the-art simulator to 25% for our model, but we achieve a speedup of several orders of magnitude. Thus, the model enables fast designspace exploration and represents a first step towards an analytic exascale system model.
Rik Jongerius, Giovanni Mariani, Andreea Anghel, Gero Dittmann, Erik Vermij, Henk Corporaal
ICCD2
2015 DeSpErate++: An Enhanced Design Space Exploration Framework Using Predictive Simulation Scheduling
abstract
Exploring the design space of computer architectures generally consists of a trial-and-error procedure where several architectural configurations are evaluated by using simulation techniques. The final goal of the multiobjective design space exploration (DSE) process is the identification of architectural configurations optimal for a set of target objective functions, typically power consumption, and performance. Simulations are computationally expensive making it rather hard to efficiently explore the design space to identify high-quality configurations in an acceptable exploration time when relying solely on a single-core machine to run simulations. To tackle this problem, engineers proposed solutions based on either: 1) the use of approximate analytic performance models to prune the suboptimal regions of the design space by reducing the number of simulations to run or 2) the use of parallel computing resources to run different simulations concurrently. In this paper we demonstrate that, to efficiently speedup the DSE process while fully exploiting the parallel computing infrastructure, we need to combine the two techniques together in a structured manner. In this paper, we investigate this issue and we propose a DSE solution that exploits approximate analytic prediction models to improve the simulation schedule on a parallel computing environment rather than to prune the number of simulations. Experimental results demonstrate that the proposed technique provides a speedup from 1.26× to 4× with respect to other parallel state-of-the art DSE techniques.
Giovanni Mariani, Gianluca Palermo, Vittorio Zaccaria, Cristina Silvano
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 DRuiD: Designing reconfigurable architectures with decision-making support
abstract
Application development for heterogeneous platforms requires to code and map functionalities on a set of different computing elements. As a consequence, the development process needs a clear understanding of both, application requirements and heterogeneous computing technologies. To support the development process, we propose a framework called DRuiD capable of learning application characteristics that make them suitable for certain computing elements. The framework is composed of an expert system that supports the designer in the mapping decision and gives hints on possible code modifications to be applied to make the functionality more suitable for a computing element. The experimental results are tailored for a heterogeneous and reconfigurable platform (the Xilinx-ml510) including two computational elements, i.e. a Virtex5 FPGA and a PowerPC. The expert system identifies 88.9% of the times what are the functionalities that are accelerated efficiently by using the FPGA, without requiring the kernel porting. Additionally, we present two case studies demonstrating the potentialities of the framework to give hints on high level code modifications for an efficient kernel mapping on the FPGA.
Giovanni Mariani, Gianluca Palermo, Roel Meeuws, Vlad Mihai Sima, Cristina Silvano, Koen Bertels
ASP-DAC1
2014 DeSpErate: Speeding-up design space exploration by using predictive simulation scheduling
abstract
The design space exploration (DSE) phase is used to tune configurable system parameters and it generally consists of a multiobjective optimization (MOO) problem. It is usually done at pre-design phase and consists of the evaluation of large design spaces where each configuration requires long simulation. Several heuristic techniques have been proposed in the past and the recent trend is reducing the exploration time by using analytic prediction models to approximate the system metrics, effectively pruning sub-optimal configurations from the exploration scope. However, there is still a missing path towards the effective usage of the underlying computing resources used by the DSE process. In this work, we will show that an alternative and almost orthogonal approach - focused on exploiting the available parallelism in terms of computing resources - can be used to better schedule the simulations and to obtain a high speedup with respect to state of the art approaches, without compromising the accuracy of exploration results. Experimental results will be presented by dealing with the DSE problem of a shared memory multi-core system considering a variable number of available parallel resources to support the DSE phase1.
Giovanni Mariani, Gianluca Palermo, Vittorio Zaccaria, Cristina Silvano
DATE1
2013 Run-time optimization of a dynamically reconfigurable embedded system through performance prediction
abstract
A key tool to increase the exploitation of dynamic reconfigurable platforms is the run-time resource manager. This system module coordinates the usage of both software and reconfigurable hardware resources in the context of a multi-programmed environment, by alleviating the operating system's induced overhead. This paper introduces a two-layers run-time resource manager for dynamic reconfigurable platforms. The upper level is composed of several application-level managers (one for each application) that provide the most suitable mapping based on resource constraints and performance prediction. The lower level is composed of a centralized system-level resource manager that assigns the HW/SW resources to each application. We present a video surveillance case study in which the proposed resource management technique outperforms the performance of the state of the art by 28% on average, introducing a computational time overhead within 2%.
Giovanni Mariani, Vlad Mihai Sima, Gianluca Palermo, Vittorio Zaccaria, Giacomo Marchiori, Cristina Silvano, Koen Bertels
FPL1
2013 ARTE: An Application-specific Run-Time managEment framework for multi-cores based on queuing models
Giovanni Mariani, Gianluca Palermo, Vittorio Zaccaria, Cristina Silvano
Parallel Comput.1
2013 Design-space exploration and runtime resource management for multicores
abstract
Application-specific multicore architectures are usually designed by using a configurable platform in which a set of parameters can be tuned to find the best trade-off in terms of the selected figures of merit (such as energy, delay, and area). This multi-objective optimization phase is called Design-Space Exploration (DSE). Among the design-time (hardware) configurable parameters we can find the memory subsystem configuration (such as cache size and associativity) and other architectural parameters such as the instruction-level parallelism of the system processors. Among the runtime (software) configurable parameters we can find the degree of task-level parallelism associated with each application running on the platform. The contribution of this article is twofold; first, we introduce an evolutionary (NSGA-II-based) methodology for identifying a hardware configuration which is robust with respect to applications and corresponding datasets. Second, we introduce a novel runtime heuristic that exploits design-time identified operating points to provide guaranteed throughput to each application. Experimental results show that the design-time/runtime combined approach improves the runtime performance of the system with respect to existing reference techniques, while meeting the overall power budget.
Giovanni Mariani, Gianluca Palermo, Vittorio Zaccaria, Cristina Silvano
ACM Trans. Embed. Comput. Syst.1
2012 Using multi-objective design space exploration to enable run-time resource management for reconfigurable architectures
abstract
Resource run-time managers have been shown particularly effective for coordinating the usage of the hardware resources by multiple applications, eliminating the necessity of a full-blown operating system. For this reason, we expect that this technology will be increasingly adopted in emerging multi-application reconfigurable systems. This paper introduces a fully automated design flow that exploits multi-objective design space exploration to enable runtime resource management for the Molen reconflgurable architecture. The entry point of the design flow is the application source code; our flow is able to heuristically determine a set of candidate hardware/software configurations of the application (i.e., operating points) that trade off the occupation of the reconflgurable fabric (in this case, an FPGA), the load of the master processor and the performance of the application itself. This information enables a run-time manager to exploit more efficiently the available system resources in the context of multiple applications. We present the results of an experimental campaign where we applied the proposed design flow to two reference audio applications mapped on the Molen architecture. The analysis proved that the overhead of the design space exploration and operating points extraction with respect to the original Molen flow is within reasonable bounds since the final synthesis time still represents the major contribution. Besides, we have found that there is a high variance in terms of execution time speedup associated with the operating points of the application (characterized by a different usage of the FPGA) which can be exploited by the run-time manager to increase/decrease the quality of service of the application depending on the available resources1.
Giovanni Mariani, Vlad Mihai Sima, Gianluca Palermo, Vittorio Zaccaria, Cristina Silvano, Koen Bertels
DATE1
2012 OSCAR: An Optimization Methodology Exploiting Spatial Correlation in Multicore Design Spaces
abstract
This paper presents OSCAR, an optimization methodology exploiting spatial correlation of multicore design spaces. This paper builds upon the observation that power consumption and performance metrics of spatially close design configurations (or points) are statistically correlated. We propose to exploit the correlation by using a response surface model (RSM), i.e., a closed-form expression suitable for predicting the quality of nonsimulated design points. This model is useful during the design space exploration (DSE) phase to quickly converge to the Pareto set of the multiobjective problem without executing lengthy simulations. To this end, we introduce a multiobjective optimization heuristic which iteratively updates and queries the RSM to identify the design points with the highest expected improvement. The RSM allows to consolidate the Pareto set by reducing the number of simulations required, thus speeding up the exploration process. We compare the proposed heuristic with state-of-the-art approaches [conventional, RSM-based, and structured design of experiments (DoEs)]. Experimental results show that OSCAR is a faster heuristic with respect to state-of-the-art techniques such as response-surface Pareto iterative refinement ReSPIR and nondominated-sorting genetic algorithm NSGA-II. In fact, OSCAR used a lower number of simulations to produce a similar solution, i.e., an average of 150 simulations instead of 320 simulations (NSGA-II) and 178 simulations (ReSPIR). When the number of design points is fixed to an average of 300, OSCAR achieves less than 0.6% in terms of average distance with respect to the reference solution while NSGA-II achieves 3.4%. Reported results also show that OSCAR can significantly improve structured DoE approaches by slightly increasing the number of experiments.
Giovanni Mariani, Gianluca Palermo, Vittorio Zaccaria, Cristina Silvano
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2010 A correlation-based design space exploration methodology for multi-processor systems-on-chip
abstract
Given the increasing complexity of multi-processor systems-on-chip, a wide range of parameters must be tuned to find the best trade-offs in terms of the selected system figures of merit (such as energy, delay and area). This optimization phase is called Design Space Exploration (DSE) consisting of a Multi-Objective Optimization (MOO) problem. In this paper, we propose an iterative design space exploration methodology exploiting the statistical properties of known system configurations to infer, by means of a correlation-based analysis, the next design points to be analyzed with low-level simulations. In fact, the knowledge of few design points is used to predict the expected improvement of unknown configurations. We show that the correlation of the configurations within the multi-processor design space can be modeled successfully with analytical functions and, thus, speed up the overall exploration phase. This makes the proposed methodology a model-assisted heuristic that, for the first time, exploits the correlation about architectural configurations to converge to the solution of the multi-objective problem.
Giovanni Mariani, Aleksandar Brankovic, Gianluca Palermo, Jovana Jovic, Vittorio Zaccaria, Cristina Silvano
DAC1
2010 An industrial design space exploration framework for supporting run-time resource management on multi-core systems
abstract
Current multi-core design methodologies are facing increasing unpredictability in terms of quality due to the actual diversity of the workloads that characterize the deployment scenario. To this end, these systems expose a set of dynamic parameters which can be tuned at run-time to achieve a specified Quality of Service (QoS) in terms of performance. A run-time manager operating system module is in charge of matching the specified QoS with the available platform resources by manipulating the overall degree of task-level parallelism of each application as well as the frequency of operation of each of the system cores. In this paper, we introduce a design space exploration framework for enabling and supporting enhanced resource management through software re-configuration on an industrial multicore platform. From one side, the framework operates at design time to identify a set of promising operating points which represent the optimal trade-off in terms of the target power consumption and performance. The operating points are used after the system has been deployed to support an enhanced resource management policy. This is done by a light-weight resource management layer which filters and selects the optimal parallelism of each application and operating frequency of each core to achieve the QoS constraints imposed by the external world and/or the user. We show how the proposed design-time and run-time techniques can be used to optimally manage the resources of a multiple-stream MPEG4 encoding chip dedicated to automotive cognitive safety tasks.
Giovanni Mariani, Prabhat Avasare, Geert Vanmeerbeeck, Chantal Ykman-Couvreur, Gianluca Palermo, Cristina Silvano, Vittorio Zaccaria
DATE1
2009 Meta-model Assisted Optimization for Design Space Exploration of Multi-Processor Systems-on-Chip
abstract
Multi-processor Systems-on-chip are currently designed by using platform-based synthesis techniques. In this approach, a wide range of platform parameters are tuned to find the best trade-offs in terms of the selected system figures of merit (such as energy, delay and area). This optimization phase is called Design Space Exploration (DSE) and it generally consists of a Multi-Objective Optimization (MOO) problem. The design space of a Multi-processor architecture is too large to be evaluated comprehensively. So far, several heuristic techniques have been proposed to address the MOO problem, but they are characterized by low efficiency to identify the Pareto front. In this paper, we address the MPSoC DSE problem by using an NSGA-II modified to be assisted by an Artificial Neural Network (ANN). In particular we exploit statistical methods to compute the prediction confidence intervals for the ANN approximations. These information are adopted in the evolution control strategy in order to carefully select which individuals should be simulated. Experimental results show that the proposed techniques is able to reduce the simulations needed for the optimization without decreasing the quality of the obtained Pareto Front. Results are compared with state of the art techniques to demonstrate that optimization time due to simulation can be speed up by adopting statistical methods during evolution control.
Giovanni Mariani, Gianluca Palermo, Cristina Silvano, Vittorio Zaccaria
DSD1
2007 Mapping and Topology Customization Approaches for Application-Specific STNoC Designs
abstract
Application-specific network-oriented communication architectures have recently become an effective solution to support high bandwidth Systems on-Chip. The Network on-Chip architectures considered so far range from regular to fully customized topologies for application-specific designs requiring high-level bandwidth. To this end, a network-centric design flow is necessary to support the design space exploration of complex SoCs with tight design constraints. This paper introduces four different approaches based on the orthogonalization of core mapping and topology customization applied to STNoC, the Network on-Chip developed by STMicroelecronics. The four methods are derived from the combination of the initial mappings to two standard topologies (ring and spidergon) with two types of topology customization based on the insertion of cross-links to reduce the network distance of standard topologies.
Gianluca Palermo, Giovanni Mariani, Cristina Silvano, Riccardo Locatelli, Marcello Coppola
ASAP2
2007 Application-Specific Topology Design Customization for STNoC
abstract
Customized network-oriented communication architectures have recently become a must to support high bandwidth SoCs. To this end, a corresponding communication design flow is necessary to support the design space exploration of complex SoCs with tight design constraints. In order to exploit the benefits introduced by the NoC approach for the on-chip communication, the paper presents a Pareto Simulated Annealing (PSA) approach for the customization of the network topology. The proposed PSA approach has been applied to STNoC, the Network on-Chip developed by STMicroelectronics. Starting from the ring topology, the proposed application-specific design flow tries to find a set of customized topologies (optimized in terms of performance and area/energy overhead) by adding custom links up to the spidergon topology.
Gianluca Palermo, Cristina Silvano, Giovanni Mariani, Riccardo Locatelli, Marcello Coppola
DSD3