Sandro Bartolini

dblp:10/679 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
5since 2021 · last 2024
0000-0002-7975-3632ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2024 DeVAS: Decoupled Virtual Address Spaces
abstract
The constant growth of workload size in modern applications is making address translation a performance bottleneck. In principle, increasing the virtual page size could be advantageous, as it would allow each cached address translation to cover a larger memory space. Nevertheless, the utilization of larger pages introduces challenges, such as issues related to memory fragmentation and physical page management. In this paper, we present Decoupled Virtual Address Spaces (DeVAS), a virtual memory proposal that enables the decoupling of address translation and memory allocation, by allowing different sizes for virtual and physical pages, aiming to exploit the benefits of both. DeVAS introduces an intermediate virtual address space allocated by the Operating System employing larger pages (e.g., 2MiB), corresponding to the virtual page size seen by the processor, and a memory controller extension devoted to their allocation in physical memory at a smaller granularity (e.g., 4KiB). We show that DeVAS achieves an average 1.13 × performance improvement over a traditional configuration (i.e., with 4KiB pages) with no specific architectural modifications and for memory-intensive benchmarks. Moreover, DeVAS strategy enables architectural modifications for increasing performance, such as simplified/optimized TLB structure and L1-cache design flexibility. When considering these adjustments, DeVAS achieves a speedup of up to 1.20 × compared to the same reference. Furthermore, it matches and even edges the performance of an ideal (i.e., not implementable in practice) virtual memory configuration based on 2MiB pages only.
Mirco Mannino, Biagio Peccerillo, Andrea Mondelli, Sandro Bartolini
SBAC-PAD4
2023 Energy and Performance Improvements for Convolutional Accelerators Using Lightweight Address Translation Support
abstract
The growing demand for deep learning applications has led to the design and development of several hardware accelerators to increase performance and energy efficiency. In particular, convolutional accelerators are among those receiving the most attention due to their applicability in many fields. Another aspect that is gaining increasing attention is the use of a shared virtual address space between processor and accelerators. It can provide several advantages such as programmability and security. The use of a shared address space relies on a time-consuming IOMMU to satisfy address translation requests. In this work, we analyze convolutional workloads in convolutional accelerators, identifying the sensitivity of performance to IOMMU activity. Additionally, based on the analysis done on convolutional workloads, we propose the use of dedicated accelerator registers (Translation Registers) to reduce costly IOMMU accesses. Translation Registers allow reducing execution time by about 20% and the energy consumption related to address translation up to about 55%.
Mirco Mannino, Biagio Peccerillo, Andrea Mondelli, Sandro Bartolini
CF4
2022 Applying Intel's oneAPI to a machine learning case study
abstract
Abstract Different technologies and approaches exist to work around the performance portability problem. Companies and academia work together to find a way to preserve performance across heterogeneous hardware using a unified language, one language to rule them all. Intel's oneAPI appears with this idea in mind. In this article, we try the new Intel solution to approach heterogeneous programming, choosing machine learning as our case study. More precisely, we choose Caffe, a machine learning framework that was created six years ago. Nevertheless, how would it be to make Caffe again from the beginning, using a fresh new technology like oneAPI? In terms of not only the ease of programming‐because only one source code would be needed to deploy Caffe to CPUs, GPUs, FPGAs, and accelerators (platforms that oneAPI currently supports)‐but also performance, where oneAPI may be capable of taking advantage of specific hardware automatically. Is Intel's oneAPI ready to take the leap?
Pablo Antonio Martínez, Biagio Peccerillo, Sandro Bartolini, José M. García 0001, Gregorio Bernabé
Concurr. Comput. Pract. Exp.3
2022 Flexible task-DAG management in PHAST library: Data-parallel tasks and orchestration support for heterogeneous systems
abstract
Summary Heterogeneous architectures proved successful in achieving unprecedented performance and energy‐efficiency. However, taking advantage of these diverse processing elements is still hard. Programmers need to code through the different approaches suitable for each target architecture and need to decide the distribution of activities on the different resources. The majority of current frameworks focuses on either performance or productivity. The former mainly provides low‐level target‐specific programming interfaces, and the latter offers high‐level tools that often fail in achieving high‐performance. In both cases, the design is usually data‐parallel, as task‐parallelism is not supported. In this work, we propose a task‐based solution within the data‐parallel heterogeneous single‐source PHAST library. Tasks can be coded in a target‐agnostic fashion, can be compiled and parallelized on multi‐core CPUs and NVIDIA GPUs automatically and support the choice of the execution platform at runtime. We evaluate the capabilities of the proposed task‐directed acyclic graph support in case of an extensive set of randomly generated task‐based applications with different sizes and characteristics. We compare it against a SYCL implementation in terms of performance and complexity metrics, highlighting that PHAST achieves about 1.56× and 2.60× speedup over SYCL for multi‐core CPU and GPU, respectively, while improving also code complexity metrics.
Biagio Peccerillo, Sandro Bartolini
Concurr. Comput. Pract. Exp.2
2022 A survey on hardware accelerators: Taxonomy, trends, challenges, and perspectives
abstract
In recent years, the limits of the multicore approach emerged in the so-called “dark silicon” issue and diminishing returns of an ever-increasing core count. Hardware manufacturers, out of necessity, switched their focus to accelerators, a new paradigm that pursues specialization and heterogeneity over generality and homogeneity. They are special-purpose hardware structures separated from the CPU with aspects that exhibit a high degree of variability. We define a taxonomy based on fourteen of these aspects, grouped in four macro-categories: general aspects, host coupling, architecture, and software aspects. According to it, we categorize around 100 accelerators of the last decade from both industry and academia, and critically analyze emerging trends. We complete our discussion with throughput and efficiency figures. Then, we discuss some prominent open challenges that accelerators are facing, analyzing state-of-the-art solutions, and suggesting prospective research directions for the future.
Biagio Peccerillo, Mirco Mannino, Andrea Mondelli, Sandro Bartolini
J. Syst. Archit.4
2019 PHAST - A Portable High-Level Modern C++ Programming Library for GPUs and Multi-Cores
abstract
A decade after the beginning of the many-core era, multi-core CPU and GPU architectures are everywhere, from mobile devices up to high-performance workstations and servers. To this day, programmers willing to harness their power need to express their code via languages and frameworks that often lack of expressivity and high-level abstractions. These solutions, despite allowing users to reach unprecedented performance, can still be a hampering factor for productivity and portability. In this paper we propose PHAST, a modern C++, STL-like, single-source programming library and approach based on multi-dimensional dynamic containers and multi-layered functors that can be targeted on NVIDIA GPUs and multi-core CPUs. Its main purpose is to let programmers write code once for different architectures at a high level of abstraction, to reach high-performance while allowing fine parameter tuning and not shielding code from low-level target-specific optimizations. To assess the value of our proposal, we consider benchmarks from different application domains, and we evaluate their PHAST implementations against CUDA, OpenCL, Kokkos, and SYCL ones from both performance and productivity points of view. We show that PHAST can significantly reduce code complexity metrics while reaching very good performance.
Biagio Peccerillo, Sandro Bartolini
IEEE Trans. Parallel Distributed Syst.2
2018 Exploring the relationship between architectures and management policies in the design of NUCA-based chip multicore systems
Sandro Bartolini, Pierfrancesco Foglia, Cosimo Antonio Prete
Future Gener. Comput. Syst.1
2018 Scalable Path-Setup Scheme for All-Optical Dynamic Circuit Switched NoCs in Cache Coherent CMPs
abstract
Nanophotonics is a promising solution for on-chip interconnection due to its intrinsic low-latency and low-power features, which can be useful for performance and energy in future Chip Multi-Processors (CMPs). This article proposes a novel arbitrated all-optical path-setup scheme for tiled CMPs adopting circuit-switched optical networks. It aims at significantly reducing path-setup latency and overall energy consumption. The proposed arbitrated scheme is able to configure multiple photonic switches simultaneously, instead of sequentially as it is done in state-of-the-art proposals. The proposed fast optical path-setup solution reduces the overhead in each transmission and, most importantly, allows optical circuit-switched networks to effectively serve cache coherence traffic, which is mainly composed of relatively small messages. Specifically, we propose a single-arbiter scheme where the whole topology is managed by a central module (single-arbiter) that takes care of the path-setup procedures. Then, to tackle scalability, we propose a logically clustered architecture (multi-arbiter) in which an arbiter is allocated in each logical core-cluster and an ad hoc distributed reservation protocol coordinates arbiters to manage inter-cluster path reservations. We show that our proposed single-arbiter architecture outperforms a state-of-the-art optical network with sequential path-setup (optical baseline) in the case of 8- and 16-core tiled CMP setups. However, due to serialization issues, the single-arbiter solution is not able to compete with a reference electronic baseline for bigger 32- and 64-core setups even if still performing much better than the optical baseline. Conversely, our multi-arbiter hierarchical solution allows us to improve performance up to almost 20% and 40% for 32- and 64-core setups, respectively, demonstrating a wide applicability of the proposed technique. Energy-wise, the analyzed solutions enable significant savings compared to both the optical baseline with sequential path setup, and to the electronic counterpart. Specifically, results show more than 25% average improvement for the single-arbiter in the 8- and 16-core cases, and more than 40% and 15% savings for the multi-arbiter in the 32- and 64-core cases, respectively.
Paolo Grani, Sandro Bartolini
ACM J. Emerg. Technol. Comput. Syst.2
2014 Assessing the energy break-even point between an optical NoC architecture and an aggressive electronic baseline
abstract
Many crossbenchmarking results reported in the open literature raise optimistic expectations on the use of optical networks-on-chip (ONoCs) for high-performance and low-power on-chip communication. However, most of those previous works ultimately fail to make a compelling case for chip-level nanophotonic NoCs, especially for the lack of aggressive electronic baselines (ENoC), and the poor accuracy in physical- and architecture-layer analysis of the ONoC. This paper aims at providing the guidelines and minimum requirements so that nanophotonic emerging technology may become of practical relevance. The key differentiating factor of this work consists of contrasting ONoC solutions with an aggressive ENoC architecture with realistic complexity, performance, and power figures, synthesized on an industrial 40nm low-power technology. At the same time, key physical design issues and network interface architecture requirements for the ONoC under test are carefully assessed, thus paving the way for a well-grounded definition of the requirements for the emerging ONoC technology to achieve the energy break-even point with respect to pure electronic interconnect solutions in future multi- and many-core systems.
Luca Ramini, Alberto Ghiribaldi, Paolo Grani, Sandro Bartolini, Hervé Tatenguem, Davide Bertozzi
DATE4
2014 Solving Graph Partitioning Problems Arising in Tagless Cache Management
Sandro Bartolini, Iacopo Casini, Paolo Detti
ISCO1
2014 Towards compelling cases for the viability of silicon-nanophotonic technology in future manycore systems
abstract
Many crossbenchmarking results reported in the open literature provide optimistic expectations on the use of optical networks-on-chip (ONoCs) for high-performance and low-power on-chip communication in future manycore systems. The goal of this paper is to highlight key methodological steps for a realistic assessment of the emerging nanophotonic technology. Building on this methodology, the paper provides an accurate energy efficiency comparison between an ONoC and an ENoC counterpart both at the level of the system interconnect and of the system as a whole. As a result, the paper points out the most promising directions for the development of the technology for the sake of practical relevance, and confirms that the technology has potential based on a characterization methodology with uncommon cross-layer visibility.
Luca Ramini, Hervé Tatenguem, Alberto Ghiribaldi, Paolo Grani, Marta Ortín-Obón, Anja Boos, Sandro Bartolini
NOCS7
2014 Exploiting silicon photonics for energy-efficient heterogeneous parallel architectures
abstract
Welcome to this special issue of the journal Concurrency and Computation: Practice and Experience on Exploiting Silicon Photonics for Energy-Efficient Heterogeneous Parallel Architectures, which contains five original manuscripts that cover a complete range of perspectives. Silicon photonics is undoubtedly expected to play a big role in the evolution in board, cross-chip, interposer-level and on-chip interconnection for low-power and/or high-performance computer systems spanning from high-end embedded devices (e.g., tablets and smartphones) and other System-on-Chips (SoCs), up to chips for the High Performance Computing (HPC) domain. The unique features of photonics (e.g., extreme low-latency, end-to-end transmission, high bandwidth density and passive long-range propagation) have the potential constitute a discontinuity element able to modify the expected shape of future computer systems from the design point of view and also from the programmability and/or runtime management perspectives. Summarizing, silicon photonics can bring innovations and benefits into current and foreseeable computing systems directly, due to their intrinsic features, but also indirectly enabling the evolution toward architectures, runtime and resource management approaches that maximize the photonic raw technological opportunities and lead to more efficient overall designs, otherwise impossible. For instance, the extreme low transmission latency (i.e., group velocity of light into silicon, about 15 ps/mm) can potentially allow a higher number of architectural modules to be close each other and thus to enable their effective tight cooperation and communication. However, computer architecture, as well as network on- and off-chip, designs needs to be adapted to extract maximum benefits from the photonic technology, which exposes other substantial differences compared to what designers are well accustomed to. For example, at the moment optical interconnection is end-to-end by nature therefore much of the knowledge and solutions based on store-and-forward paradigm cannot be directly transferred and exploited. However, propagation into a silicon waveguide can occur with limited losses (e.g., even less than 1 dB/cm) over on-chip or interposer distances without signal regeneration needs. In brief, in this arena, new ad-hoc solutions need to be pursued. Then, despite optical communication is very well established and all the involved elements (such as modulators, detectors, waveguides and resonators) have been extensively researched on, silicon photonics applied to computing systems is still in its infancy. Consequently, researchers have already highlighted a deep interaction between design choices at very different layers of abstraction. For this reason, nowadays the whole spectrum of layers, from physical concerns about optical structures on silicon (e.g., module layout and modeling to expose interactions and, for instance, to evaluate and limit insertion losses) up to network issues (e.g., connectivity and topologies, bandwidth and latency) and even computer architecture choices (e.g., memory coherency and consistency models, memory hierarchy and parallelism management), need to be studied with a strong multidisciplinary approach. This special issue contributes to this promising field with extended and carefully reviewed versions of selected papers from the First International Workshop on Exploiting Silicon Photonics for Energy-Efficient Heterogeneous Parallel Architectures (SiPhotonics'14), which was held in Vienna (Austria) as part of the 9th HiPEAC conference on High Performance and Embedded Architecture and Compilers. Therefore, due to the peculiarity of the depicted scenario, the papers of this number address a complete range of perspectives to silicon photonics applied to computing systems, from raw technology issues and solutions up to studies at the overall system level of modern multi-/many-core systems, both from academic and industrial researchers working in this area. We start this special issue with the paper entitled Optical Crossbars on Chip, A Comparative Study based on Worst-Case Losses. In this paper, Le Beux et al., 1 study the worst-case losses for possible crossbar implementations depending on three key design factors: network topology, considered layout and insertion losses induced by the fabrication process. They compare different implementations relying on matrix, multistage and ring-based network topologies, finding that ring-based networks yield the most power-efficient solution. The paper Capturing the Sensitivity of Optical Network Quality Metrics to its Network Interface Parameters by Ortin et al., 2 addresses the network interface architecture (NI) required to support optical communications on the silicon chip. The paper proposes a complete network interface architecture for wavelength-routed optical NoCs, by coping with the intricacy of some specific issues such as flow control, buffering strategy and dual-clock domains, deadlock avoidance, serialization, and above all, the co-design around the requirements of a cache coherency protocol. The most important conclusion is that NI design and optimization perhaps has now higher priority over the relentless search for improvements in individual optical devices. The emerging of circuit-level simulators for photonic integrated circuits (PICs) is driven by recent developments in technologies for integration of large-scale monolithic PICs in both, silicon and InP technologies. Arellano et al., 3 present their solution for modeling PICs in the framework of the circuit-level simulation tool VPIcomponentMakerTMPhotonic Circuits. In their paper The Power of Circuit Simulations for Designing Photonic Integrated Circuits, they demonstrate the combination of different simulation approaches in time domain, frequency domain and time-and-frequency domain (TFDM) for fast and accurate simulations. This is particularly crucial for being able to model and design more and more complex optical circuits like the ones that could be needed to be employed in, and/or between, chip multiprocessors. In the paper entitled Managing Resources Dynamically in Hybrid Photonic-Electronic NoCs, García-Guirado et al., 4 present novel fine-grain policies to manage the photonic resources in a tiled-CMP scenario. The objective is to maintain the optical channel in the load condition that allows it to deliver best performance. Their policies are dynamic and base their decisions on parameters such as message size, ring availability and distance between endpoints, at the message level. The resulting network behavior is also fairer to all cores, reducing processor idle time thanks to faster thread synchronization, improving performance and reducing both the overall network latency and energy consumption when compared to the same CMP without the photonic ring. Finally, the paper Towards Zero Latency Photonic Switching in Shared Memory Networks explores techniques which intelligently use information from the memory hierarchy to predict communication in order to setup photonic circuits with reduced or eliminated arbitration latency in case of reconcilable optical networks. Madarbux et al., 5 present a switch scheduling algorithm which arbitrates on a per memory transaction basis and holds open photonic circuits to exploit temporal locality, showing that this can reduce the average arbitration latency overhead and eliminate arbitration latency altogether for many of memory transactions. We would like to thank all the authors, reviewers and editors involved in the elaboration of this special issue, including also the reviewers that were involved in the SiPhotonics'14 workshop, where short versions of the papers were previously selected. We are especially grateful to Profs. Geoffrey C. Fox and David W. Walker, editors of the journal, for approving this special issue and for his help along the process of its preparation.
Sandro Bartolini, José M. García 0001
Concurr. Comput. Pract. Exp.1
2014 Managing resources dynamically in hybrid photonic-electronic networks-on-chip
abstract
SUMMARY Nanophotonics promises to solve the scalability problems of current electrical interconnects thanks to its low sensitivity to distance in terms of latency and energy consumption. Before this technology reaches maturity, hybrid photonic‐electronic networks will be a viable alternative. Ideally, ordinary electrical meshes and ring‐based photonic networks should cooperate to minimize overall latency and energy consumption, but currently, we lack mechanisms to do this efficiently. In this paper, we present novel fine‐grain policies to manage the photonic resources in a tiled chip multiprocessor (CMP) scenario. Our policies are dynamic and base their decisions on parameters such as message size, ring availability, and distance between endpoints, at the message level. The resulting network behavior is also fairer to all cores, reducing processor idle time thanks to faster thread synchronization. All these policies improve performance when compared to the same CMP without the photonic ring, and the most elaborate ones reduce the overall network latency by 50%, execution time by 36%, and network energy consumption by 52% on average, in a 16‐core CMP for the PARSEC benchmark suite. Larger hybrid networks with 64 endpoints for 256‐core CMPs, based on Corona and Firefly designs, also show far superior throughput and lower latency if managed by one of the proposed policies. Copyright © 2014 John Wiley & Sons, Ltd.
Antonio García-Guirado, Ricardo Fernández-Pascual, José M. García 0001, Sandro Bartolini
Concurr. Comput. Pract. Exp.4
2014 Design Options for Optical Ring Interconnect in Future Client Devices
abstract
Nanophotonic is a promising solution for on-chip interconnection due to its intrinsic low-latency and low-power features. Future tiled chip multiprocessors (CMPs) for rich client devices can receive energy benefits from this technology but we show that great care has to be put in the integration of the various involved facets to avoid queuing and serialization issues and obtain the rated potential advantages. We evaluate different management strategies for accessing a simple, shared photonic path (ring), working in conjunctions with a standard electronic mesh or alone, in a tiled CMP. Our results highlight that a careful selection of the most latency-critical messages to be routed in photonics and the use of a conflict-free access scheme is crucial for obtaining performance/power advantages when the available bandwidth is limited. We identify the design point where all the traffic can be routed on the photonic path and thus the electronic network can be suppressed. At this point, the ring achieves 20--25% speedup and 84% energy consumption improvement over the electronic baseline. Then we investigate the same trade-offs when the number of rings is increased up to eight, allowing to raise performance benefits up to 40% or reaching up to 80% energy reduction. We finally explore the effects of deploying a given optical parallelism split between a higher number of waveguides for further improving energy savings.
Paolo Grani, Sandro Bartolini
ACM J. Emerg. Technol. Comput. Syst.2
2013 Contrasting wavelength-routed optical NoC topologies for power-efficient 3D-stacked multicore processors using physical-layer analysis
abstract
Optical networks-on-chip (ONoCs) are currently still in the concept stage, and would benefit from explorative studies capable of bridging the gap between abstract analysis frameworks and the constraints and challenges posed by the physical layer. This paper aims to go beyond the traditional comparison of wavelength-routed ONoC topologies based only on their abstract properties, and for the first time assesses their physical implementation efficiency in an homogeneous experimental setting of practical relevance. As a result, the paper can demonstrate the significant and different deviation of topology layouts from their logic schemes under the effect of placement constraints on the target system. This becomes then the preliminary step for the accurate characterization of technology-specific metrics such as the insertion loss critical path, and to derive the ultimate impact on power efficiency and feasibility of each design.
Luca Ramini, Paolo Grani, Sandro Bartolini, Davide Bertozzi
DATE3
2013 Olympic: A Hierarchical All-Optical Photonic Network for Low-Power Chip Multiprocessors
abstract
The continuous increase of the number of cores in tiled chip-multi-processors (CMP) will prevent traditional electronic networks on chip (NoC) to maintain an acceptable tradeoff between performance and power consumption. Recent advances in silicon-photonics open new opportunities for fast and low-energy on-chip interconnections but specific design and tuning is needed. This paper proposes Olympic, an all-optical NoC architecture using a hierarchical topology made up of replicated and cascaded simple photonic building blocks (rings). Local rings connect tiles within clusters directly and a global ring glues together local ones and enables inter-cluster communications. The all-optical approach allows to achieve a low-energy solution, very important for future embedded CMPs. The cost of these benefits resides mainly in the additional optical-electronic-optical conversions needed for inter-cluster transmissions and in this paper we single out promising design tradeoffs using the PARSEC benchmark suite. We show that a careful design of our Olympic clustered architecture can achieve 65% energy reduction with only 1% slowdown compared to a full 2D mesh, for a 16-core CMP.
Sandro Bartolini, Luca Lusnig, Enrico Martinelli
DSD1
2012 A Simple On-Chip Optical Interconnection for Improving Performance of Coherency Traffic in CMPs
abstract
Nanophotonic interconnection is a promising solution for inter-core communication in future chip multiprocessors (CMPs). Main benefits derive from its intrinsic low-latency and high-bandwidth, especially when employing wavelength division multiplexing (WDM), as well as reduced power requirements when compared to electronic NoCs. Existing works on optical NoCs (ONoC) mainly concentrate on relatively complex proposals needed to host the whole CMP traffic. In some proposals complexity is increased also from the need of an electronic network for preliminary pathsetup in the optical one. This paper proposes to enhance a conventional NoC with only a simple photonic structure, a ring, and aims at investigating its suitability to support the low-latency transmission of small latency-critical coherency control messages as to improve performance of multithreaded applications. In particular, our proposed scheme supports fast multicast transmission of invalidation messages. We have simulated Parsec benchmarks on an 8 core full-system CMP. Results show that a careful selection of coherency control messages to be forwarded to the photonic ring allows improving execution time up to 19%, with an average of 6% across all considered benchmarks. We discuss how different selections of messages, i.e. related to read and/or write operations, affect results and single out the most profitable set. Moreover, we show that the sharing behavior of benchmarks has a central role in the final performance.
Sandro Bartolini, Paolo Grani
DSD1
2011 Link-time optimization for power efficiency in a tagless instruction cache
abstract
The instruction cache is a critical component in any microprocessor. It must have high performance to enable fetching of instructions on every cycle. However, current designs waste a large amount of energy on each access as tags and data banks from all cache ways are consulted in parallel to fetch the correct instructions as quickly as possible. Existing approaches to reduce this overhead remove unnecessary accesses to the data banks or to the ways that are not likely to hit. However, tag hunks still need to be checked. This paper considers a new hybrid hardware and linker-assisted approach to tagless instruction caching. Our novel cache architecture, supported by the compilation toolchain, removes the need for tag checks entirely for the majority of cache accesses. The linker places frequently-executed instructions in specific program regions that are then mapped into the cache without the need for tag checks. This requires minor hardware modifications, no ISA changes and works across cache configurations. Our approach keeps the software and hardware independent, resulting in both backward and forward compatibility. evaluation on a superscalar processor with and without SMI' support shows power savings of 66% within the instruction cache with no loss of performance. This translates to a 49% saving when considering the combined power of the instruction cache and translation lookaside buffer, which is involved in managing our tagless scheme.
Timothy M. Jones 0001, Sandro Bartolini, Jonas Maebe, Dominique Chanet
CGO2
2010 Feedback-Driven Restructuring of Multi-threaded Applications for NUCA Cache Performance in CMPs
abstract
This paper addresses feedback-directed restructuring techniques tuned to Non Uniform Cache Architectures (NUCA) in CMPs running multi-threaded applications. Access time to NUCA caches depends on the location of the referred block, so the locality and cache mapping of the application influence the overall performance. We show techniques for altering the distribution of applications into the cache space as to achieve improved average memory access time. In CMPs running multi-threaded applications, the aggregated accesses (and locality) of the processors form the actual cache load and pose specific issues. We consider a number of Splash-2 and Parsec benchmarks on an 8 processor system and we show that a relatively simple remapping algorithm is able to improve the average Static-NUCA (SNUCA) cache access time by 5.5% and allows an SNUCA cache to surpass the performance of a more complex dynamic-NUCA (DNUCA) for most benchmarks. Then, we present a more sophisticated remapping algorithm, relying on cache geometry information and on the access distribution statistics from individual processors, that reduces the average cache access time by 10.2% and is very stable across all benchmarks.
Sandro Bartolini, Pierfrancesco Foglia, Marco Solinas, Cosimo Antonio Prete
SBAC-PAD1
2008 Instruction Cache Energy Saving Through Compiler Way-Placement
abstract
Fetching instructions from a set-associative cache in an embedded processor can consume a large amount of energy due to the tag checks performed. Recent proposals to address this issue involve predicting or memoizing the correct way to access. However, they also require significant hardware storage which negates much of the energy saving. This paper proposes way-placement to save instruction cache energy. The compiler places the most frequently executed instructions at the start of the binary and at runtime these are mapped to explicit ways within the cache. We compare with a state-of-the-art hardware technique and show that our scheme saves almost 50% of the instruction cache energy compared to 32% for the hardware approach. We report results on a variety of cache sizes and associativities, achieving 59% instruction cache energy savings and an ED product of 0.80 in the best configuration with negligible hardware overhead and no ISA changes.
Timothy M. Jones 0001, Sandro Bartolini, Bruno De Bus, John Cavazos, Michael F. P. O'Boyle
DATE2
2008 Effects of Instruction-Set Extensions on an Embedded Processor: A Case Study on Elliptic Curve Cryptography over GF(2m)
abstract
Elliptic-Curve cryptography (ECC) is promising for enabling information security in constrained embedded devices. In order to be efficient on a target architecture, ECCs require accurate choice/tuning of the algorithms that perform the underlying mathematical operations. This paper contributes with a cycle-level analysis of the dependencies of ECC performance from the interaction between the features of the mathematical algorithms and the actual architectural and microarchitectural features of an ARM-based Intel XScale processor. Another contribution is the cycle-level analysis of a modified ARM processor that includes a word-level finite field polynomial multiplier (poly_mul) in its data path. This extension constitutes a good trade-off between applicability in a number of contexts, the simplicity of integration within the processor, and performance. This paper points out the most advantageous mix of elliptic curve (EC) parameters both for the standard ARM-based Intel XScale platform and for the one equipped with the polyjnul unit. In particular, the latter case allows for more than 41 percent execution time reduction on the considered benchmarks. Last, this paper investigates the correlation between the possible architectural organizations of a processor equipped with poly_mul unit(s) and EC benchmark performance. For instance, only superscalar pipelines can exploit the features of out-of-order execution and only very complex organizations (for example, four way superscalar) can exploit a high number of available ALUs. Conversely, we show that there are no benefits in endowing the processor with more than one poly_mul, and we point out a possible trade-off between performance and complexity increase: A two-way in-order/out-of-order pipeline allows +50 percent and +90 percent of Instructions per Cycle (IPC), respectively. Finally, we show that there are no critical constraints on the latency and pipelining capability of the polyjnul unit for the basic EC point multiplication.
Sandro Bartolini, Irina Branovic, Roberto Giorgi, Enrico Martinelli
IEEE Trans. Computers1
2007 Inclusion of a Montgomery Multiplier Unit into an Embedded Processor's Datapath to Speed-up Elliptic Curve Cryptography
abstract
This paper analyzes the effects of including a full-width GF(2m) Montgomery multiplier within the datapath of an existing embedded processor, aiming to speed-up elliptic curve cryptography (ECC). This approach tends to exploit the tight coupling between the new and the other processor modules while maintaining both software compatibility and high flexibility to adapt to different ECC parameters and algorithms. In addition, the present work focuses on the effects on performance due to the interaction between the new unit and the other processor parts. We show that the modified ARM processor runs the ECC critical operation (kP) 9-times faster than in pure software and up to 14-times faster using 3 units and optimized instruction scheduling. Moreover, the improved processor achieves the same performance with 1/4 sized caches thanks to more than 93% reduction of memory traffic.
Sandro Bartolini, Cinzia Castagnini, Enrico Martinelli
IAS1
2005 Optimizing instruction cache performance of embedded systems
abstract
In the embedded domain, the gap between memory and processor performance and the increase in application complexity need to be supported without wasting precious system resources: die size, power, etc. For these reasons, effective exploitation of small and simple cache memories is of the utmost importance. However, programs running on such caches can experience serious inefficiencies due to cache conflicts.We present a new Cache-Aware Code Allocation Technique (CAT), which transforms the structure of programs so that their behavior toward memory can meet the locality features the cache is able to exploit. The proposed approach uses detailed information of program execution to place program areas into memory and employs the new idea of “look-forward estimation” that helps to seek better global layouts during the placement of each area. CAT-optimized programs outperform the original ones achieving the same miss rate on two times, and sometimes four times, smaller caches. Moreover, CAT improves the instruction miss rate by more than 40% if compared to the best procedure-reordering algorithm. CAT performances derive from the increased number of cache lines that support the execution of optimized applications and from a more balanced load on them.
Sandro Bartolini, Cosimo Antonio Prete
ACM Trans. Embed. Comput. Syst.1
2004 A Performance Evaluation of ARM ISA Extension for Elliptic Curve Cryptography over Binary Finite Fields
abstract
In this paper, we present an evaluation of possible ARM instruction set extension for elliptic curve cryptography (ECC) over binary finite fields GF(2/sup m/). The use of elliptic curve cryptography is becoming common in embedded domain, where its reduced key size at a security level equivalent to standard public-key methods (such as RSA) allows for power consumption savings and more efficient operation. ARM processor was selected because it is widely used for embedded system applications. We developed an ECC benchmark set with three widely used public-key algorithms: Diffie-Hellman for key exchange, digital signature algorithm, as well as El-Gamal method for encryption/decryption. We analyzed the major bottlenecks at function level and evaluated the performance improvement, when we introduce some simple architectural support in the ARM ISA. Results of our experiments show that the use of a word-level multiplication instruction over binary field allows for an average 33% reduction of the total number of dynamically executed instructions, while execution time improves by the same amount when projective coordinates are used.
Sandro Bartolini, Irina Branovic, Roberto Giorgi, Enrico Martinelli
SBAC-PAD1
2002 A cache-aware program transformation technique suitable for embedded systems
Sandro Bartolini, Cosimo Antonio Prete
Inf. Softw. Technol.1