Hamid Tabani

dblp:135/6712 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-6061-7470ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Hardware support for contention tracking in CPU and GPU last-level cache
abstract
Modern MPSoCs increasingly rely on resource utilization to improve application performance with different computation needs. The last-level cache (LLC) is one of the main shared resources, contributing to the improvement of aggregated performance. However, LLC sharing also increases individual application performance variability, which is undesirable in scenarios where performance guarantees are required. While deploying cache partitioning mechanisms allows regaining predictability, they negatively affect aggregated performance. This confronts system designers with the dire conundrum of choosing between aggregated performance and predictability. We contend that adding hardware support to track contention among tasks (kernels) in the LLC enables it to be shared, removing shortcomings brought by partitioning while providing a clear view of how tasks (kernels) affect each other in the LLC of the CPUs and GPUs. This approach enables achieving the desired balance between performance and predictability. Thus, we propose a low-overhead hardware mechanism, called demotion counters (DC), that tightly estimates the contention tasks (kernels) generate on each other in the shared LLC, outperforming other solutions that build on existing hardware contention-tracking proposals which suffer an average workload breakdown deviation (wbd) over 0.13. Our results also show that DC introduces 0.66% area overhead. 1
Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla
J. Syst. Archit.3
2023 An automotive case study on the limits of approximation for object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001, Francesc Moll, Enric Morancho, Ramon Canal, Josep Altet, Antonio Calomarde, Francisco J. Cazorla, Antonio Rubio 0001, Pau Fontova, Jordi Fornt
J. Syst. Archit.2
2023 Dynamic and execution views to improve validation, testing, and optimization of autonomous driving software
Miguel Alcon, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla
Softw. Qual. J.2
2023 Vector Extensions in COTS Processors to Increase Guaranteed Performance in Real-Time Systems
abstract
The need for increased application performance in high-integrity systems such as those in avionics is on the rise as software continues to implement more complex functionalities. The prevalent computing solution for future high-integrity embedded products is multi-processor systems-on-chip (MPSoC) processors. MPSoCs include central processing unit (CPU) multicores that enable improving performance via thread-level parallelism. MPSoCs also include generic accelerators (graphics processing units [GPUs]) and application-specific accelerators. However, the data processing approach (DPA) required to exploit each of these underlying parallel hardware blocks carries several open challenges to enable the safe deployment in high-integrity domains. The main challenges include the qualification of its associated runtime system and the difficulties in analyzing programs deploying the DPA with out-of-the-box timing analysis and code coverage tools. In this work, we perform a thorough analysis of vector extensions (VExts) in current commercial off-the-shelf (COTS) processors for high-integrity systems. We show that VExts prevent many of the challenges arising with parallel programming models and GPUs. Unlike other DPAs, VExts require no runtime support, prevent design race conditions that might arise with parallel programming models, and have minimum impact on the software ecosystem, enabling the use of existing code coverage and timing analysis tools. We develop vectorized versions of neural network kernels and show that the NVIDIA Xavier VExts provide a reasonable increase in guaranteed application performance of up to 2.7x. Our analysis contends that VExts are the DPA approach with arguably the fastest path for adoption in high-integrity systems.
Roger Pujol, Josep Jorba 0002, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla
ACM Trans. Embed. Comput. Syst.3
2022 Contention Tracking in GPU Last-Level Cache
abstract
The Last-level cache (LLC) is one of the main GPU’s shared resources that contributes to improve performance but also increases individual kernel’s performance variability. This is detrimental in scenarios in which some level of performance predictability is required. While predictability can be regained by deploying cache partitioning (isolation) mechanisms, isolation negatively affects performance efficiency. This work shows that not partitioning the LLC and providing the ability to track the contention that kernels generate on each other allows them to share LLC space, hence increasing efficiency, while the system designer obtains a clear view of how each kernel affects each other in the LLC so as to balance performance and predictability goals. In this line, we propose GPU demotion counters (GDC), a low-overhead hardware mechanism to track contention that kernels generate on each other in the shared LLC.
Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla
ICCD3
2022 At-scale evaluation of weight clustering to enable energy-efficient object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001
J. Syst. Archit.2
2021 Empirical Evidence for MPSoCs in Critical Systems: The Case of NXP's T2080 Cache Coherence
abstract
The adoption of complex MPSoCs in critical realtime embedded systems mandates a detailed analysis of their architecture to facilitate certification. This analysis is hindered by the lack of a thorough understanding of the MPSoC system due to the unobvious and/or insufficiently documented behavior of some key hardware features. Confidence in those features can only be regained by building specific tests to both, assess whether their behavior matches specifications and unveil their behavior when it is not fully known a priori. In this line, in this work we develop a thorough understanding of the cache coherence protocol in the avionics-relevant NXP T2080 architecture.
Roger Pujol, Hamid Tabani, Jaume Abella 0001, Mohamed Hassan 0002, Francisco J. Cazorla
DATE2
2021 Enabling Unit Testing of Already-Integrated AI Software Systems: The Case of Apollo for Autonomous Driving
abstract
The advanced AI-based software used for autonomous driving comprises multiple highly-coupled modules that are data and control dependent. Deploying those already-integrated software frameworks makes unit testing, a fundamental step in the validation process of critical software, very challenging in safety-critical systems. To tackle this issue, in this paper, we show the steps we followed to develop standalone versions of the modules in an industry-level autonomous driving framework (Apollo) by applying several modifications to its architectural design. We show how the standalone modules have the same functional behavior as their integrated counterpart modules. We exemplify the benefits of standalone modules by performing incremental analysis of the software timing requirements of each module running on a heterogeneous System on Chip (SoC). This is a mandatory step to consolidate and integrate software modules guaranteeing timing constraints (e.g. related to freedom from interference) while maximizing SoC utilization.
Miguel Alcon, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla
DSD2
2021 Improving the Efficiency of Transformers for Resource-Constrained Devices
abstract
Transformers provide promising accuracy and have become popular and used in various domains such as natural language processing and computer vision. However, due to their massive number of model parameters, memory and computation requirements, they are not suitable for resource-constrained low-power devices. Even with high-performance and specialized devices, the memory bandwidth can become a performance-limiting bottleneck. In this paper, we present a performance analysis of state-of-the-art vision transformers on several devices. We propose to reduce the overall memory footprint and memory transfers by clustering the model parameters. We show that by using only 64 clusters to represent model parameters, it is possible to reduce the data transfer from the main memory by more than 4x, achieve up to 22% speedup and 39% energy savings on mobile devices with less than 0.1% accuracy loss.
Hamid Tabani, Ajay Balasubramaniam, Shabbir Marzban, Elahe Arani, Bahram Zonooz
DSD1
2021 Performance Analysis and Optimization Opportunities for NVIDIA Automotive GPUs
Hamid Tabani, Fabio Mazzocchetti, Pedro Benedicte, Jaume Abella 0001, Francisco J. Cazorla
J. Parallel Distributed Comput.1
2020 A Cross-Layer Review of Deep Learning Frameworks to Ease Their Optimization and Reuse
abstract
Machine learning and especially Deep Learning (DL) approaches are at the heart of many domains, from computer vision and speech processing to predicting trajectories in autonomous driving and data science. Those approaches mainly build upon Neural Networks (NNs), which are compute-intensive in nature. A plethora of frameworks, libraries and platforms have been deployed for the implementation of those NNs, but end users often lack guidance on what frameworks, platforms and libraries to use to obtain the best implementation for their particular needs. This paper analyzes the DL ecosystem providing a structured view of some of the main frameworks, platforms and libraries for DL implementation. We show how those DL applications build ultimately on some form of linear algebra operations such as matrix multiplication, vector addition, dot product and the like. This analysis allows understanding how optimizations of specific linear algebra functions for specific platforms can be effectively leveraged to maximize specific targets (e.g. performance or power-efficiency) at application level reusing components across frameworks and domains.
Hamid Tabani, Roger Pujol, Jaume Abella 0001, Francisco J. Cazorla
ISORC1
2020 Timing of Autonomous Driving Software: Problem Analysis and Prospects for Future Solutions
abstract
The software used to implement advanced functionalities in critical domains (e.g. autonomous operation) impairs software timing. This is not only due to the complexity of the underlying high-performance hardware deployed to provide the required levels of computing performance, but also due to the complexity, non-deterministic nature, and huge input space of the artificial intelligence (AI) algorithms used. In this paper, we focus on Apollo, an industrial-quality Autonomous Driving (AD) software framework: we statistically characterize its observed execution time variability and reason on the sources behind it. We discuss the main challenges and limitations in finding a satisfactory software timing analysis solution for Apollo and also show the main traits for the acceptability of statistical timing analysis techniques as a feasible path. While providing a consolidated solution for the software timing analysis of Apollo is a huge effort far beyond the scope of a single research paper, our work aims to set the basis for future and more elaborated techniques for the timing analysis of AD software.
Miguel Alcon, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla
RTAS2
2019 Assessing the Adherence of an Industrial Autonomous Driving Framework to ISO 26262 Software Guidelines
abstract
The complexity and size of Autonomous Driving (AD) software are comparably higher than that of software implementing other (standard) functionalities in the car. To make things worse, a big fraction of AD software is not specifically designed for the automotive (or any other critical) domain, but the mainstream market. This brings uncertainty on to which extent AD software adheres to guidelines in safety standards. In this paper, we present our experience in applying ISO 26262 -- the applicable functional safety standard for road vehicles -- software safety guidelines to industrial AD software, in particular, Apollo, a heterogeneous Autonomous Driving framework used extensively in industry. We provide quantitative and qualitative metrics of compliance for many ISO 26262 recommendations on software design, implementation, and testing.
Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla, Guillem Bernat
DAC1
2019 Generating and Exploiting Deep Learning Variants to Increase Heterogeneous Resource Utilization in the NVIDIA Xavier
abstract
Deep learning-based solutions and, in particular, deep neural networks (DNNs) are at the heart of several functionalities in critical-real time embedded systems (CRTES) from vision-based perception (object detection and tracking) systems to trajectory planning. As a result, several DNN instances simultaneously run at any time on the same computing platform. However, while modern GPUs offer a variety of computing elements (e.g. CPUs, GPUs, and specific accelerators) in which those DNN tasks can be executed depending on their computational requirements and temporal constraints, current DNNs are mainly programmed to exploit one of them, namely, regular cores in the GPU. This creates resource imbalance and under-utilization of GPU resources when executing several DNN instances, causing an increase in DNN tasks' execution time requirements. In this paper, (a) we develop different variants (implementations) of well-known DNN libraries used in the Apollo Autonomous Driving (AD) software for each of the computing elements of the latest NVIDIA Xavier SoC. Each variant can be configured to balance resource requirements and performance: the regular CPU core implementation that can run on 2, 4, and 6 cores; the GPU regular and Tensor core variants that can run in 4 or 8 GPU’s Streaming Multiprocessors (SM); and 1 or 2 NVIDIA’s Deep Learning Accelerators (NVDLA); (b) we show that each particular variant/configuration offers a different resource utilization/performance point; finally, (c) we show how those heterogeneous computing elements can be exploited by a static scheduler to sustain the execution of multiple and diverse DNN variants on the same platform.
Roger Pujol, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla
ECRTS2
2019 Performance Analysis and Optimization of Automotive GPUs
abstract
Advanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD) have drastically increased the performance demands of automotive systems. Suitable high-performance platforms building upon Graphic Processing Units (GPUs) have been developed to respond to this demand, being NVIDIA Jetson TX2 a relevant representative. However, whether high-performance GPU configurations are appropriate for automotive setups remains as an open question. This paper aims at providing light on this question by modelling an automotive GPU (Jetson TX2), analyzing its microarchitectural parameters against relevant benchmarks, and identifying specific configurations able to meaningfully increase performance within similar cost envelopes, or to decrease costs preserving original performance levels. Overall, our analysis opens the door to the optimization of automotive GPUs for further system efficiency.
Fabio Mazzocchetti, Pedro Benedicte, Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla
SBAC-PAD3
2018 A Novel Register Renaming Technique for Out-of-Order Processors
abstract
Modern superscalar processors support a large number of in-flight instructions, which requires sizeable register files. Conventional register renaming techniques allocate a new storage location, i.e. physical register, for every instruction whose destination is a logical register in order to remove false dependences. Physical registers are released in a conservative manner when the same logical register is redefined. For this reason, many cycles may happen between the last read and the release of a physical register, leading to suboptimal utilization of the register file. We have observed that for more than 50% of the instructions in SPECfp and more than 30% of the instructions in SPECint that have a destination register, the produced value has only a single consumer. In this case, the RAW dependence guarantees that the producer-consumer instructions pair will be executed in program order and, hence, the same physical register can be used to store the value produced by both instructions. In this paper, we propose a renaming technique that exploits this property to reduce the pressure on the register file. Our technique leverages physical register sharing by introducing minor changes in the register map table and the issue queue. We also describe how our renaming scheme supports precise exceptions. We evaluated our renaming technique on top of a modern out-of-order processor. Our experimental results show that it provides 6% speedup on average for the SPEC2006 benchmarks. Alternatively, our renaming scheme achieves the same performance while reducing the number of physical registers by 10.5%.
Hamid Tabani, José-María Arnau, Jordi Tubella, Antonio González 0001
HPCA1
2017 An Ultra Low-Power Hardware Accelerator for Acoustic Scoring in Speech Recognition
abstract
Accurate, real-time Automatic Speech Recognition (ASR) comes at a high energy cost, so accuracy has often to be sacrificed in order to fit the strict power constraints of mobile systems. However, accuracy is extremely important for the end-user, and today's systems are still unsatisfactory for many applications. The most critical component of an ASR system is the acoustic scoring, as it has a large impact on the accuracy of the system and takes up the bulk of execution time. The vast majority of ASR systems implement the acoustic scoring by means of Gaussian Mixture Models (GMMs), where the acoustic scores are obtained by evaluating multidimensional Gaussian distributions.In this paper, we propose a hardware accelerator for GMM evaluation that reduces the energy required for acoustic scoring by three orders of magnitude compared to solutions based on CPUs and GPUs. Our accelerator implements a lazy evaluation scheme where Gaussians are computed on demand, avoiding 50% of the computations. Furthermore, it employs a novel clustering scheme to reduce the size of the acoustic model, which results in 8x memory bandwidth savings with a negligible impact on accuracy. Finally, it includes a novel memoization scheme that avoids 74.88% of floating-point operations. The end design provides a 164x speedup and 3532x energy reduction when compared with a highly-tuned implementation running on a modern mobile CPU. Compared to a state-of-the-art mobile GPU, the GMM accelerator achieves 5.89x speedup over a highly optimized CUDA implementation, while reducing energy by 241x.
Hamid Tabani, José-María Arnau, Jordi Tubella, Antonio González 0001
PACT1
2013 Optimal task allocation for maximizing reliability in distributed real-time systems
abstract
Distributed system has been developed as a platform for huge computations. Reliability is one of the prominent issues in such systems. Many studies have been recently done to improve reliability by proper task allocation in distributed systems, but they have only considered some system constraints such as processing load, memory capacity, and communication rate. In this paper, we consider time constraint in form of task deadline to above-mentioned constraints in order to model and analyze reliability in distributed real-time systems. To maximize reliability besides satisfying the constraints, we proposed a new offline task allocation algorithm. The algorithm is Systematic Memory-based Simulated Annealing (SMSA) which uses a monotonic cooling schedule and limited memory to store recently visited solutions to prevent cycling. In addition, an effective greedy heuristic algorithm intensifies SMSA. For evaluating the algorithm, SMSA is compared with Genetic Algorithm (GA) and Simulated Annealing (SA). Results have shown that in contrast to SA and GA, SMSA obtains satisfactory reliability in reasonable execution time. Meanwhile, SMSA meets all deadlines same as SA and GA. Furthermore, SMSA results have low deviation from average reliability.
Hamid Reza Faragardi, Reza Shojaee, Mohammad Amin Keshtkar, Hamid Tabani
ICIS4
2013 An analytical model to evaluate reliability of cloud computing systems in the presence of QoS requirements
abstract
Cloud computing is widely referred as the next generation of computing systems. Reliability is a key metric for assessing performance in such systems. Redundancy and diversity are prevalent approaches to enhance reliability in Cloud Computing Systems (CCS). Proper resource allocation is an alternative approach to reliability improvement in such systems. In contrast to redundancy, appropriate resource allocation can improve system reliability without imposing extra cost. On the other hand, contemplating reliability irrespective of Quality of Service (QoS) requirements may be undesirable in most of CCSs. In this paper, we focus on resource allocation approach and introduce an analytical model in order to analyze system reliability besides considering application and resource constraints. Task precedence structure and QoS are taken into account as the application constraints. Memory and storage limitation of each server as well as maximum communication load on each link are considered as the principle resource constraints. In addition, effect of network topology on system reliability is discussed in detail and the model is extended to cover various network topologies.
Hamid Reza Faragardi, Reza Shojaee, Hamid Tabani, Aboozar Rajabi
ICIS3