Hamid Noori

dblp:78/4018 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0003-1410-6781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 MDSSD-MobV2: An embedded deconvolutional multispectral pedestrian detection based on SSD-MobileNetV2
Fereshteh Aghaee, Ehsan Fazl Ersi, Hamid Noori
Multim. Tools Appl.3
2022 Characterizing energy and performance of soft-core-based homogeneous multiprocessor systems
Farshid Samsami Khodadad, Hamid Noori
J. Supercomput.2
2022 Online energy-efficient fair scheduling for heterogeneous multi-cores considering shared resource contention
Bagher Salami, Hamid Noori, Mahmoud Naghibzadeh
J. Supercomput.2
2021 A novel approach based on genetic algorithm to speed up the discovery of classification rules on GPUs
abstract
This paper proposes a new approach to produce classification rules based on evolutionary computation with novel crossover and mutation operators customized for execution on graphics processing unit (GPU). Also, a novel method is presented to define the fitness function, i.e. the function which measures quantitatively the accuracy of the rule. The proposed fitness function is benefited from parallelism due to the parallel execution of data instances. To this end, two novel concepts; coverage matrix and reduction vectors are used and an altered form of the reduction vector is compared with previous works. Our CUDA program performs operations on coverage matrix and reduction vector in parallel. Also these data structures are used for evaluation of fitness function and calculation of genetic operators in parallel. We proposed a vector called average coverage to handle crossover and mutation properly. Our proposed method obtained a maximum accuracy of 99.74% for Hepatitis C Virus (HCV) dataset, 95.73% for Poker dataset, and 100% for COVID-19 dataset. Our speedup is higher than 20% for HCV and COVID-19, and 50% for Poker, compared to using single core processors.
Mohamad Beheshti Roui, Mariam Zomorodi Moghadam, Masoomeh Sarvelayati, Moloud Abdar, Hamid Noori, Pawel Plawiak, Ryszard Tadeusiewicz, Xujuan Zhou, Abbas Khosravi, Saeid Nahavandi, U. Rajendra Acharya
Knowl. Based Syst.5
2021 Fairness-Aware Energy Efficient Scheduling on Heterogeneous Multi-Core Processors
abstract
Heterogeneous multi-core processors (HMP) with the same instruction set architecture (ISA) integrate complex high performance big cores with power efficient small cores on the same chip. In comparison with homogeneous architectures, HMPs have been shown to significantly increase energy efficiency. However, current techniques to exploit the energy efficiency of HMPs do not consider fair usage of resources that leads to reduced performance predictability, a longer makespan, starvation, and QoS degradation. The effect of different cluster voltage and frequency levels on fairness is another issue neglected by previous task scheduling algorithms. The present study investigates both the fairness problem and energy efficiency in HMPs. This article proposes a heterogeneous fairness-aware energy efficient framework (HFEE) that employs DVFS to meet fairness constraints and provide energy efficient scheduling. The proposed framework is implemented and evaluated on a real heterogeneous multi-core processor. The experimental results indicate that the introduced technique can significantly improve energy efficiency and fairness when compared to Linux standard scheduler and two energy efficient and fairness-aware schedulers.
Bagher Salami, Hamid Noori, Mahmoud Naghibzadeh
IEEE Trans. Computers2
2021 PEPS: predictive energy-efficient parallel scheduler for multi-core processors
Zeinab Maghsoud, Hamid Noori, Saadat Pour Mozaffari
J. Supercomput.2
2020 Efficient scheduling of streams on GPGPUs
Mohamad Beheshti Roui, S. Kazem Shekofteh, Hamid Noori, Ahad Harati
J. Supercomput.3
2020 cCUDA: Effective Co-Scheduling of Concurrent Kernels on GPUs
abstract
While GPUs are meantime omnipresent for many scientific and technical computations, they still continue to evolve as processors. An important recent feature is the ability to execute multiple kernels concurrently via queue streams. However, experiments show that different parameters including the behavior of kernels, the order of kernel launches and other execution configurations, e.g., the number of concurrent thread blocks, may result in different execution time for concurrent kernel execution. Since kernels may have different resource requirements, they can be classified into different classes, which are traditionally assumed as either memory-bound or compute-bound. However, a kernel may belong to the different classes on different hardware according to the hardware resources. In this paper, the definition of kernel mix intensity is introduced. Based on this, a scheduling framework called concurrent CUDA (cCUDA) is proposed to co-schedule the concurrent kernels more efficiently. It first profiles and ranks kernels with different execution behaviors and then takes the kernel resource requirements into account to partition thread blocks of different kernels and overlap them to better utilize the GPU resources. Experimental results on real hardware demonstrate performance improvement in terms of execution time of up to 1.86x, and an average speedup of 1.28x for a wide range of kernels. cCUDA is available at https://github.com/kshekofteh/cCUDA.
S. Kazem Shekofteh, Hamid Noori, Mahmoud Naghibzadeh, Holger Fröning, Hadi Sadoghi Yazdi
IEEE Trans. Parallel Distributed Syst.2
2019 Metric Selection for GPU Kernel Classification
abstract
Graphics Processing Units (GPUs) are vastly used for running massively parallel programs. GPU kernels exhibit different behavior at runtime and can usually be classified in a simple form as either “compute-bound” or “memory-bound.” Recent GPUs are capable of concurrently running multiple kernels, which raises the question of how to most appropriately schedule kernels to achieve higher performance. In particular, co-scheduling of compute-bound and memory-bound kernels seems promising. However, its benefits as well as drawbacks must be determined along with which kernels should be selected for a concurrent execution. Classifying kernels can be performed online by instrumentation based on performance counters. This work conducts a thorough analysis of the metrics collected from various benchmarks from Rodinia and CUDA SDK. The goal is to find the minimum number of effective metrics that enables online classification of kernels with a low overhead. This study employs a wrapper-based feature selection method based on the Fisher feature selection criterion. The results of experiments show that to classify kernels with a high accuracy, only three and five metrics are sufficient on a Kepler and a Pascal GPU, respectively. The proposed method is then utilized for a runtime scheduler. The results show an average speedup of 1.18× and 1.1× compared with a serial and a random scheduler, respectively.
S. Kazem Shekofteh, Hamid Noori, Mahmoud Naghibzadeh, Hadi Sadoghi Yazdi, Holger Fröning
ACM Trans. Archit. Code Optim.2
2017 Performance evaluation metrics for ring-oscillator-based temperature sensors on FPGAs: A quality factor
Navid Rahmanikia, Amirali Amiri, Hamid Noori, Farhad Mehdipour
Integr.3
2016 Physical-aware predictive dynamic thermal management of multi-core processors
Bagher Salami, Hamid Noori, Farhad Mehdipour, Mohammadreza Baharani
J. Parallel Distributed Comput.2
2015 Exploring Efficiency of Ring Oscillator-Based Temperature Sensor Networks on FPGAs (Abstract Only)
abstract
Due to technology advances and complexity of designs, thermal issue is a bottleneck in electronics designs. Various dynamic thermal management techniques have been proposed to address this issue. To effectively apply thermal management techniques, providing an accurate thermal map of chips is highly required. For this goal, a network of temperature sensors ought to be provided. There are various implementations for temperature sensors and network of sensors on Field Programmable Gate Arrays (FPGAs). This work defines and formulates four metrics and criteria, in terms of area, thermal, and power overheads and thermal map accuracy for exploring and evaluating efficiency of different implementations of Ring Oscillator-based Temperature Sensor (ROTS) networks on FPGAs and reports the comparison results for 12 networks with various sensor configurations. According to our metrics and experiments, the sensor that it is composed of NOT gates with open latches and RNS ring counter has lower thermal and power overheads compared to other configurations. Moreover, in this work, a new ROTS is presented that occupies 25% less resources than the most compact temperature sensor. Also, it provides 1.72 times higher sensitivity than the best sensitive ROTS design.
Navid Rahmanikia, Amirali Amiri, Hamid Noori, Farhad Mehdipour
FPGA3
2015 Dynamic Task Priority Scaling for Thermal Management of Multi-core Processors with Heavy Workload
abstract
This paper presents a task priority scaling algorithm for dynamic thermal management of multi-core processors. The unique features of this algorithm include: 1) enabling task-level Dynamic Frequency Scaling (DFS) capability through software, 2) reducing task migration and provide load balancing using dynamic task priority scaling, 3) targeting DTM for systems with high workload. This algorithm is evaluated on a commercial quad-core processor. The experimental results indicate that the proposed approach can decrease the average and peak temperature by 9.73% and 7.1%, respectively, compared to Linux standard scheduler.
Saadat Pour-Mozafari, Hamid Noori, Farhad Mehdipour
ACM Great Lakes Symposium on VLSI3
2015 Critical path-aware voltage island partitioning and floorplanning for hard real-time embedded systems
Aminollah Mahabadi, Ahmad Khonsari, Behnam Khodabandeloo, Hamid Noori, Alireza Majidi
Integr.4
2014 Physical-aware task migration algorithm for dynamic thermal management of SMT multi-core processors
abstract
This paper presents a task migration algorithm for dynamic thermal management of Simultaneous Multi-Threading (SMT) multi-core processors. The unique features of this algorithm include: 1) considering SMT capability of processors for dynamic thermal management via task scheduling, 2) using adaptive task migration threshold, and 3) considering cores physical features. This algorithm is evaluated on a commercial SMT quad-core processor. The experimental results indicate that our technique can significantly decrease the average and peak temperature compared to Linux standard scheduler, and two well-known thermal management techniques.
Bagher Salami, Mohammadreza Baharani, Hamid Noori, Farhad Mehdipour
ASP-DAC3
2014 High-level design space exploration of locally linear neuro-fuzzy models for embedded systems
Mohammadreza Baharani, Hamid Noori, Mohammad Aliasgari, Zainalabedin Navabi
Fuzzy Sets Syst.2
2014 Proactive task migration with a self-adjusting migration threshold for dynamic thermal management of multi-core processors
Bagher Salami, Mohammadreza Baharani, Hamid Noori
J. Supercomput.3
2012 Improving performance and energy efficiency of embedded processors via post-fabrication instruction set customization
Hamid Noori, Farhad Mehdipour, Koji Inoue, Kazuaki J. Murakami
J. Supercomput.1
2011 Instruction and data cache peak temperature reduction using cache access balancing in embedded processors
abstract
In this work we study cache peak temperature variation under different cache access patterns. In particular we show that unbalanced cache access results in higher cache peak temperature. This is the result of frequent accesses made to overused cache sets. Moreover we study cache peak temperature under cache access balancing techniques and show that exploiting such techniques not only reduces cache miss rate but also results in lower peak temperature. Our study shows that balancing cache access reduces peak temperature by up to 20% and 12% for instruction and data caches respectively. This temperature reduction reduces peak temperature in neighbor components by up to 7%.
Mohsen Taherian, Amirali Baniasadi, Hamid Noori
AICCSA3
2011 Securing Embedded Processors against Power Analysis Based Side Channel Attacks Using Reconfigurable Architecture
abstract
Power analysis based side channel attacks are significant security risk in embedded applications. Reconfigurable architecture has already been proposed as a security improvement method for run time monitoring systems or implementing critical parts of cryptographic applications. Here we propose reconfigurable architecture as a hardware countermeasure against power analysis based side channel attacks. We augment an embedded processor with a reconfigurable functional unit (RFU). By random execution of custom instructions on the RFU we mask power analysis based side channel attacks. Moreover we devised an automatic design flow to generate RFU and its configuration bits from cryptographic algorithms' source code. Obfuscation of base processor power traces as our primary goal is complied with RFUs covering in average about 30% of object code. We also report the power traces for our secure processor and the overall timing and area overhead. The correlation coefficient is calculated for AES and SHA cryptographic algorithms and experimental results show our method produces power traces close to random traces. Our approach is completely generic and can be used for any cryptographic application. Compared to previous methods, our work costs no runtime overhead and an average of 27% area overhead.
Sahar Abbaspour Seyyedi, Mehdi Kamal, Hamid Noori, Saeed Safari
EUC3
2010 Dual-purpose custom instruction identification algorithm based on Particle Swarm Optimization
abstract
Extending instruction set architecture (ISA) of embedded processors is an effective way to enhance performance and energy efficiency. The typical approaches for identifying custom instructions (CIs) limit the maximum number of input and output (I/O) operands to the available register file port. Recently, there are several work that explore CI candidates without imposing a limit on the number of input and output operands. In this paper, we present a new algorithm based on Particle Swarm Optimization (PSO) to identify CIs within a given data flow graph (DFG) and evaluate it for both categories of CI identification approaches (with and without I/O constrains). By novel evolving strategy, we enhance the quality of the results in our partitioning algorithm. Experimental results show that in most cases CI identification with I/O constraints based on PSO finds better or the same CIs in terms of performance compared to genetic algorithm (GA)[1] and ISEGEN [2] (96% and 90%, respectively). Comparing our proposed algorithm with [12] and [13] reveals that ours has a shorter run-time several order of magnitudes for large DFGs and is independent of the number of forbidden nodes. Moreover, we propose a modified version of PSO called Wrapper PSO that is up to 100× and 500× faster than GA and ISEGEN in large DFGs, respectively.
Mehdi Kamal, Neda Kazemian Amiri, Arezoo Kamran, Seyyed Alireza Hoseini, Masoud Dehyadegari, Hamid Noori
ASAP6
2009 A combined analytical and simulation-based model for performance evaluation of a reconfigurable instruction set processor
abstract
Performance evaluation is a serious challenge in designing or optimizing reconfigurable instruction set processors. The conventional approaches based on synthesis and simulations are very time consuming and need a considerable design effort. A combined analytical and simulation-based model (CAnSO*) is proposed and validated for performance evaluation of a typical reconfigurable instruction set processor. The proposed model consists of an analytical core that incorporates statistics gathered from cycle-accurate simulation to make a reasonable evaluation and provide a valuable insight. Compared to cycle-accurate simulation results, CAnSO proves almost 2% variation in the speedup measurement.
Farhad Mehdipour, Hamid Noori, Bahman Javadi, Hiroaki Honda, Koji Inoue, Kazuaki J. Murakami
ASP-DAC2
2008 Design space exploration for a coarse grain accelerator
abstract
In the design process of a reconfigurable accelerator employing in an embedded system, multitude parameters may result in remarkable complexity and a large design space. Design space exploration as an alternative to the quantitative approach can be employed to find a right balance between the different design parameters. In this paper, a hybrid approach is introduced to analytically explore the design space for a coarse grain accelerator and determine a wise design point exploiting data extracted from applications, quantitatively. It also provides flexibility for taking into account new design constraints as well as new characteristics of applications. Furthermore, this approach is a methodological approach which reduces the design time and results in a point which satisfies the design goals.
Farhad Mehdipour, Hamid Noori, Morteza Saheb Zamani, Koji Inoue, Kazuaki J. Murakami
ASP-DAC2
2008 Variation-Aware Software Techniques for Cache Leakage Reduction Using Value-Dependence of SRAM Leakage Due to Within-Die Process Variation
Maziar Goudarzi, Tohru Ishihara, Hamid Noori
HiPEAC3
2008 Enhancing energy efficiency of processor-based embedded systems through post-fabrication ISA extension
abstract
Application-specific instruction set extension is an effective technique for reducing accesses to components such as on- and off-chip memories, register file and enhancing the energy efficiency. However, the addition of custom functional units to the base processor is required for supporting custom instructions, which due to the increase of manufacturing and design costs in new nanometer-scale technologies and shorter time-to-market, is becoming an issue. To address above issues, in our proposed approach, an optimized reconfigurable functional unit is used instead, and instruction set customization is done after chip-fabrication. Therefore, while maintaining the flexibility of a conventional microprocessor, the low-energy feature of customization is applicable. Experimental results show that the maximum and average energy savings are 67% and 22%, respectively for our proposed architecture framework.
Hamid Noori, Farhad Mehdipour, Koji Inoue, Kazuaki J. Murakami
ISLPED1
2008 An architecture framework for an adaptive extensible processor
Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Morteza Saheb Zamani
J. Supercomput.1
2007 Interactive presentation: Generating and executing multi-exit custom instructions for an adaptive extensible processor
abstract
To improve the performance of embedded processors, an effective technique is collapsing critical computation subgraphs as application-specific instruction set extensions and executing them on custom functional units. The problems of this approach are immense cost and long time of designing. To address these issues, an adaptive extensible processor was proposed in which custom instructions (CIs) are generated and added after chip-fabrication. To support this feature, custom functional units are replaced by a reconfigurable matrix of functional units with the capability of conditional execution. Unlike previous proposed CIs, it can include multiple exits. Experimental results show that multi-exit CIs enhance the performance by 46% in average compared to CIs limited to one basic block. A maximum speedup of 2.89 compared to a 4-issue in-order RISC processor, and a speedup of 1.66 in average, was achieved on MiBench benchmark suite
Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Maziar Goudarzi
DATE1
2007 The effect of temperature on cache size tuning for low energy embedded systems
abstract
Energy consumption is a major concern in embedded computing systems. Several studies have shown that cache memories account for about 40% or more of the total energy consumed in these systems. In older technology nodes, active power was the primary contributor to total power dissipation of a CMOS design. However, with the scaling of feature sizes, the share of leakage in total power consumption of digital systems continues to grow. Temperature is a factor which exponentially increases the leakage current. In this paper, we show the effects of temperature on the selection of optimal cache size for low energy embedded systems. Our results show that for a given application, the optimal cache size selection is affected by the temperature. Our experiments have been done for 100nm technology. Our study reveals that the cache size selection for different temperatures depends on the rate at which cache miss increases when reducing the cache size. When the miss rate increases sharply the optimal point is the same for all examined temperatures, however when it becomes smoother, the optimal point for different temperatures begin to get farther.
Hamid Noori, Maziar Goudarzi, Koji Inoue, Kazuaki J. Murakami
ACM Great Lakes Symposium on VLSI1
2006 Custom Instruction Generation Using Temporal Partitioning Techniques for a Reconfigurable Functional Unit
Farhad Mehdipour, Hamid Noori, Morteza Saheb Zamani, Kazuaki J. Murakami, Koji Inoue, Mehdi Sedighi
EUC2
2006 A Reconfigurable Functional Unit for an Adaptive Dynamic Extensible Processor
abstract
This paper presents a reconfigurable functional unit (RFU) for an adaptive dynamic extensible processor. The processor can tune its extended instructions to the target applications, after chip-fabrication. The custom instructions (CIs) are generated deploying the hot basic blocks during the training mode. In the normal mode, CIs are executed on the RFU. A quantitative approach was used for designing the RFU. The RFU is a matrix of functional units with 8 inputs and 6 outputs. Performance is enhanced up to 1.25 using the proposed RFU for 22 applications of Mibench. This processor needs no extra opcodes for CIs, new compiler, source code modification and recompilation.
Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Morteza Saheb Zamani
FPL1
2005 Efficient Host-Independent Coprocessor Architecture for Speech Coding Algorithms
abstract
The recent growth of cellular phone systems, voice over IP devices, and other multimedia applications has created a considerable need for efficient voice coding algorithms. These algorithms usually require intensive amount of signal processing capabilities and demand significant signal processing power. The current market trend of integrating multiple voice channels into a single die has further intensified the need for more powerful hardware platforms. Some new design ideas such as vocoder-specialized DSP architectures, combined RISC/DSP platforms, and adding hardware accelerators or coprocessors to the general-purpose processors have been proposed. In this paper, a new hardware accelerator design has been proposed which executes macro instructions (MIs). The proposed coprocessor can be added to each processor type that can support at least one coprocessor without modifying the compiler and redesigning the processor. It can handle computationally intensive loops in speech coding algorithms parallel with the main processor. The coprocessor along with software optimization reduces clock cycles required for G.723.1 by 80% and G.729 by 64% while MIPS R3000 RISC is used as the host.
Hamid Safizadeh, Hamid Noori, Mehdi Sedighi, Ali Jahanian 0001, Neda Zolfaghari
DSD2
2001 Motivation from a Full-Rate Specific Design to a DSP Core Approach for GSM Vocoders
Shervin Sheidaei, Hamid Noori, A. Akbariazirani, Hossein Pedram
FPL2