EDBT 2026 Demo / reviewers in the wild / expert
Masoumeh Ebrahimi
dblp:74/7355
· DBLP profile ↗
56ranked-venue papers
20as first author
11since 2021 · last 2022
0000-0001-7877-6712ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 14 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 4 first-authorComputer networks · 2 · 2 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Inference Time Reduction of Deep Neural Networks on Embedded Devices: A Case StudyabstractFrom object detection to semantic segmentation, deep learning has achieved many groundbreaking results in recent years. However, due to the increasing complexity, the execution of neural networks on embedded platforms is greatly hindered. This has motivated the development of several neural network minimisation techniques, amongst which pruning has gained a lot of focus. In this work, we perform a case study on a series of methods with the goal of finding a small model that could run fast on embedded devices. First, we suggest a simple, but effective, ranking criterion for filter pruning called Mean Weight. Then, we combine this new criterion with a threshold-aware layer-sensitive filter pruning method, called T-sensitive pruning, to gain high accuracy. Further, the pruning algorithm follows a structured filter pruning approach that removes all selected filters and their dependencies from the DNN model, leading to less computations, and thus low inference time in lower-end CPUs. To validate the effectiveness of the proposed method, we perform experiments on three different datasets (with 3, 101, and 1000 classes) and two different deep neural networks (i.e., SICK-Net and MobileNet V1). We have obtained speedups of up to 13x on lower-end CPUs (Armv8) with less than 1% drop in accuracy. This satisfies the goal of transferring deep neural networks to embedded hardware while attaining a good trade-off between inference time and accuracy. Isma-Ilou Sadou, Seyed Morteza Nabavinejad, Zhonghai Lu, Masoumeh Ebrahimi |
DSD | 4 |
| 2022 | Policy-Induced Unsupervised Feature Selection: A Networking Case StudyabstractA promising approach for leveraging the flexibility and mitigating the complexity of future telecom systems is the use of machine learning (ML) models that can analyze the network performance, as well as taking proactive actions. A key enabler for ML models is timely access to reliable data, in terms of features, which require pervasive measurement points throughout the network. However, excessive monitoring is associated with network overhead. Considering domain knowledge may provide clues to find a balance between overhead reduction and meeting requirements on future ML use cases by monitoring just enough features. In this work, we propose a method of unsupervised feature selection that provides a structured approach in incorporation of the domain knowledge in terms of policies. Policies are provided to the method in form of must-have features defined as the features that need to be monitored at all times. We name such family of unsupervised feature selection the policy-induced unsupervised feature selection as the policies inform selection of the latent features. We evaluate the performance of the method on two rich sets of data traces collected from a data center and a 5G-mmWave testbed. Our empirical evaluations point at the effectiveness of the solution. Jalil Taghia, Farnaz Moradi 0001, Hannes Larsson, Xiaoyu Lan, Masoumeh Ebrahimi, Andreas Johnsson |
INFOCOM | 5 |
| 2022 | Exploring Approaches for Heterogeneous Transfer Learning in Dynamic NetworksabstractMaintaining machine-learning models for prediction of service performance is challenging, especially in dynamic network and cloud environments where route changes occur, and execution environments can be scaled and migrated. Recently, transfer learning has been proposed as an approach for leveraging already learned knowledge in a new environment. The challenge is that the new environment may be significantly different from the one the model is trained in, and transferred from, with respect to data distributions and dimensionality.In this paper, we introduce heterogeneous transfer learning in the context of dynamic environments and show its efficiency in predicting service performance. We propose two heterogeneous transfer-learning approaches and evaluate them on several neural-network architectures and scenarios. The scenarios are a natural consequence of network and cloud infrastructure reorchestration. We quantify the transfer gain, and empirically show positive gain in a majority of cases for both approaches. Furthermore, we study the impact of neural-network configurations on the transfer gain, providing tradeoff insights. The evaluation of the approaches is performed using data traces collected from a cloud testbed that runs two services under multiple realistic load conditions. Fernando García Sanz, Masoumeh Ebrahimi, Andreas Johnsson |
NOMS | 2 |
| 2022 | Coordinated Batching and DVFS for DNN Inference on GPU AcceleratorsabstractEmploying hardware accelerators to improve the performance and energy-efficiency of DNN applications is on the rise. One challenge of using hardware accelerators, including the GPU-based ones, is that their performance is limited by internal and external factors, such as power caps. A common approach to meet the power cap constraint is using the Dynamic Voltage Frequency Scaling (DVFS) technique. However, the functionally of this technique is limited and platform-dependent. To tackle this challenge, we propose a new control knob, which is the size of input batches fed to the GPU accelerator in DNN inference applications. We first evaluate the impact of batch size on power consumption and performance of DNN inference. Then, we introduce the design and implementation of a fast and lightweight runtime system, called BatchDVFS. Dynamic batching is implemented inBatchDVFSto adaptively change the batch size, and hence, trade-off throughput with power consumption. It employs an approach based on binary search to find the proper batch size within a short period of time. Combining dynamic batching with the DVFS technique,BatchDVFScan control the power consumption in wider ranges, and hence, yield higher throughput in the presence of power caps. To find near-optimal solution for long-running jobs that can afford a relatively significant profiling overhead, compared withBatchDVFSoverhead, we also design an approach, called BOBD, that employs Bayesian Optimization to wisely explore the vast state space resulted by combination of the batch size and DVFS solutions. Conducting several experiments using a modern GPU and several DNN models and input datasets, we show that ourBatchDVFScan significantly surpass the techniques solely based on DVFS or batching, regarding throughput (up to 11.2x and 2.2x, respectively), while successfully meeting the power cap. Seyed Morteza Nabavinejad, Sherief Reda, Masoumeh Ebrahimi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | BatchSizer: Power-Performance Trade-off for DNN InferenceabstractGPU accelerators can deliver significant improvement for DNN processing; however, their performance is limited by internal and external parameters. A well-known parameter that restricts the performance of various computing platforms in real-world setups, including GPU accelerators, is the power cap imposed usually by an external power controller. A common approach to meet the power cap constraint is using the Dynamic Voltage Frequency Scaling (DVFS) technique. However, the functionally of this technique is limited and platform-dependent. To improve the performance of DNN inference on GPU accelerators, we propose a new control knob, which is the size of input batches fed to the GPU accelerator in DNN inference applications. After evaluating the impact of this control knob on power consumption and performance of GPU accelerators and DNN inference applications, we introduce the design and implementation of a fast and lightweight runtime system, called BatchSizer. This runtime system leverages the new control knob for managing the power consumption of GPU accelerators in the presence of the power cap. Conducting several experiments using a modern GPU and several DNN models and input datasets, we show that our BatchSizer can significantly surpass the conventional DVFS technique regarding performance (up to 29%), while successfully meeting the power cap. Seyed Morteza Nabavinejad, Sherief Reda, Masoumeh Ebrahimi |
ASP-DAC | 3 |
| 2021 | Hierarchical Fault Simulation of Deep Neural Networks on Multi-Core SystemsabstractIn this paper, a hierarchical fault simulation technique for neural networks is proposed, supporting both permanent and temporary faults. In the proposed technique, different levels of hierarchy are used, forming a mixed-level simulation environment. In such an environment, the pre-synthesis behavioral specification of the network and the post-synthesis gate-level model are co-simulated. To accelerate the fault simulation process, faults are injected in the gate-level specification of the selected neurons while the behavioral model in different levels of abstraction is used to simulate the remaining neurons. Further speedup is obtained through event-driven simulation and parallelization. Experimental results confirm the time efficiency of the proposed fault simulation technique. Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Antonio Miele, Hannu Tenhunen, Juha Plosila |
ETS | 3 |
| 2021 | On Heterogeneous Transfer Learning for Improved Network Service Performance PredictionabstractTransfer learning has been proposed as an approach for leveraging already learned knowledge in a new environment, especially when the amount of training data is limited. However, due to the dynamic nature of future networks and cloud infrastructures, a new environment may differ from the one the model is trained and transferred from. In this paper, we propose and evaluate an approach based on neural networks for heterogeneous transfer learning that addresses model transfer between environments with different input feature sets, which is a natural consequence of network and cloud re-orchestration. We quantify the transfer gain, and empirically show positive gain in a majority of cases. Further, we study the impact of neural-network architectures on the transfer gain, providing tradeoff insights for multiple cases. The evaluation of the approach is performed using data traces collected from a testbed that runs a Video-on-Demand service and a Key-Value Store under various load conditions. Fernando García Sanz, Masoumeh Ebrahimi, Andreas Johnsson |
GLOBECOM | 2 |
| 2021 | SRAM Gauge: SRAM Health Monitoring via Cells RaceabstractBy shrinking transistors’ dimensions and, consequently, reducing the operating voltage in nano-scale CMOS technologies, the stability of SRAM cells has become a major reliability concern. SRAM cells’ robustness against undesirable bit-flips is commonly measured by Static Noise Margin (SNM). Degradation in SNM is mainly because of the gradual variations in transistors’ parameters due to aging. This work proposes a built-in SRAM health sensor capable of monitoring the SNM of individual SRAM cells in a memory block. The sensor is composed of extra non-operational sensor cells with different predefined SNMs. These sensor cells are put in a race with operational SRAM cells to determine their strength. The precision, sensing range, and robustness of the proposed sensor against process variation are adjustable at the cost of small area overhead. In our simulation setup, with the area overhead of 0.29%, the sensor monitors a wide range of SNMs from 275mV to 325mV, with a precision of 5mV. Nezam Rohbani, Masoumeh Ebrahimi |
ISLPED | 2 |
| 2021 | Application Characterization for Near Memory ProcessingabstractData movement between memory subsystem and processor unit is a crippling performance and energy bottleneck for data-intensive applications. Near Memory Processing (NMP) is a promising solution to alleviate the data movement bottleneck. The introduction of 3D-stacked memories and more importantly hybrid memory systems enable the long-wished NMP capability. This work explores the feasibility and efficacy of having NMP on the hybrid memory system for a given set of applications. In this paper, we first redefine a set of NMP-centric performance metrics in order to analyze the efficacy of a given processing unit. Leveraging the proposed metrics, we characterize various sets of applications to assess the suitability of a processing unit in terms of performance. Specifically, in this work we motivate the efficiency of NMP subsystems to process memory-intensive applications when 3D-NVM technologies are employed. Maryam S. Hosseini, Masoumeh Ebrahimi, Pooria M. Yaghini, Nader Bagherzadeh |
PDP | 2 |
| 2021 | High-Performance Parallel Fault Simulation for Multi-Core SystemsabstractFault simulation is a time-consuming process that requires customized methods and techniques to accelerate it. Multi-threading and Multi-core approaches are two promising techniques that can be exploited to accelerate the fault simulation process by using different parts of the hardware at the same time. However, an efficient parallelization is obtained only by the refinement of software with respect to the hardware platform. In this paper, a parallel multi-thread fault simulation technique is proposed to accelerate the simulation process on multi-core platforms. In this approach, the gate input values are independently assigned to each thread. Each input value carries the information of several parallel simulation processes. This provides a multithread parallel fault simulation environment. The experimental results show that the proposed technique can efficiently use the hardware platform. In a single-core platform, the proposed technique can reduce the time by 25% while in a dual-core increasing the thread approximately halves the execution time. Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Hamid Nejatollahi, Hannu Tenhunen, Juha Plosila |
PDP | 3 |
| 2021 | ECDR$^{2}$2: Error Corrector and Detector Relocation Router for Network-on-ChipabstractNetwork-on-chip (NoC) is commonly used in modern many-core systems due to their high bandwidth and flexibility. As the manufacturing process keeps scaling, the reliability challenge in NoCs increases as well. The error correction code (ECC) is widely adopted in error correction NoCs to improve the data correctness. At the same time, extra stages are introduced in the router pipeline to improve the error correction capability. As a result, conventional error correction routers suffer from high network latency. Motivated by this limitation, i.e., we remove the extra pipeline stages delicately introduced for error correction. We propose an error correction router, called error corrector and detector relocation router (ECDR2), whose architecture optimizes the pipeline flow of the router. As a result, it can achieve both low latency and high error correction. Experimental results show that, compared with the baseline design, ECDR2obtains 13.67 and 39.4 percent less average latency under the uniform traffic pattern and Dedup benchmark, respectively, in an 8 × 8 mesh NoC. The circuit area of ECR is also 7.9 percent less than that of the baseline design under 45-nm technology. Letian Huang, Chikun Yuan, Junshi Wang, Masoumeh Ebrahimi, Qiang Li 0021 |
IEEE Trans. Computers | 4 |
| 2019 | Online Path-Based Test Method for Network-on-ChipabstractA considerable amount of routers and links remains idle after each mapping application onto the Network-on-Chip based many-core systems. Online path-based test method is a kind of self-test for these idle components. In this paper, a path-based fabric for NoC is firstly proposed. A path serves as the basic component, covering one link and its associated control logic in the routers. One possibility is to apply fault detection on the idle paths, while the other paths continue to operate normally. Moreover, this paper details the hardware implementation, targeting the stuck-at and bridging faults. It suggests a good trade-off between fault coverage, hardware overhead and test time. Experimental results show that the approach achieves 93% of the stuck-at faults in control unit and cover 100% of the stuck-at and bridging faults on the global link within 256 clock cycles. Junkai Zhan, Letian Huang, Junshi Wang, Masoumeh Ebrahimi, Qiang Li 0021 |
ISCAS | 4 |
| 2019 | NoC-based DNN accelerator: a future design paradigmabstractDeep Neural Networks (DNN) have shown significant advantages in many domains such as pattern recognition, prediction, and control optimization. The edge computing demand in the Internet-of-Things era has motivated many kinds of computing platforms to accelerate the DNN operations. The most common platforms are CPU, GPU, ASIC, and FPGA. However, these platforms suffer from low performance (i.e., CPU and GPU), large power consumption (i.e., CPU, GPU, ASIC, and FPGA), or low computational flexibility at runtime (i.e., FPGA and ASIC). In this paper, we suggest the NoC-based DNN platform as a new accelerator design paradigm. The NoC-based designs can reduce the off-chip memory accesses through a flexible interconnect that facilitates data exchange between processing elements on the chip. We first comprehensively investigate conventional platforms and methodologies used in DNN computing. Then we study and analyze different design parameters to implement the NoC-based DNN accelerator. The presented accelerator is based on mesh topology, neuron clustering, random mapping, and XY-routing. The experimental results on LeNet, MobileNet, and VGG-16 models show the benefits of the NoC-based DNN accelerator in reducing off-chip memory accesses and improving runtime computational flexibility. Kun-Chih Chen, Masoumeh Ebrahimi, Ting-Yi Wang, Yuch-Chi Yang |
NOCS | 2 |
| 2019 | Testing aware dynamic mapping for path-centric network-on-chip test
Shuyan Jiang, Junkai Zhan, Junshi Wang, Masoumeh Ebrahimi, Letian Huang |
Integr. | 6 |
| 2019 | Efficient Design-for-Test Approach for Networks-on-ChipabstractTo achieve high reliability in on-chip networks, it is necessary to test the network continuously with Built-in Self-Tests (BIST) so that the faults can be detected quickly and the number of affected packets can be minimized. However, BIST causes significant performance loss due to data dependencies. We propose EsyTest, a comprehensive test strategy with minimized influence on system performance. EsyTest tests the data path and the control path separately. The data path test starts periodically, but the actual test performs in the free time slots to avoid deactivating the router for testing. A reconfigurable router architecture and an adaptive fault-tolerant routing algorithm are proposed to guarantee the access to the processing core when the associated router is under test. During the whole test procedure of the network, all processing cores are accessible, and thus the system performance is maintained during the test. At the same time, EsyTest provides a full test coverage for the NoC and a better hardware compatibility comparing with the existing test strategies. Under the PARSEC benchmark and different test frequencies, the execution time increases less than 5 percent at the cost of 9.9 percent more area and 4.6 percent more power in comparison with the execution where no test procedure is applied. Junshi Wang, Masoumeh Ebrahimi, Letian Huang, Qiang Li 0021, Guangjun Li, Axel Jantsch |
IEEE Trans. Computers | 2 |
| 2018 | A lifetime-aware mapping algorithm to extend MTTF of Networks-on-ChipabstractFast aging of components has become one of the major concerns in Systems-on-Chip with further scaling of the submicron technology. This problem accelerates when combined with improper working conditions such as unbalanced components' utilization. Considering the mapping algorithms in the Networks-on-Chip domain, some routers/links might be frequently selected for mapping while others are underutilized. Consequently, the highly utilized components may age faster than others which results in disconnecting the related cores from the network. To address this issue, we propose a mapping algorithm, called lifetime-aware neighborhood allocation (LaNA), that takes the aging of components into account when mapping applications. The proposed method is able to balance the wear-out of NoC components, and thus extending the service time of NoC. We model the lifetime as a resource consumed over time and accordingly define the lifetime budget metric. LaNA selects a suitable node for mapping which has the maximum lifetime budget. Experimental results show that the lifetime-aware mapping algorithm could improve the minimal MTTF of NoC around 72.2%, 58.3%, 46.6% and 48.2% as compared to NN, CoNA, WeNA and CASqA, respectively. Letian Huang, Masoumeh Ebrahimi, Junshi Wang, Shuyan Jiang, Qiang Li 0021 |
ASP-DAC | 4 |
| 2018 | Optimizing dynamic mapping techniques for on-line NoC testabstractWith the aggressive scaling of submicron technology, intermittent faults are becoming one of the limiting factors in achieving a high reliability in Network-on-Chip (NoC). Increasing test frequency is necessary to detect intermittent faults, which in turn interrupts the execution of applications. On the other hand, the main goal of traditional mapping algorithms is to allocate applications to the NoC platform, ignoring about the test requirement. In this paper, we propose a novel testing-aware mapping algorithm (TAMA) for NoC, targeting intermittent faults on the paths between crossbars. In this approach, the idle links are identified and the components between two crossbars are tested when the application is mapped to the platform. The components can be tested if there is enough time from when the application leaves the platform and a new application enters it. The mapping algorithm is tuned to give a higher priority to the tested paths in the next application mapping. This leaves enough time to test the links and the belonging components that have not been tested in the expected time. Experiment results show that the proposed testing-aware mapping algorithm leads to a significant improvement over FF, NN, CoNA, and WeNA. Shuyan Jiang, Junshi Wang, Masoumeh Ebrahimi, Letian Huang, Qiang Li 0021 |
ASP-DAC | 5 |
| 2018 | First-Last: A Cost-Effective Adaptive Routing Solution for TSV-Based Three-Dimensional Networks-on-Chipabstract3D integration opens up new opportunities for future multiprocessor chips by enabling fast and highly scalable 3D Network-on-Chip (NoC) topologies. However, in an aim to reduce the cost of Through-silicon via (TSV), partially vertically connected NoCs, in which only a few vertical TSV links are available, have been gaining relevance. To reliably route packets under such conditions, we introduce a lightweight, efficient and highly resilient adaptive routing algorithm targeting partially vertically connected 3D-NoCs named First-Last. It requires a very low number of virtual channels (VCs) to achieve deadlock-freedom (2 VCs in the East and North directions and 1 VC in all other directions), and guarantees packet delivery as long as one healthy TSV connecting all layers is available anywhere in the network. An improved version of our algorithm, named Enhanced-First-Last is also introduced and shown to dramatically improve performance under low TSV availability while still using less virtual channels than state-of-the-art algorithms. A comprehensive evaluation of the cost and performance of our algorithms is performed to demonstrate their merits with respects to existing solutions. Amir Charif, Alexandre Coelho, Masoumeh Ebrahimi, Nader Bagherzadeh, Nacer-Eddine Zergainoh |
IEEE Trans. Computers | 3 |
| 2018 | LEAD: An Adaptive 3D-NoC Routing Algorithm with Queuing-Theory Based Analytical Verificationabstract2D-NoCs have been the mainstream approach used to interconnect multi-core systems. 3D-NoCs have emerged to compensate for deficiencies of 2D-NoCs such as long latency and power overhead. A low-latency routing algorithm for 3D-NoC is designed to accommodate high-speed communication between cores. Both simulation and analytical models are applied to estimate the communication latency of NoCs. Generally, simulations are time-consuming and slow down the design process. Analytical models provide, within a fraction of the time, nearly accurate results which can be used by simulation to fine-tune the design. In this paper, a high performance and adaptive routing algorithm has been proposed for partially connected 3D-NoCs. Latency of the routing algorithm under different traffic patterns, different number of elevators and different elevator assignment mechanisms are reported. An analytical model, tailored to the adaptivity of the algorithm and under low traffic scenarios, has been developed and the results have been verified by simulation. According to the results, simulation and analytical results are consistent within a 10 percent margin. Ronak Salamat, Misagh Khayambashi, Masoumeh Ebrahimi, Nader Bagherzadeh |
IEEE Trans. Computers | 3 |
| 2017 | A reliable weighted feature selection for auto medical diagnosisabstractFeature selection is a key step in data analysis. However, most of the existing feature selection techniques are serial and inefficient to be applied to massive data sets. We propose a feature selection method based on a multi-population weighted intelligent genetic algorithm to enhance the reliability of diagnoses in e-Health applications. The proposed approach, called PIGAS, utilizes a weighted intelligent genetic algorithm to select a proper subset of features that leads to a high classification accuracy. In addition, PIGAS takes advantage of multi-population implementation to further enhance accuracy. To evaluate the subsets of the selected features, the KNN classifier is utilized and assessed on UCI Arrhythmia dataset. To guarantee valid results, leave-one-out validation technique is employed. The experimental results show that the proposed approach outperforms other methods in terms of accuracy and efficiency. The results of the 16-class classification problem indicate an increase in the overall accuracy when using the optimal feature subset. Accuracy achieved being 99.70% indicating the potential of the algorithm to be utilized in a practical auto-diagnosis system. This accuracy was obtained using only half of features, as against an accuracy of66.76% using all the features. Golnaz Sahebi, Amin Majd, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen |
INDIN | 3 |
| 2017 | EbDa: A New Theory on Design and Verification of Deadlock-free Interconnection Networks
Masoumeh Ebrahimi, Masoud Daneshtalab |
ISCA | 1 |
| 2017 | Non-blocking BIST for continuous reliability monitoring of Networks-on-ChipabstractTo achieve high reliability in on-chip networks, frequent runs of Built-in Self-Test allow the detection of and recovery from faults before they affect packets and the system functionality. However, to test routers, wrappers isolate cores from the network which leads to execution blocking and performance loss. In this paper, we propose a design-for-test reconfigurable router with two alternative bypassing channels. The router architecture allows maintaining the connection between cores and the network during the testing procedure by utilizing the bypassing channels. With the help of an adaptive routing algorithm and a testing strategy, networks can be fully tested at a high testing frequency with <;15% increase of execution time. Junshi Wang, Letian Huang, Masoumeh Ebrahimi, Qiang Li 0021, Guangjun Li, Axel Jantsch |
ISCAS | 3 |
| 2016 | Non-Blocking Testing for Network-on-ChipabstractTo achieve high reliability in on-chip networks, it is necessary to test the network as frequently as possible to detect physical failures before they lead to system-level failures. A main obstacle is that the circuit under test has to be isolated, resulting in network cuts and packet blockage which limit the testing frequency. To address this issue, we propose a comprehensive network-level approach which could test multiple routers simultaneously at high speed without blocking or dropping packets. We first introduce a reconfigurable router architecture allowing the cores to keep their connections with the network while the routers are under test. A deadlock-free and highly adaptive routing algorithm is proposed to support reconfigurations for testing. In addition, a testing sequence is defined to allow testing multiple routers to avoid dropping of packets. A procedure is proposed to control the behavior of the affected packets during the transition of a router from the normal to the testing mode and vice versa. This approach neither interrupts the execution of applications nor has a significant impact on the execution time. Experiments with the PARSEC benchmarks on an 8x8 NoC-based chip multiprocessors show only 3 percent execution time increase with four routers simultaneously under test. Letian Huang, Junshi Wang, Masoumeh Ebrahimi, Masoud Daneshtalab, Xiaofan Zhang 0004, Guangjun Li, Axel Jantsch |
IEEE Trans. Computers | 3 |
| 2016 | A Resilient Routing Algorithm with Formal Reliability Analysis for Partially Connected 3D-NoCsabstract3D ICs can take advantage of a scalable communication platform, commonly referred to as the Networks-on-Chip (NoC). In the basic form of 3D-NoC, all routers are vertically connected. Partially connected 3D-NoC has emerged because of physical limitations of using vertical links. Routing is of great importance in such partially connected architectures. A high-performance, fault-tolerant and adaptive routing strategy with respect to the communication flow among the cores is crucial while freedom from livelock and deadlock has to be guaranteed. In this paper we introduce a new routing algorithm for partially connected 3D-NoCs. The routing algorithm is adaptive and tolerates the faults on vertical links as compared to the predesigned routing algorithms. Our results show a$40-50\%$improvement in the fraction of intact inter-level communications when the fault tolerant algorithm is used. This routing algorithm is light-weight and has only one virtual channel along the$Y$dimension. Ronak Salamat, Misagh Khayambashi, Masoumeh Ebrahimi, Nader Bagherzadeh |
IEEE Trans. Computers | 3 |
| 2015 | Dynamic Application Mapping Algorithm for Wireless Network-on-ChipabstractBecause of high bandwidth, low latency and flexible topology configurations provided by wireless NoC, this emerging technology is gaining momentum to be a promising future on-chip interconnection paradigm. However, congestion occurrence in wireless routers reduces the benefit of high speed wireless links and significantly increases the network latency, therefore, in this paper, a Dynamic Application Mapping Algorithm (DAMA) is introduced for wireless NoCs in order to reduce both internal and external congestion. DAMA has three key steps: finding the first node to map, choosing the first task to be mapped onto the first node, and allocation of the remaining tasks to the remaining nodes. Simulation results show significant gain in the mapping cost functions compared to state-of-the-art works. Amin Rezaei 0001, Masoud Daneshtalab, Danella Zhao, Farshad Safaei, Xiaohang Wang 0001, Masoumeh Ebrahimi |
PDP | 6 |
| 2015 | An Adaptive, Low Restrictive and Fault Resilient Routing Algorithm for 3D Network-on-ChipabstractThe cost and reliability issues of TSVs move 3D-NoCs toward heterogonous designs with limited number of TSVs. However, designing a deadlock-free routing algorithm for such heterogonous architectures is extremely challenging due to the increased possibilities of forming cycles between and within layers for 3D designs. In this paper, we target designing a routing algorithm for heterogeneous 3D-NoCs with the capability of working under the technical limit in which there is just one TSV in the network. This algorithm is light-weight and provides adaptivity by using only one virtual channel along the Y dimension. Ronak Salamat, Masoumeh Ebrahimi, Nader Bagherzadeh |
PDP | 2 |
| 2015 | A Routing-Level Solution for Fault Detection, Masking, and Tolerance in NoCsabstractFaults may occur in numerous locations of a router in a NoC platform. Compared with the faults in the data path, faults in the control path may cause more severe effects which may result in crashing the entire system. Most of the current efforts in literature focus on disabling a router when a fault is detected. Considering this level of coarse-granularity, the functioning parts of a router have to be unnecessarily disabled which may severely affect the performance or functionality of the on-chip network. To cope with this problem, in this paper we propose a mechanism to tolerate faults in the control path which largely avoid disabling a router as long as the fault is not severe. This mechanism is called DMT, standing for three distinguishing characteristics of the proposed method as fault Detection, fault Masking and fault Tolerance. The proposed mechanism can efficiently detect the faults expressed as illegal turns while it has the capability to tolerate faults without a prior knowledge on where and why a fault has happened. Xiaofan Zhang 0004, Masoumeh Ebrahimi, Letian Huang, Guangjun Li, Axel Jantsch |
PDP | 2 |
| 2015 | In-order delivery approach for 2D and 3D NoCs
Masoud Daneshtalab, Masoumeh Ebrahimi, Sergei Dytckov, Juha Plosila |
J. Supercomput. | 2 |
| 2014 | Parameterized AES-Based Crypto Processor for FPGAsabstractIn this paper, we propose a parameterized crypto co-processor based on Advanced Encryption Standard (AES). This parameterized AES module is combined with a 32-bit general purpose 5-stage pipelined MIPS processor. The AES module used in this paper is fully pipelined. The processor fetches an instruction from the instruction memory and sends it to the decode stage. If the instruction is the crypto instruction it is pushed into the AES module during the decode stage. However if the instruction belongs to the MIPS processor, the remaining cycles will be completed on the MIPS processor. The parameterized AES module has different latencies on different rounds of AES according to the application requirements. The effects of different number of rounds on latency, memory, and area are studied and reported. Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen, Sergei Dytckov, Giovanni Beltrame |
DSD | 3 |
| 2014 | Efficient STDP Micro-Architecture for Silicon Spiking Neural NetworksabstractSpiking neural networks (SNNs) are the closest approach to biological neurons in comparison with conventional artificial neural networks (ANN). SNNs are composed of neurons and synapses which are interconnected with a complex pattern. As communication in such massively parallel computational systems is getting critical, the network-on-chip (NoC) becomes a promising solution to provide a scalable and robust interconnection fabric. However, using NoC for large-scale SNNs arises a trade-off between scalability, throughput, neuron/router ratio (cluster size), and area overhead. In this paper, we tackle the trade-off using a clustering approach and try to optimize the synaptic resource utilization. An optimal cluster size can provide the lowest area overhead and power consumption. For the learning purposes, a phenomenon known as spike-timing-dependent plasticity (STDP) is utilized. The micro-architectures of the network, clusters, and the computational neurons are also described. The presented approach suggests a promising solution of integrating NoCs and STDP-based SNNs for the optimal performance based on the underlying application. Sergei Dytckov, Masoud Daneshtalab, Masoumeh Ebrahimi, Hassan Anwar, Juha Plosila, Hannu Tenhunen |
DSD | 3 |
| 2014 | Integration of AES on Heterogeneous Many-Core SystemabstractIncreasing in the transistor density in a single chip makes it possible for many-core systems to utilize design space for implementing complex embedded systems. In this paper, we propose an architecture for heterogeneous many-core system to integrate block cipher algorithm which is based on Advanced Encryption Standard (AES). In order to implement AES as a crypto-core along with heterogeneous many-core system two different approaches are proposed. In the first approach, the platform is composed of an AES module, a crypto-core, a network interface, and an internal memory which are managed through a controller. In the second approach, the platform is composed of a Direct Memory Access (DMA), network interface, an internal memory, and a microprocessor in which the AES module is integrated as a crypto-core. Both approaches have been analyzed and compared in terms of area overhead and performance. Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Marco Ramírez 0001, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2014 | Path-Based Partitioning Methods for 3D Networks-on-Chip with Minimal Adaptive RoutingabstractCombining the benefits of 3D ICs and Networks-on-Chip (NoCs) schemes provides a significant performance gain in Chip Multiprocessors (CMPs) architectures. As multicast communication is commonly used in cache coherence protocols for CMPs and in various parallel applications, the performance of these systems can be significantly improved if multicast operations are supported at the hardware level. In this paper, we present several partitioning methods for the path-based multicast approach in 3D mesh-based NoCs, each with different levels of efficiency. In addition, we develop novel analytical models for unicast and multicast traffic to explore the efficiency of each approach. In order to distribute the unicast and multicast traffic more efficiently over the network, we propose the Minimal and Adaptive Routing (MAR) algorithm for the presented partitioning methods. The analytical and experimental results show that an advantageous method named Recursive Partitioning (RP) outperforms the other approaches. RP recursively partitions the network until all partitions contain a comparable number of switches and thus the multicast traffic is equally distributed among several subsets and the network latency is considerably decreased. The simulation results reveal that the RP method can achieve performance improvement across all workloads while performance can be further improved by utilizing the MAR algorithm. Nineteen percent average and 42 percent maximum latency reduction are obtained on SPLASH-2 and PARSEC benchmarks running on a 64-core CMP. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, José Flich, Hannu Tenhunen |
IEEE Trans. Computers | 1 |
| 2014 | Adaptive load balancing in learning-based approaches for many-core embedded systems
Fahimeh Farahnakian, Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila |
J. Supercomput. | 2 |
| 2013 | MD: Minimal path-based fault-tolerant routing in on-Chip NetworksabstractThe communication requirements of many-core embedded systems are convened by the emerging Network-on-Chip (NoC) paradigm. As on-chip communication reliability is a crucial factor in many-core systems, the NoC paradigm should address the reliability issues. Using fault-tolerant routing algorithms to reroute packets around faulty regions will increase the packet latency and create congestion around the faulty region. On the other hand, the performance of NoC is highly affected by the network congestion. Congestion in the network can increase the delay of packets to route from a source to a destination, so it should be avoided. In this paper, a minimal and defect-resilient (MD) routing algorithm is proposed in order to route packets adaptively through the shortest paths in the presence of a faulty link, as long as a path exists. To avoid congestion, output channels can be adaptively chosen whenever the distance from the current to destination node is greater than one hop along both directions. In addition, an analytical model is presented to evaluate MD for two-faulty cases. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Farhad Mehdipour |
ASP-DAC | 1 |
| 2013 | CARS: congestion-aware request scheduler for network interfaces in NoC-based manycore systemsabstractNetwork congestion is a critical issue of memory parallelism in network-based manycore systems where multiple memories can be accessed simultaneously. Therefore, a congestion-aware method is necessitated to deal with the network congestion. In this paper, we present a streamlined method in order to reduce the network congestion. The idea is to use the global congestion information as a metric in network interfaces to reduce the congestion level of highly congested areas. Network interfaces connected to memory modules are equipped with an adaptive scheduler using the global congestion information to reduce additional traffic to congested areas. Experimental results with synthetic test cases demonstrate that the on-chip network utilizing the proposed adaptive scheduler presents up to 23% improvement in average latency. Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen |
DATE | 2 |
| 2013 | Fault-tolerant routing algorithm for 3D NoC using Hamiltonian path strategyabstractWhile Networks-on-Chip (NoC) have been increasing in popularity with industry and academia, it is threatened by the decreasing reliability of aggressively scaled transistors. In this paper, we address the problem of faulty elements by the means of routing algorithms. Commonly, fault-tolerant algorithms are complex due to supporting different fault models while preventing deadlock. When moving from 2D to 3D network, the complexity increases significantly due to the possibility of creating cycles within and between layers. In this paper, we take advantages of the Hamiltonian path to tolerate faults in the network. The presented approach is not only very simple but also able to support almost all one-faulty unidirectional links in 2D and 3D NoCs. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila |
DATE | 1 |
| 2013 | Minimal-path fault-tolerant approach using connection-retaining structure in Networks-on-ChipabstractThere are many fault-tolerant approaches presented both in off-chip and on-chip networks. Regardless of all varieties, there has always been a common assumption between them. Most of all known fault-tolerant methods are based on rerouting packets around faults. Rerouting might take place through nonminimal paths which affect the performance significantly not only by taking longer paths but also by creating hotspot around a fault. In this paper, we present a fault-tolerant approach based on using the shortest paths. This method maintains the performance of Networks-on-Chip in the presence of faults. To avoid using non-minimal paths, the router architecture is slightly modified. In the new form of architecture, there is an ability to connect the horizontal and orthogonal links of a faulty router such that healthy routers are kept connected to each other. Based on this architecture, a fault-tolerant routing algorithm is presented which is obviously much simpler than traditional fault-tolerant routing algorithms. According to this algorithm, only the shortest paths are used by packets in the presence of fault. This results retains the performance of NoCs in faulty situations. This algorithm is highly reliable, for an instance, the reliability is more than 99.5% when there are six faulty routers in an 8×8 mesh network. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen |
NOCS | 1 |
| 2013 | DyXYZ: Fully Adaptive Routing Algorithm for 3D NoCsabstractTraditional methods in 3D NoCs simply use a deterministic routing algorithm to deliver packets from a source to a destination node. However, deterministic methods are unable to distribute the traffic load over the network, which results in degrading the performance. In this paper, we present a fully adaptive routing algorithm for 3D NoCs, named DyXYZ. In DyXYZ, the congestion information at the input buffer of the neighboring routers is used as congestion metric to select among the output channels. This algorithm is proven to be deadlock free by using 4, 4, and 2 virtual channels along the X, Y, and Z dimensions, respectively. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen |
PDP | 1 |
| 2013 | High Performance Fault-Tolerant Routing Algorithm for NoC-Based Many-Core SystemsabstractNetworks-on-Chip (NoCs) has become a promising approach for the on-chip communication infrastructure of many-core Systems-on-Chip (SoCs). Faults may occur in the NoC both at the router and link level. There are many fault-tolerant approaches presented both in the off-chip and on-chip networks. Some approaches disable some healthy components in order to form a specific shape and others not. Regardless of all varieties, there has always been a common assumption among them. Most of all traditional fault-tolerant methods are based on rerouting packets around a faulty node or region. These approaches affect the performance significantly not only by taking longer paths but also by creating hotspot around a fault. The focus of this paper is to maintain the performance of NoC in the presence of faults. The presented method takes advantage of a fully adaptive routing algorithm using one and two virtual channels along the X and Y dimensions. This method is able to tolerate all cases of one-faulty node without losing the performance of NoC. According to the experimental results, this presented fault-tolerant routing algorithm is able to support up to six faulty nodes in the 8×8 mesh network by up to 98% reliability. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila |
PDP | 1 |
| 2013 | Cluster-based topologies for 3D Networks-on-Chip using advanced inter-layer bus architecture
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Comput. Syst. Sci. | 1 |
| 2013 | A systematic reordering mechanism for on-chip networks using efficient congestion-aware method
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Syst. Archit. | 2 |
| 2013 | Fuzzy-based Adaptive Routing Algorithm for Networks-on-Chip
Masoumeh Ebrahimi, Hannu Tenhunen, Masoud Dehyadegari |
J. Syst. Archit. | 1 |
| 2012 | CATRA- congestion aware trapezoid-based routing algorithm for on-chip networksabstractCongestion occurs frequently in Networks-on-Chip when the packets demands exceed the capacity of network resources. Congestion-aware routing algorithms can greatly improve the network performance by balancing the traffic load in adaptive routing. Commonly, these algorithms either rely on purely local congestion information or take into account the congestion conditions of several nodes even though their statuses might be out-dated for the source node, because of dynamically changing congestion conditions. In this paper, we propose a method to utilize both local and non-local network information to determine the optimal path to forward a packet. The non-local information is gathered from the nodes that not only are more likely to be chosen as intermediate nodes in the routing path but also provide up-to-date information to a given node. Moreover, to collect and deliver the non-local information, a distributed propagation system is presented. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DATE | 1 |
| 2012 | MAFA: Adaptive Fault-Tolerant Routing Algorithm for Networks-on-ChipabstractWhile Networks-on-Chip have been increasing in popularity with industry and academia, it is threatened by the decreasing reliability of aggressively scaled transistors. This level of failure has architectural level ramifications, as it may cause an entire on-chip network to fail. Traditional fault-tolerant routing algorithms can overcome the faulty links or routers by rerouting packets around faulty regions. These approaches increase the packet latency and create congestion around the faulty region. In this paper, we present a novel fault-tolerant method that is able to route packets through shortest paths in the presence of faulty links, as long as a path exists. Although the same idea can be applied to a network with any number of virtual channels, we utilize two virtual channels to tolerate all one and two faulty links. Finally, the method is extended to support multiple faulty links by fully utilizing all allowable turns in the network. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen |
DSD | 1 |
| 2012 | HARAQ: Congestion-Aware Learning Model for Highly Adaptive Routing Algorithm in On-Chip NetworksabstractThe occurrence of congestion in on-chip networks can severely degrade the performance due to increased message latency. In mesh topology, minimal methods can propagate messages over two directions at each switch. When shortest paths are congested, sending more messages through them can deteriorate the congestion condition considerably. In this paper, we present an adaptive routing algorithm for on-chip networks that provide a wide range of alternative paths between each pair of source and destination switches. Initially, the algorithm determines all permitted turns in the network including 180-degree turns on a single channel without creating cycles. The implementation of the algorithm provides the best usage of all allowable turns to route messages more adaptively in the network. On top of that, for selecting a less congested path, an optimized and scalable learning method is utilized. The learning method is based on local and global congestion information and can estimate the latency from each output channel to the destination region. Masoumeh Ebrahimi, Masoud Daneshtalab, Fahimeh Farahnakian, Juha Plosila, Pasi Liljeberg, Maurizio Palesi, Hannu Tenhunen |
NOCS | 1 |
| 2012 | LEAR - A Low-Weight and Highly Adaptive Routing Method for Distributing Congestions in On-chip NetworksabstractCongestion-aware routing algorithms can improve network throughput by avoiding packets to be routed through congested areas. In this paper, we propose a minimal/non-minimal routing algorithm to alleviate congestion in the network by making use of all available paths between sources and destinations. The simplicity of the proposed algorithm provides a cost and power efficient solution for Networks-on-Chip while the high degree of adaptive ness, achieved by using an additional virtual channel along the Y dimension, leads to an increased performance. In this method, different restrictions are imposed on the use of each virtual channel, so that the prohibited turns in one virtual channel are permitted in the other one. By fully exploiting of the eligible turns in the network, a large number of output channels can be provided by the proposed method. Based on this method, a packet is routed along the non-minimal path when the neighboring routers in the minimal path are congested. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 1 |
| 2012 | Memory-Efficient On-Chip Network With Adaptive InterfacesabstractTo achieve higher memory bandwidth in network-based multiprocessor architectures, multiple dynamic random access memories can be accessed simultaneously. In such architectures, not only resource utilization and latency are the critical issues but also a reordering mechanism is required to deliver the response transactions of concurrent memory accesses in-order. In this paper, we present a memory-efficient on-chip network architecture to cope with these issues efficiently. Each node of the network is equipped with a novel network interface (NI) to deal with out-of-order delivery, and a priority-based router to decrease the network latency. The proposed NI exploits a streamlined reordering mechanism to handle the in-order delivery and utilizes the advance extensible interface transaction-based protocol to maintain compatibility with existing intellectual property cores. To improve the memory utilization and reduce the memory latency, an optimized memory controller is integrated in the presented NI. Experimental results with synthetic test cases demonstrate that the proposed on-chip network architecture provides significant improvements in average network latency (16%), average memory access latency (19%), and average memory utilization (22%). Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Exploring partitioning methods for 3D Networks-on-Chip utilizing adaptive routing modelabstractThree-Dimensional (3D) integration is a solution to the interconnect bottleneck in Two-Dimensional (2D) MultiProcessor System on Chip (MPSoC). 3D IC design improves performance and decreases power consumption by replacing long horizontal interconnects with shorter vertical ones. As the multicast communication is utilized commonly in various parallel applications, the performance can be significantly improved by supporting of multicast operations at the hardware level. In this paper, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method named Recursive Partitioning (RP) in which the network is recursively partitioned until all partitions contain comparable number of nodes. By this approach, the multicast traffic is distributed among several subsets and the network latency is considerably decreased. We also present Minimal Adaptive Routing (MAR) algorithm for the unicast and multicast traffic in 3D-mesh Networks-on-Chip (NoCs). The idea behind the MAR algorithm is utilizing the Hamiltonian path to provide a set of alternative paths. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 1 |
| 2011 | Agent-based on-chip network using efficient selection methodabstractCongestion in on-chip networks may cause many drawbacks in multiprocessor systems including throughput reduction, increase in latency, and additional power consumption. Furthermore, conventional congestion control methods, employed for on-chip networks, cannot efficiently collect congestion information and distribute them over the on-chip network. In this paper, we present a novel structure for on-chip networks, named Agent-based Network-on-Chip (ANoC), to diagnose the congested areas. In addition to the presented structure, an efficient Congestion-Aware Selection (CAS) method is proposed to reduce overall network latency. CAS is capable of selecting an appropriate output channel to route packets along a less congested path. 29% average and 35% maximum latency reduction are achieved on SPLASH-2 and PARSEC benchmarks running on a 36-core Chip Multi-Processor. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
VLSI-SoC | 1 |
| 2011 | A generic adaptive path-based routing method for MPSoCs
Masoud Daneshtalab, Masoumeh Ebrahimi, Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen |
J. Syst. Archit. | 2 |
| 2010 | Partitioning methods for unicast/multicast traffic in 3D NoC architectureabstractAs the scale of integration grows, the interconnection problem becomes one of the major design considerations of Multi Processor System on Chip (MPSoC). In recent years, many researchers have conducted studies on 3D IC designs stacking multiple layers on top of each other. In order to decrease the transmission delay of unicast/multicast messages in a network based multicore system, the network is divided into several partitions. In this paper, we first introduce a novel idea of balanced partitioning that allows the network to be partitioned effectively. Then, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method based on the idea of balanced partitioning to provide a high degree of parallelism with a considerable reduction of packet delay in unicast/multicast traffic. Simulations are provided to evaluate and compare the performance of proposed methods. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
DDECS | 1 |
| 2010 | A Low-Latency and Memory-Efficient On-chip NetworkabstractUsing multiple SDRAMs in MPSoCs and NoCs to increase memory parallelism is very common nowadays. In-order delivery, resource utilization, and latency are the most critical issues in such architectures. In this paper, we present a novel network interface architecture to cope with these issues efficiently. The proposed network interface exploits a resourceful reordering mechanism to handle the in-order delivery and to increase the resource utilization. A brilliant memory controller is efficiently integrated into this network interface to improve the memory utilization and reduce both memory and network latencies. In addition, to bring compatibility with existing IP cores the proposed network interface utilizes AXI transaction based protocol. Experimental results with synthetic test cases demonstrate that the proposed architecture gives significant improvements in average network latency (12%), average memory access latency (19%), and average memory utilization (22%). Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 2 |
| 2010 | A High-Performance Network Interface Architecture for NoCs Using Reorder Buffer SharingabstractIncreasing memory parallelism in MPSoCs to provide higher memory bandwidth is achieved by accessing multiple memories simultaneously. Inasmuch as the response transactions of concurrent memory accesses must be in-order, a reordering mechanism is required. To our knowledge the resource utilization of conventional reordering mechanisms is low. In this paper, we present a novel network interface architecture for on-chip networks to increase the resource utilization and to improve overall performance. Also, based on the proposed architecture, a hybrid network interface is presented to integrate both memory and processor in a tile. The proposed architecture exploits AXI transaction based protocol to be compatible with existing IP cores. Experimental results with synthetic test cases demonstrate that the proposed architecture outperforms the conventional architecture in terms of latency. Also, the cost of the presented architecture is evaluated with UMC 0.09 ¿ m technology. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 1 |
| 2010 | HAMUM - A Novel Routing Protocol for Unicast and Multicast Traffic in MPSoCsabstractMany parallel applications in MPSoCs take advantage of multicast communication. Several multicast schemes such as path-based, tree-based, and unicast-based have been proposed in interconnection networks. Path-based multicast scheme has been proven to be more efficient than the other schemes in on-chip interconnection network. A new adaptive routing model based on Hamiltonian path for both the multicast and unicast traffics, called Hamiltonian Adaptive Multicast and Unicast Model (HAMUM), is presented. Results obtained in both multicast and mixed traffic models show that the proposed adaptive algorithm for multicast aspect has lower latency and power dissipation compared to previously proposed path-based multicasting algorithms with less than 0.5% hardware overhead. Additionally, for the unicast aspect the proposed adaptive model outperforms the other unicast turn models. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
PDP | 1 |
| 2009 | An efficent dynamic multicast routing protocol for distributing traffic in NOCsabstractNowadays, in MPSoCs and NoCs, multicast protocol is significantly used for many parallel applications such as cache coherency in distributed shared-memory architectures, clock synchronization, replication, or barrier synchronization. Among several multicast schemes proposed in on chip interconnection networks, path-based multicast scheme has been proven to be more efficient than the tree-based, and unicast-based. In this paper a low distance path-based multicast scheme is proposed. The proposed method takes advantage of the network partitioning, and utilizing of an efficient destination ordering algorithm. The results in performance, and power consumption show that the proposed method outstands the previous on chip path-based multicasting algorithms. Masoumeh Ebrahimi, Masoud Daneshtalab, Mohammad Hossein Neishaburi, Siamak Mohammadi, Ali Afzali-Kusha, Juha Plosila, Hannu Tenhunen |
DATE | 1 |
| 2009 | An Adaptive Unicast/Multicast Routing Algorithm for MPSoCsabstractSeveral parallel applications in MPSoCs take advantage of multicast communication. Path-based multicast scheme has been proven to be more efficient than the others multicast schemes in on-chip interconnection network. We present a new adaptive path based model for both the multicast and unicast wormhole routing protocols. The proposed model under mixed traffic models has lower latency than the previous path-based methods with negligible hardware overhead. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
DSD | 1 |