VLDB 2026 Research / reviewers in the wild / expert
Masoud Daneshtalab
dblp:53/771
· DBLP profile ↗
130ranked-venue papers
13as first author
39since 2021 · last 2026
0000-0001-6289-1521ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 84 · 11 first-author · 21 since 2021Artificial intelligence and machine learning · 13 · 9 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Theory of computation · 2Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepFedNAS: Pareto Optimal Supernet Training for Improved and Predictor-Free Federated Neural Architecture SearchabstractFederated Neural Architecture Search (FedNAS) is hindered by unguided supernet training and costly post-training search pipelines.We introduce DeepFedNAS, a two-phase framework that resolves these issues.We propose Federated Pareto Optimal Supernet Training, using a pre-computed path of elite architectures as an intelligent curriculum to train a superior supernet.Subsequently, our Predictor-Free Search Method uses a principled fitness function as a zero-cost proxy for accuracy, finding optimal subnets in seconds.DeepFedNAS achieves state-of-the-art accuracy, superior robustness to data heterogeneity, and a ∼ 61× search pipeline speedup, making FedNAS practical and efficient. Bostan Khan, Masoud Daneshtalab |
ESANN | 2 |
| 2026 | Physics-Informed Recurrent Architecture with Embedded Thermodynamic Dynamics for Robust Sequence ModelingabstractPhysics-informed machine learning has shown strong potential in improving generalisation under limited or noisy data, but most existing approaches treat physical priors only as soft regularisation terms on the loss.This work introduces a physics-structured recurrent architecture where thermodynamic differential equations are embedded directly into LSTM state updates.Adaptive physical parameters are learned through auxiliary multilayer perceptrons, forming a differentiable hybrid dynamical system that fuses physics priors with sequence learning.Experiments on industrial datasets show improved robustness under unseen fault conditions, outperforming conventional LSTMs and PINN-style models.The framework offers a scalable and generalizable approach to physics-aware recurrent modeling. Zafer Yigit, Håkan Forsberg, Masoud Daneshtalab |
ESANN | 3 |
| 2026 | Correcting and Quantifying Systematic Errors in 3D Box Annotations for Autonomous DrivingabstractAccurate ground truth annotations are critical to supervised learning and evaluating the performance of autonomous vehicle systems. These vehicles are typically equipped with active sensors, such as LiDAR, which scan the environment in predefined patterns. 3D box annotation based on data from such sensors is challenging in dynamic scenarios, where objects are observed at different timestamps, hence different positions. Without proper handling of this phenomenon, systematic errors are prone to being introduced in the box annotations. Our work is the first to discover such annotation errors in widely used, publicly available datasets. Through our novel offline estimation method, we correct the annotations so that they follow physically feasible trajectories and achieve spatial and temporal consistency with the sensor data. For the first time, we define metrics for this problem; and we evaluate our method on the Argoverse 2, MAN TruckScenes, and our proprietary datasets. Our approach increases the quality of box annotations by more than 17% in these datasets. Furthermore, we quantify the annotation errors in them and find that the original annotations are misplaced by up to 2.5 m, with highly dynamic objects being the most affected. Finally, we test the impact of the errors in benchmarking and find that the impact is larger than the improvements that state-of-the-art methods typically achieve w.r.t. the previous state-of-the-art methods; showing that accurate annotations are essential for correct interpretation of performance. Our code is available at https://github.com/alexandre-justo-miro/annotation-correction-3D-boxes. Alexandre Justo Miro, Ludvig af Klinteberg, Bogdan Timus, Aron Asefaw, Ajinkya Khoche, Thomas Gustafsson, Sina Sharif Mansouri, Masoud Daneshtalab |
WACV | 8 |
| 2026 | Frequency domain complex-valued convolutional neural networkabstract• Formulation of fully complex-valued building blocks for complex-valued CNNs, ensuring consistent operation in the frequency domain. • Lightweight, computationally efficient, fully complex-valued residual CNN operating on complex data in the frequency domain. • Novel log-magnitude activation function that maintains phase information while introducing effective non-linearity, along with a detailed comparative analysis with the complex ReLU variant and the Cardioid activation function. • We have shown that complex-valued CNNs can outperform real-valued CNNs, enhancing performance and generalization while reducing computational demands. Complex-valued convolutional neural networks have demonstrated promising results in reducing space, time, and computational complexity compared to real-valued models, particularly in signal and image processing. Despite their strong representational capacity and theoretical benefits, complex-valued CNNs remain limited due to theabsence of simplified theoretical and practical formulations for fully complex-valued building blocks. Existing studies often depend on fast Fourier transforms (FFT/IFFT) for domain transitions between layers due to the lack of well-established complex-valued activation functions or filter parameters initialization. Additionally, many earlier works adapt complex versions of the real-valued activation functions in a split-type manner, which might distort phase information and weaken generalization. To overcome these challenges, we propose a lightweight fully complex-valued residual CNN that operates entirely on complex data in the frequency domain. Our design simplifies fully complex building blocks and introduces a Log-Magnitude activation function that preserves phase information, outperforming traditional complex ReLU variants and the Cardioid activation function. Experimental validation across diverse multi-modal datasets, including MNIST, SVHN, MIT-BIH Arrhythmia, PTB Diagnostic ECG, DIAT- μ RadHAR , and DIAT- μ SAT , demonstrates the superior performance of our fully complex-valued CNNs over real-valued models. Mainak Chakraborty, Masood Aryapoor, Masoud Daneshtalab |
Expert Syst. Appl. | 3 |
| 2026 | Adaptive and efficient federated distillation with selective homomorphic encryption for edge AI
Dadmehr Rahbari, Masoud Daneshtalab, Maksim Jenihhin |
Expert Syst. Appl. | 2 |
| 2025 | Bridging Quantization and Deployment: A Fixed-Point Workflow for FPGA AcceleratorsabstractDeploying deep learning models on resource-constrained hardware like Field-Programmable Gate Arrays (FPGAs) remains challenging despite advancements in quantization techniques, which often fail to map optimally to target hardware. This study proposes an end-to-end workflow for fixed-point quantization targeting FPGA accelerators, integrating hardware emulation within quantization-aware training to bridge software-hardware co-design. This ensures quantized models are optimized for real-world deployment. Evaluations on CIFAR-10 and ImageNet datasets using ResNet and VGG models show competitive performance, with up to 2% improvement in accuracy over state-of-the-art methods. This work provides insights to resolve the discrepancies that often occur between software-based quantization and hardware deployment. Our methodology effectively bridges quantization and deployment, providing a practical solution for edge device applications. Obed M. Mogaka, Håkan Forsberg, Masoud Daneshtalab |
DDECS | 3 |
| 2025 | ProARD: Progressive Adversarial Robustness Distillation: Provide Wide Range of Robust StudentsabstractAdversarial Robustness Distillation (ARD) has emerged as an effective method to enhance the robustness of lightweight deep neural networks against adversarial attacks. Current ARD approaches have leveraged a large robust teacher network to train one robust lightweight student. However, due to the diverse range of edge devices and resource constraints, current approaches require training a new student network from scratch to meet specific constraints, leading to substantial computational costs and increased CO2emissions.This paper proposes Progressive Adversarial Robustness Distillation (ProARD), enabling the efficient one-time training of a dynamic network that supports a diverse range of accurate and robust student networks without requiring retraining. We first make a dynamic deep neural network based on dynamic layers by encompassing variations in width, depth, and expansion in each design stage to support a wide range of architectures (> 1019). Then, we consider the student network with the largest size as the dynamic teacher network. ProARD trains this dynamic network using a weight-sharing mechanism to jointly optimize the dynamic teacher network and its internal student networks. However, due to the high computational cost of calculating exact gradients for all the students within the dynamic network, a sampling mechanism is required to select a subset of students. We show that random student sampling in each iteration fails to produce accurate and robust students. ProARD employs a progressive sampling strategy that gradually reduces the size of student networks in three steps during training while applying robustness distillation between the dynamic teacher network and the selected students. Finally, we leverage a multi-objective evolutionary algorithm based on a proposed accuracy-robustness predictor to identify optimal architectures that balance accuracy, robustness, and efficiency.Through the experiments, we show that ProARD reduces the computational cost by 60× and improves accuracy and robustness by 13% and 14%, respectively, compared to random sampling. We also demonstrate that our accuracy-robustness predictor can estimate the accuracy and robustness of test student networks with root mean squared errors of 0.0073 and 0.0072, respectively. Seyedhamidreza Mousavi, Seyed Ali Mousavi, Masoud Daneshtalab |
IJCNN | 3 |
| 2025 | Experimental Evaluation of a CAN-to-TSN Gateway ImplementationabstractThe increasing complexity of modern embedded systems highlights the limitations of Controller Area Network (CAN) in terms of transmission speed and scalability. The IEEE Time-Sensitive Networking (TSN) task group developed a set of standards to enhance switched Ethernet with high bandwidth, low jitter, and deterministic communication. Despite these advances, CAN will likely co-exist with TSN in, e.g., the automotive industry due to factors such as cost-effectiveness and legacy of CAN. This paper presents an experimental evaluation of a CAN-toTSN gateway implementation, focusing on the impact of different forwarding and scheduling strategies on network performance. We analyze various queuing techniques and scheduling mechanisms in a realistic experimental setup and assess their impact on end-to-end delay and TSN bandwidth utilization. The evaluation results demonstrate that encapsulating only a single CAN frame within a TSN frame effectively minimizes the end-to-end delay of CAN frames, in particular when a high-speed TSN network is used. Furthermore, we perform a comparative evaluation of the Time-Aware Shaper (TAS) and Weighted Round Robin (WRR) mechanisms in the TSN network. Interestingly, WRR leads to lower delays for CAN frames in the TSN network compared to TAS, which we attribute to the lack of synchronization between CAN and TSN. Aldin Berisa, Benjamin Kraljusic, Nejla Zahirovic, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
ISORC | 5 |
| 2025 | Machine Learning-Based Prognostic Approaches for Construction Equipment Powertrain SystemsabstractConstruction equipment has important roles in industries such as construction and mining. Any downtime because of failures increase cost. Traditional diagnostic systems detect failures only after they occur, making it difficult to take precautions and prolonging repair times. This paper is the first to address the analysis of machine learning-powered Prognostic and Health Management (PHM) systems specifically for predicting failures in diesel engine air intake systems, focusing on two common issues: air leakage and Exhaust Gas Recirculation (EGR) blockage. This study compares various machine learning and deep learning models for anomaly detection and fault classification using real-world sensor data from controlled engine tests. The results demonstrate that ensemble and neural network-based machine learning methods, such as Random Forest, XGBoost, and LSTM, achieve highly successful predictions for anomaly detection and fault classification. Zafer Yigit, Håkan Forsberg, Masoud Daneshtalab |
IV | 3 |
| 2025 | FMCW Radar-Based Human Activity Recognition: A Machine Learning Approach for Elderly CareabstractIn this paper, we propose a novel system prototype for human activity recognition using a low-cost, low-power millimeter-wave (mmWave) frequency-modulated continuous wave (FMCW) radar. Our approach applies the Fast Fourier Transform on the slow time axis and employs a Capon filter to generate range-Doppler, range-azimuth, and range-elevation maps, respectively. It can also effectively mitigate noise and multipath effects. We then use principal component analysis for feature reduction, reducing the dimensionality of the feature vectors extracted from these maps, which can be used to train conventional machine learning classifiers. This approach aims to achieve a balance between computational complexity, accuracy, and overall system performance. Our proposed system demon-strates promising recognition rates and robustness across varying levels of activity granularity, achieving recognition rates from 90.28% for four activities up to 70.97% for seven fine-grained ac-tivities. These findings highlight the potential of millimeter wave radar and suggested range maps combined with conventional machine learning classifiers for noninvasive, privacy-preserving activity recognition, with significant implications for healthcare, elderly care, and ambient assisted living. Mohammadreza Mashhadigholamali, Ali Samimi Fard, Samaneh Zolfaghari, Hajar Abedi, Mainak Chakraborty, Luigi Borzì, Masoud Daneshtalab, George Shaker |
WCNC | 7 |
| 2024 | FORTUNE: A Negative Memory Overhead Hardware-Agnostic Fault TOleRance TechniqUe in DNNsabstractThis paper presents FORTUNE, a hardware-agnostic fault tolerance technique for DNNs that leverages quantization to enhance reliability without significant performance overhead. Unlike conventional methods like Triple Modular Redundancy (TMR), which are computationally expensive, the proposed approach uses memory savings from quantization to protect the critical Most Significant Bit, improving fault tolerance in Deep Neural Networks (DNNs). Memory utilization has been reduced by 37.5% across all networks, with vulnerability in AlexNet reduced by 56% compared to the 8-bit version and 84% compared to the unprotected 3-bit version. These improvements come with only a minor increase in execution time of less than 3%. Using AlexNet as an example demonstrates how our approach effectively enhances memory utilization and resilience while causing only a minimal increase in execution time. Samira Nazari, Mahdi Taheri, Ali Azarpeyvand, Mohsen Afsharchi, Tara Ghasempouri, Christian Herglotz, Masoud Daneshtalab, Maksim Jenihhin |
ATS | 7 |
| 2024 | Autonomous Realization of Safety- and Time-Critical Embedded Artificial IntelligenceabstractThere is an evident need to complement embedded critical control logic with AI inference, but today's AI-capable hardware, software, and processes are primarily targeted towards the needs of cloud-centric actors. Telecom and defense airspace industries, which make heavy use of specialized hardware, face the challenge of manually hand-tuning AI workloads and hardware, presenting an unprecedented cost and complexity due to the diversity and sheer number of deployed instances. Furthermore, embedded AI functionality must not adversely affect real-time and safety requirements of the critical business logic. To address this, end-to-end AI pipelines for critical platforms are needed to automate the adaption of networks to fit into resource-constrained devices under critical and real-time constraints, while remaining interoperable with de-facto standard AI tools and frameworks used in the cloud. We present two industrial applications where such solutions are needed to bring AI to critical and resource-constrained hardware, and a generalized end-to-end AI pipeline that addresses these needs. Crucial steps to realize it are taken in the industry-academia collaborative FASTER-AI project. Joakim Lindén, Andreas Ermedahl, Hans Salomonsson, Masoud Daneshtalab, Björn Forsberg, Paris Carbone |
DATE | 4 |
| 2024 | SAFFIRA: a Framework for Assessing the Reliability of Systolic-Array-Based DNN AcceleratorsabstractSystolic array has emerged as a prominent archi-tecture for Deep Neural Network (DNN) hardware accelerators, providing high-throughput and low-latency performance essen-tial for deploying DNNs across diverse applications. However, when used in safety-critical applications, reliability assessment is mandatory to guarantee the correct behavior of DNN accelerators. While fault injection stands out as a well-established practical and robust method for reliability assessment, it is still a very time-consuming process. This paper addresses the time efficiency issue by introducing a novel hierarchical software-based hardware-aware fault injection strategy tailored for systolic array-based DNN accelerators. The uniform Recurrent Equations system is used for software modeling of the systolic-array core of the DNN accelerators. The approach demonstrates a reduction of the fault injection time up to 3 × compared to the state-of-the-art hybrid (software/hardware) hardware-aware fault injection frameworks and more than 2000 × compared to RT-level fault injection frameworks - without compromising accuracy. Additionally, we propose and evaluate a new reliability metric through experimental assessment. The performance of the framework is studied on state-of-the-art DNN benchmarks. Mahdi Taheri, Masoud Daneshtalab, Jaan Raik, Maksim Jenihhin, Salvatore Pappalardo, Paul Jiménez, Bastien Deveautour, Alberto Bosio |
DDECS | 2 |
| 2024 | AdAM: Adaptive Fault-Tolerant Approximate Multiplier for Edge DNN AcceleratorsabstractMultiplication is the most resource-hungry operation in the neural network’s processing elements. In this paper, we propose an architecture of a novel adaptive fault-tolerant approximate multiplier tailored for ASIC-based DNN accelerators. AdAM employs an adaptive adder relying on an unconventional use of the leading one position value of the inputs for fault detection through the optimization of unutilized adder resources. The proposed architecture uses a lightweight fault mitigation technique that sets the detected faulty bits to zero. The hardware resource utilization and the DNN accelerator’s reliability metrics are used to compare the proposed solution against the triple modular redundancy (TMR) in multiplication, unprotected exact multiplication, and unprotected approximate multiplication. It is demonstrated that the proposed architecture enables a multiplication with a reliability level close to the multipliers protected by TMR utilizing 63.54% less area and having 39.06% lower power-delay product compared to the exact multiplier. Mahdi Taheri, Natalia Cherezova, Samira Nazari, Ahsan Rafiq, Ali Azarpeyvand, Tara Ghasempouri, Masoud Daneshtalab, Jaan Raik, Maksim Jenihhin |
ETS | 7 |
| 2024 | Cost-Effective Fault Tolerance for CNNs Using Parameter Vulnerability Based Hardening and PruningabstractConvolutional Neural Networks (CNNs) have become integral in safety-critical applications, thus raising concerns about their fault tolerance. Conventional hardwaredependent fault tolerance methods, such as Triple Modular Redundancy (TMR), are computationally expensive, imposing a remarkable overhead on CNNs. Whereas fault tolerance techniques can be applied either at the hardware level or at the model levels, the latter provides more flexibility without sacrificing generality. This paper introduces a model-level hardening approach for CNNs by integrating error correction directly into the neural networks. The approach is hardwareagnostic and does not require any changes to the underlying accelerator device. Analyzing the vulnerability of parameters enables the duplication of selective filters/neurons so that their output channels are effectively corrected with an efficient and robust correction layer. The proposed method demonstrates fault resilience nearly equivalent to TMR-based correction but with significantly reduced overhead. Nevertheless, there exists an inherent overhead to the baseline CNNs. To tackle this issue, a cost-effective parameter vulnerability based pruning technique is proposed that outperforms the conventional pruning method, yielding smaller networks with a negligible accuracy loss. Remarkably, the hardened pruned CNNs perform up to $\mathbf{2 4 \%}$ faster than the hardened un-pruned ones. Mohammad Hasan Ahmadilivani, Seyedhamidreza Mousavi, Jaan Raik, Masoud Daneshtalab, Maksim Jenihhin |
IOLTS | 4 |
| 2024 | Contrastive Learning for Lane Detection via cross-similarityabstractDetecting lane markings in road scenes poses a significant challenge due to their intricate nature, which is susceptible to unfavorable conditions. While lane markings have strong shape priors, their visibility is easily compromised by varying lighting conditions, adverse weather, occlusions by other vehicles or pedestrians , road plane changes, and fading of colors over time. The detection process is further complicated by the presence of several lane shapes and natural variations, necessitating large amounts of high-quality and diverse data to train a robust lane detection model capable of handling various real-world scenarios. In this paper, we present a novel self-supervised learning method termed Contrastive Learning for Lane Detection via Cross-Similarity (CLLD) to enhance the resilience and effectiveness of lane detection models in real-world scenarios, particularly when the visibility of lane markings are compromised. CLLD introduces a novel contrastive learning (CL) method that assesses the similarity of local features within the global context of the input image. It uses the surrounding information to predict lane markings. This is achieved by integrating local feature contrastive learning with our newly proposed operation, dubbed cross-similarity . The local feature CL concentrates on extracting features from small patches, a necessity for accurately localizing lane segments. Meanwhile, cross-similarity captures global features, enabling the detection of obscured lane segments based on their surroundings. We enhance cross-similarity by randomly masking portions of input images in the process of augmentation. Extensive experiments on TuSimple and CuLane benchmark datasets demonstrate that CLLD consistently outperforms state-of-the-art contrastive learning methods, particularly in visibility-impairing conditions like shadows, while it also delivers comparable results under normal conditions. When compared to supervised learning, CLLD still excels in challenging scenarios such as shadows and crowded scenes, which are common in real-world driving. Ali Zoljodi, Sadegh Abadijou, Mina Alibeigi, Masoud Daneshtalab |
Pattern Recognit. Lett. | 4 |
| 2023 | FARMUR: Fair Adversarial Retraining to Mitigate Unfairness in Robustness
Seyed Ali Mousavi, Masoud Daneshtalab |
ADBIS | 3 |
| 2023 | NeuroPIM: Felxible Neural Accelerator for Processing-in-Memory ArchitecturesabstractThe performance of microprocessors under many modern workloads is mainly limited by the off-chip memory bandwidth. The emerging process-in-memory paradigm present a unique opportunity to reduce data movement overheads by moving computation closer to memory. State-of-the-art processing-in-memory proposals stack a logic layer on top of one or multiple memory layers in a 3D fashion and leverage the logic layer to build near-memory processing units. Such processing units are either application-specific accelerators or general-purpose cores. In this paper, we present NeuroPIM, a new processing-in-memory architecture that uses a neural network as the memory-side general-purpose accelerator. This design is mainly motivated by the observation that in many real-world applications, some program regions, or even the entire program, can be replaced by a neural network that is learned to approximate the program’s output. NeuroPIM benefits from both the flexibility of general-purpose processors and superior performance of application-specific accelerators. Experimental results show that NeuroPIM provides up to 41% speedup over a processor-side neural network accelerator and up to 8x speedup over a general-purpose processor. Ali Monavari Bidgoli, Sepideh Fattahi, Seyyed Hossein Seyyedaghaei Rezaei, Mehdi Modarressi, Masoud Daneshtalab |
DDECS | 5 |
| 2023 | APPRAISER: DNN Fault Resilience Analysis Employing Approximation ErrorsabstractNowadays, the extensive exploitation of Deep Neural Networks (DNNs) in safety-critical applications raises new reliability concerns. In practice, methods for fault injection by emulation in hardware are efficient and widely used to study the resilience of DNN architectures for mitigating reliability issues already at the early design stages. However, the state-of-the-art methods for fault injection by emulation incur a spectrum of time-, design-and control-complexity problems. To overcome these issues, a novel resiliency assessment method called APPRAISER is proposed that applies functional approximation for a non-conventional purpose and employs approximate computing errors for its interest. By adopting this concept in the resiliency assessment domain, APPRAISER provides thousands of times speed-up in the assessment process, while keeping high accuracy of the analysis. In this paper, APPRAISER is validated by comparing it with state-of-the-art approaches for fault injection by emulation in FPGA. By this, the feasibility of the idea is demonstrated, and a new perspective in resiliency evaluation for DNNs is opened. Mahdi Taheri, Mohammad Hasan Ahmadilivani, Maksim Jenihhin, Masoud Daneshtalab, Jaan Raik |
DDECS | 4 |
| 2023 | Comparative Evaluation of Various Generations of Controller Area Network Based on Timing AnalysisabstractThis paper performs a comparative evaluation of various generations of Controller Area Network (CAN), including the classical CAN, CAN Flexible Data-Rate (FD), and CAN Extra Long (XL). We utilize response-time analysis for the evaluation. In this regard, we identify that the state of the art lacks the response-time analysis for CAN XL. Hence, we discuss the worst-case transmission times calculations for CAN XL frames and incorporate them to the existing analysis for CAN to support response-time analysis of CAN XL frames. Using the extended analysis, we perform a comparative evaluation of the three generations of CAN by analyzing an automotive industrial use case. In crux, we show that using CAN FD is more advantageous than the classical CAN and CAN XL when using frames with payloads of up to 8 bytes, despite the fact that CAN XL supports higher bit rates. For frames with 12-64 bytes payloads, CAN FD performs better than CAN XL when running at the same bit rate, but CAN XL performs better when running at a higher bit rate. Additionally, we discovered that CAN XL performs better than the classical CAN and CAN FD when the frame payload is over 64 bytes, even if it runs at the same or higher bit rates than CAN FD. Aldin Berisa, Adis Panjevic, Imran Kovac, Hans Lyngbäck, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
ETFA | 6 |
| 2023 | DeepVigor: VulnerabIlity Value RanGes and FactORs for DNNs' Reliability AssessmentabstractDeep Neural Networks (DNNs) and their accelerators are being deployed ever more frequently in safety-critical applications leading to increasing reliability concerns. A traditional and accurate method for assessing DNNs’ reliability has been resorting to fault injection, which, however, suffers from prohibitive time complexity. While analytical and hybrid fault injection-/analytical-based methods have been proposed, they are either inaccurate or specific to particular accelerator architectures.In this work, we propose a novel accurate, fine-grain, metric-oriented, and accelerator-agnostic method called DeepVigor that provides vulnerability value ranges for DNN neurons’ outputs. An outcome of DeepVigor is an analytical model representing vulnerable and non-vulnerable ranges for each neuron that can be exploited to develop different techniques for improving DNNs’ reliability. Moreover, DeepVigor provides reliability assessment metrics based on vulnerability factors for bits, neurons, and layers using the vulnerability ranges.The proposed method is not only faster than fault injection but also provides extensive and accurate information about the reliability of DNNs, independent from the accelerator. The experimental evaluations in the paper indicate that the proposed vulnerability ranges are 99.9% to 100% accurate even when evaluated on previously unseen test data. Also, it is shown that the obtained vulnerability factors represent the criticality of bits, neurons, and layers proficiently. DeepVigor is implemented in the PyTorch framework and validated on complex DNN benchmarks. Mohammad Hasan Ahmadilivani, Mahdi Taheri, Jaan Raik, Masoud Daneshtalab, Maksim Jenihhin |
ETS | 4 |
| 2023 | Investigating and Analyzing CAN-to-TSN Gateway Forwarding TechniquesabstractController Area Network (CAN) and Ethernet network are expected to co-exist in automotive industry as Ethernet provides a high-bandwidth communication, while CAN is a legacy cost-effective solution. Due to the shortcomings of conventional switched Etherent, such as determinism, IEEE Time Sensitive Networking (TSN) task group developed a set of standards to enhance the switched Ethernet technology providing low-jitter and deterministic communication. Considering these two network domains, we investigate various design approaches for a gateway that connects a CAN domain to a TSN domain. We present three gateway forwarding techniques and we develop end-to-end delay analysis methods for them. Via the analysis methods and applying them to synthetic use cases we show that the intuitive existing approach of encapsulating multiple CAN frames into a single Ethernet frame is not necessarily an efficient solution. In fact, we demonstrate several cases where it is preferable to encapsulate only one CAN frame into a TSN frame, in particular when we use a high speed TSN network. The results have a significant impact on developing such gateways as the implementation of the one-to-one frame encapsulation is considerably simpler than other complex gateway-forwarding techniques. Aldin Berisa, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
ISORC | 3 |
| 2023 | End-to-end Timing Modeling and Analysis of TSN in Component-Based Vehicular SoftwareabstractIn this paper, we present an end-to-end timing model to capture timing information from software architectures of distributed embedded systems that use network communication based on the Time-Sensitive Networking (TSN) standards. Such a model is required as an input to perform end-to-end timing analysis of these systems. Furthermore, we present a methodology that aims at automated extraction of instances of the end-to-end timing model from component-based software architectures of the systems and the TSN network configurations. As a proof of concept, we implement the proposed end-to-end timing model and the extraction methodology in the Rubus Component Model (RCM) and its tool chain Rubus-ICE that are used in the vehicle industry. We demonstrate the usability of the proposed model and methodology by modeling a vehicular industrial use case and performing its timing analysis. Bahar Houtan, Mehmet Onur Aybek, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, John Lundbäck, Saad Mubeen |
ISORC | 4 |
| 2023 | Special Session: Approximation and Fault Resiliency of DNN AcceleratorsabstractDeep Learning, and in particular, Deep Neural Network (DNN) is nowadays widely used in many scenarios, including safety-critical applications such as autonomous driving. In this context, besides energy efficiency and performance, reliability plays a crucial role since a system failure can jeopardize human life. As with any other device, the reliability of hardware architectures running DNNs has to be evaluated, usually through costly fault injection campaigns. This paper explores approximation and fault resiliency of DNN accelerators. We propose to use approximate (AxC) arithmetic circuits to agilely emulate errors in hardware without performing fault injection on the DNN. To allow fast evaluation of AxC DNN, we developed an efficient GPU-based simulation framework. Further, we propose a fine-grain analysis of fault resiliency by examining fault propagation and masking in networks. Mohammad Hasan Ahmadilivani, Mario Barbareschi, Salvatore Barone, Alberto Bosio, Masoud Daneshtalab, Salvatore Della Torca, Gabriele Gavarini, Maksim Jenihhin, Jaan Raik, Annachiara Ruospo, Ernesto Sánchez 0001, Mahdi Taheri |
VTS | 5 |
| 2023 | Supporting end-to-end data propagation delay analysis for TSN-based distributed vehicular embedded systemsabstractIn this paper, we identify that the existing end-to-end data propagation delay analysis for distributed embedded systems can calculate pessimistic (over-estimated) analysis results when the nodes are synchronized. This is particularly the case of the Scheduled Traffic (ST) class in Time-sensitive Networking (TSN), which is scheduled offline according to the IEEE 802.1Qbv standard and the nodes are synchronized according to the IEEE 802.1AS standard. We present a comprehensive system model for distributed embedded systems that incorporates all of the above mentioned aspect as well as all traffic classes in TSN. We extend the analysis to support both synchronization and non-synchronization among the ECUs as well as offline schedules on the networks. The extended analysis can now be used to analyze all traffic classes in TSN when the nodes are synchronized without introducing any pessimism in the analysis results. We evaluate the proposed model and the extended analysis on a vehicular industrial use case. Bahar Houtan, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
J. Syst. Archit. | 3 |
| 2023 | A comprehensive systematic review of integration of time sensitive networking and 5G communicationabstractMany industrial real-time applications in various domains, e.g., automotive, industrial automation, industrial IoT, and industry 4.0, require ultra-low end-to-end network latency, often in the order of 10 milliseconds or less. The IEEE 802.1 time-sensitive networking (TSN) is a set of standards that supports the required low-latency wired communication with ultra-low jitter. The flexibility of such a wired connection can be increased if it is integrated with a mobile wireless network. The fifth generation of cellular networks (5G) is capable of supporting the required levels of network latency with the Ultra-Reliable Low Latency Communication (URLLC) service. To fully utilize the potential of these two technologies (TSN and 5G) in industrial applications, seamless integration of the TSN wired-based network with the 5G wireless-based network is needed. In this article, we provide a comprehensive and well-structured snapshot of the existing research on TSN-5G integration. In this regard, we present the planning, execution, and analysis results of the systematic review. We also identify the trends, technical characteristics, and potential gaps in the state of the art, thus highlighting future research directions in the integration of TSN and 5G communication technologies. We notice that 73% of the primary studies address the time synchronization in the integration of TSN and 5G technologies, introducing approaches with an accuracy starting from the levels of hundred nanoseconds to one microsecond. Majority of primary studies aim at optimizing communication latency in their approach, which is a key quality attribute in automotive and industrial automation applications today. Zenepe Satka, Mohammad Ashjaei, Hossein Fotouhi, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
J. Syst. Archit. | 4 |
| 2023 | DASS: Differentiable Architecture Search for Sparse Neural NetworksabstractThe deployment of Deep Neural Networks (DNNs) on edge devices is hindered by the substantial gap between performance requirements and available computational power. While recent research has made significant strides in developing pruning methods to build a sparse network for reducing the computing overhead of DNNs, there remains considerable accuracy loss, especially at high pruning ratios. We find that the architectures designed for dense networks by differentiable architecture search methods are ineffective when pruning mechanisms are applied to them. The main reason is that the current methods do not support sparse architectures in their search space and use a search objective that is made for dense networks and does not focus on sparsity. This paper proposes a new method to search for sparsity-friendly neural architectures. It is done by adding two new sparse operations to the search space and modifying the search objective. We propose two novel parametric SparseConv and SparseLinear operations in order to expand the search space to include sparse operations. In particular, these operations make a flexible search space due to using sparse parametric versions of linear and convolution operations. The proposed search objective lets us train the architecture based on the sparsity of the search space operations. Quantitative analyses demonstrate that architectures found through DASS outperform those used in the state-of-the-art sparse networks on the CIFAR-10 and ImageNet datasets. In terms of performance and hardware effectiveness, DASS increases the accuracy of the sparse version of MobileNet-v2 from 73.44% to 81.35% (+7.91% improvement) with a 3.87× faster inference time. Mohammad Loni, Mina Alibeigi, Masoud Daneshtalab |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | TAS: Ternarized Neural Architecture Search for Resource-Constrained Edge DevicesabstractTernary Neural Networks (TNNs) compress network weights and activation functions into 2-bit representation resulting in remarkable network compression and energy efficiency. However, there remains a significant gap in accuracy between TNNs and full-precision counterparts. Recent advances in Neural Architectures Search (NAS) promise opportunities in automated optimization for various deep learning tasks. Unfortunately, this area is unexplored for optimizing TNNs. This paper proposes TAS, a framework that drastically reduces the accuracy gap between TNNs and their full-precision counterparts by integrating quantization into the network design. We experienced that directly applying NAS to the ternary domain provides accuracy degradation as the search settings are customized for full-precision networks. To address this problem, we propose (i) a new cell template for ternary networks with maximum gradient propagation; and (ii) a novel learnable quantizer that adaptively relaxes the ternarization mechanism from the distribution of the weights and activation functions. Experimental results reveal that TAS delivers 2.64% higher accuracy and ≃2.8 ×memory saving over competing methods with the same bit-width resolution on the CIFAR-10 dataset. These results suggest that TAS is an effective method that paves the way for the efficient design of the next generation of quantized neural networks. Mohammad Loni, Mohammad Riazati, Masoud Daneshtalab, Mikael Sjödin |
DATE | 4 |
| 2022 | End-to-end Timing Model Extraction from TSN-Aware Distributed Vehicle SoftwareabstractExtraction of end-to-end timing information from software architectures of vehicular systems to support their timing analysis is a daunting challenge. To address this challenge, this paper presents a systematic method to extract this information from vehicular software architectures that can be distributed over several electronic control units connected by Time-Sensitive Networking (TSN) networks. As a proof of concept, the proposed extraction method is applied to an industrial component model, namely the Rubus Component Model (RCM), and its toolchain. Furthermore, the usability of the proposed method is demonstrated in an industrial use case from the vehicular domain. Bahar Houtan, Mehmet Onur Aybek, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
SEAA | 4 |
| 2022 | 3DLaneNAS: Neural Architecture Search for Accurate and Light-Weight 3D Lane Detection
Ali Zoljodi, Mohammad Loni, Sadegh Abadijou, Mina Alibeigi, Masoud Daneshtalab |
ICANN (1) | 5 |
| 2022 | QoS-MAN: A Novel QoS Mapping Algorithm for TSN-5G FlowsabstractIntegrating wired Ethernet networks, such as Time-Sensitive Networks (TSN), to 5G cellular network requires a flow management technique to efficiently map TSN traffic to 5G Quality-of-Service (QoS) flows. The 3GPP Release 16 provides a set of predefined QoS characteristics, such as priority level, packet delay budget, and maximum data burst volume, which can be used for the 5G QoS flows. Within this context, mapping TSN traffic flows to 5G QoS flows in an integrated TSN-5G network is of paramount importance as the mapping can significantly impact on the end-to-end QoS in the integrated network. In this paper, we present a novel and efficient mapping algorithm to map different TSN traffic flows to 5G QoS flows. To the best of our knowledge, this is the first QoS-aware mapping algorithm based on the application constraints used to exchange flows between TSN and 5G network domains. We evaluate the proposed mapping algorithm on synthetic scenarios with random sets of constraints on deadline, jitter, bandwidth, and packet loss rate. The evaluation results show that the proposed mapping algorithm can fulfill over 90% of the applications’ constraints. Zenepe Satka, Mohammad Ashjaei, Hossein Fotouhi, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
RTCSA | 4 |
| 2022 | Developing a Translation Technique for Converged TSN-5G CommunicationabstractTime Sensitive Networking (TSN) is a set of IEEE standards based on switched Ethernet that aim at meeting high-bandwidth and low-latency requirements in wired communication. TSN implementations typically do not support integration of wireless networks, which limits their applicability to many industrial applications that need both wired and wire-less communication. The development of 5G and its promised Ultra-Reliable and Low-Latency Communication (URLLC) in-tegrated with TSN would offer a promising solution to meet the bandwidth, latency and reliability requirements in these industrial applications. In order to support such an integration, we propose a technique to translate the traffic between TSN and 5G communication technologies. As a proof of concept, we implement the translation technique in a well-known TSN simulator, namely NeSTiNg, that is based on the OMNeT ++ tool. Furthermore, we evaluate the proposed technique using an automotive industrial use case. Zenepe Satka, David Pantzar, Alexander Magnusson, Mohammad Ashjaei, Hossein Fotouhi, Mikael Sjödin, Masoud Daneshtalab, Saad Mubeen |
WFCS | 7 |
| 2022 | FastStereoNet: A Fast Neural Architecture Search for Improving the Inference of Disparity Estimation on Resource-Limited PlatformsabstractConvolutional neural networks (CNNs) provide the best accuracy for disparity estimation. However, CNNs are computationally expensive, making them unfavorable for resource-limited devices with real-time constraints. Recent advances in neural architectures search (NAS) promise opportunities in automated optimization for disparity estimation. However, the main challenge of the NAS methods is the significant amount of computing time to explore a vast search space [e.g.,$1.6\times 10^{29}$] and costly training candidates. To reduce the NAS computational demand, many proxy-based NAS methods have been proposed. Despite their success, most of them are designed for comparatively small-scale learning tasks. In this article, we propose a fast NAS method, called FastStereoNet, to enable resource-aware NAS within an intractably large search space. FastStereoNet automatically searches for hardware-friendly CNN architectures based on late acceptance hill climbing (LAHC), followed by simulated annealing (SA). FastStereoNet also employs a fine-tuning with a transferred weights mechanism to improve the convergence of the search process. The collection of these ideas provides competitive results in terms of search time and strikes a balance between accuracy and efficiency. Compared to the state of the art, FastStereoNet provides$5.25\times $reduction in search time and$44.4\times $reduction in model size. These benefits are attained while yielding a comparable accuracy that enables seamless deployment of disparity estimation on resource-limited devices. Finally, FastStereoNet significantly improves the perception quality of disparity estimation deployed on field-programmable gate array and Intel Neural Compute Stick 2 accelerator in a significantly less onerous manner. Mohammad Loni, Ali Zoljodi, Amin Majd, Byung Hoon Ahn, Masoud Daneshtalab, Mikael Sjödin, Hadi Esmaeilzadeh |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | Network-on-ReRAM for Scalable Processing-in-Memory Architecture DesignabstractThe non-volatile metal-oxide resistive random access memory (ReRAM) is an emerging alternative for the current memory technologies. The unique capability of ReRAM to perform analog and digital arithmetic and logic operations has enabled this technology to incorporate both computation and memory capabilities on the same unit. Due to this interesting property, there is a growing trend in recent years to implement emerging data-intensive applications on ReRAM structures. A typical ReRAM-based processing-in-memory architecture may consist tens to hundreds of ReRAM units (mats) that can either store or process data. To support such large-scale ReRAM structure, this paper proposes a scalable network-on-ReRAM architecture. The proposed network employs a novel associative router architecture, designed based on the ReRAM-based content-addressable memories. With the in-memory packet processing capability, this router yields higher throughput and resource utilization levels than a conventional router. This router is technology compatible with ReRAM and as our evaluations show, employing it to build a network-on-ReRAM makes the emerging ReRAM-based processing-in-memory architectures more scalable and performance-efficient. Bita Dabiri, Mehdi Modarressi, Masoud Daneshtalab |
DSD | 3 |
| 2021 | Schedulability Analysis of Best-Effort Traffic in TSN NetworksabstractThis paper presents a schedulability analysis for the Best-Effort (BE) traffic class within Time Sensitive Networking (TSN) networks. The presented analysis considers several features in the TSN standards, including the Credit-Based Shaper (CBS), the Time-Aware Shaper (TAS) and the frame preemption. Although the BE class in TSN is primarily used for the traffic with no strict timing requirements, some industrial applications prefer to utilize this class for the non-hard real-time traffic instead of classes that use the CBS. The reason mainly lies in the fact that the complexity of TSN configuration becomes significantly high when the time-triggered traffic via the TAS and other classes via the CBS are used altogether. We demonstrate the applicability of the presented analysis on a vehicular application use case. We show that a network designer can get information on the schedulability of the BE traffic, based on which the network configuration can be further refined with respect to the application requirements. Bahar Houtan, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Sara Afshar, Saad Mubeen |
ETFA | 3 |
| 2021 | RoCo-NAS: Robust and Compact Neural Architecture SearchabstractDeep model compression has been studied widely, and state-of-the-art methods can now achieve high compression ratios with minimum accuracy loss. Recent advances in adversarial attacks reveal the inherent vulnerability of deep neural networks to slightly perturbed images called adversarial examples. Since then, extensive efforts have been performed to enhance deep networks' robustness via specialized loss functions and learning algorithms. Previous works suggest that network size and robustness against adversarial examples contradict on most occasions. In this paper, we investigate how to optimize compactness and robustness to adversarial attacks of neural network architectures while maintaining the accuracy using multi-objective neural architecture search. We propose the use of previously generated adversarial examples as an objective to evaluate the robustness of our models in addition to the number of floating-point operations to assess model complexity i.e. compactness. Experiments on some recent neural architecture search algorithms show that due to their limited search space they fail to find robust and compact architectures. By creating a novel neural architecture search (RoCo-NAS), we were able to evolve an architecture that is up to 7% more accurate against adversarial samples than its more complex architecture counterpart. Thus, the results show inherently robust architectures regardless of their size. This opens up a new range of possibilities for the exploration and design of deep neural networks using automatic architecture search. Vahid Geraeinejad, Sima Sinaei, Mehdi Modarressi, Masoud Daneshtalab |
IJCNN | 4 |
| 2021 | ELC-ECG: Efficient LSTM Cell for ECG Classification Based on Quantized ArchitectureabstractLong Short-Term Memory (LSTM) is one of the most popular and effective Recurrent Neural Network (RNN) models used for sequence learning in applications such as ECG signal classification. Complex LSTMs could hardly be deployed on resource-limited bio-medical wearable devices due to the huge amount of computations and memory requirements. Binary LSTMs are introduced to cope with this problem. However, naive binarization leads to significant accuracy loss in ECG classification. In this paper, we propose an efficient LSTM cell along with a novel hardware architecture for ECG classification. By deploying 5-level binarized inputs and just 1- level binarization for weights, output, and in-memory cell activations, the delay of one LSTM cell operation is reduced 50x with about 0.004% accuracy loss in comparison with full precision design of ECG classification. Seyed Ahmad Mirsalari, Najmeh Nazari, Seyed Ali Ansarmohammadi, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
ISCAS | 6 |
| 2021 | Time-Sensitive Networking in automotive embedded systems: State of the art and research opportunitiesabstractThe functionality advancements and novel customer features that are currently found in modern automotive systems require high-bandwidth and low-latency in-vehicle communications, which become even more compelling for autonomous vehicles. In a recent effort to meet these requirements, the IEEE Time-Sensitive Networking (TSN) task group has developed a set of standards that introduce novel features in Switched Ethernet. TSN standards offer, for example, a common notion of time through accurate and reliable clock synchronization, delay bounds for real-time traffic, time-driven transmissions, improved reliability, and much more. In order to fully utilize the potential of these novel protocols in the automotive domain, TSN should be seamlessly integrated into the state-of-the-art and state-of-practice model-based development processes for automotive embedded systems. Some of the core phases in these processes include software architecture modeling, timing predictability verification, simulation, and hardware realization and deployment. Moreover, throughout the development of automotive embedded systems, the safety and security requirements specified on these systems need to be duly taken into account. In this context, this work provides an overview of TSN in automotive applications and discusses the recent technological developments relevant to the adoption of TSN in automotive embedded systems. The work also points at the open challenges and future research directions. Mohammad Ashjaei, Lucia Lo Bello, Masoud Daneshtalab, Gaetano Patti, Sergio Saponara, Saad Mubeen |
J. Syst. Archit. | 3 |
| 2021 | Guest Editorial: Special issue on parallel, distributed, and network-based processing in next-generation embedded systems
Saad Mubeen, Lucia Lo Bello, Masoud Daneshtalab, Sergio Saponara |
J. Syst. Archit. | 3 |
| 2020 | DenseDisp: Resource-Aware Disparity Map Estimation by Compressing Siamese Neural ArchitectureabstractStereo vision cameras are flexible sensors due to providing heterogeneous information such as color, luminance, disparity map (depth), and shape of the objects. Today, Convolutional Neural Networks (CNNs) present the highest accuracy for the disparity map estimation [1]. However, CNNs require considerable computing capacity to process billions of floating-point operations in a real-time fashion. Besides, commercial stereo cameras produce huge size images (e.g., 10 Megapixels [2]), which impose a new computational cost to the system. The problem will be pronounced if we target resource-limited hardware for the implementation. In this paper, we propose DenseDisp, an automatic framework that designs a Siamese neural architecture for disparity map estimation in a reasonable time. DenseDisp leverages a meta-heuristic multi-objective exploration to discover hardware-friendly architectures by considering accuracy and network FLOPS as the optimization objectives. We explore the design space with four different fitness functions to improve the accuracy-FLOPS trade-off and convergency time of the DenseDisp. According to the experimental results, DenseDisp provides up to 39. 1x compression rate while losing around 5% accuracy compared to the state-of-the-art results. Mohammad Loni, Ali Zoljodi, Daniel Maier 0002, Amin Majd, Masoud Daneshtalab, Mikael Sjödin, Ben H. H. Juurlink, Reza Akbari |
CEC | 5 |
| 2020 | SHiLA: Synthesizing High-Level Assertions for High-Speed Validation of High-Level DesignsabstractIn the past, assertions were mostly used to validate the system through the design and simulation process. Later, a new method known as assertion synthesis was introduced, which enabled the designers to use the assertions for high-speed hardware emulation and safety and reliability insurance after tape-out. Although the synthesis of the assertions at the register transfer level is proposed and implemented in several works, none of them can be adopted for high-level assertions. In this paper, we propose the SHiLA framework and a detailed implementation guide by which assertion synthesis can also be applied to the high-level design processes. The proposed method, which is fully tool independent, is not only an enabler to highspeed assertion-assisted simulation but can also be used in other scenarios that need assertion synthesis, as it has the minimum possible effect on the main design's performance. Mohammad Riazati, Masoud Daneshtalab, Mikael Sjödin, Björn Lisper |
DDECS | 2 |
| 2020 | Adjustable self-healing methodology for accelerated functions in heterogeneous systemsabstractSelf-healing is a promising approach for designing reliable digital systems. It refers to the ability of a system to detect faults and automatically fixing them to avoid total failure. With the development of digital systems, heterogeneous systems, in which some parts of the system are executed on the programmable logic, and some other parts run on the processing elements (CPU), are becoming more prevalent. In this work, we propose an adjustable self-healing method that is applicable to heterogeneous systems with accelerated functions and enables the designers to add the self-healing feature to the design. In this method, by manipulating the software codes that are being executed on the processing element, we add the ability to verify the accelerated functions on the programmable logic and heal the possible failures to the system. This is done not only in a straightforward manner but also without being forced to choose a specific reliability-overhead point. The designer will have the option to select the optimum configuration for a desired reliability level. Experimental results on a large design including several accelerated functions are provided and show 42% improvement of reliability by having 27% overhead, as an example of the reliability-overhead point. Mohammad Riazati, Tara Ghasempouri, Masoud Daneshtalab, Jaan Raik, Mikael Sjödin, Björn Lisper |
DSD | 3 |
| 2020 | MuBiNN: Multi-Level Binarized Recurrent Neural Network for EEG Signal ClassificationabstractRecurrent Neural Networks (RNN) are widely used for learning sequences in applications such as EEG classification. Complex RNNs could be hardly deployed on wearable devices due to their computation and memory-intensive processing patterns. Generally, reduction in precision leads much more efficiency and binarized RNNs are introduced as energy-efficient solutions. However, naive binarization methods lead to significant accuracy loss in EEG classification. In this paper, we propose a multi-level binarized LSTM, which significantly reduces computations whereas ensuring an accuracy pretty close to the full precision LSTM. Our method reduces the delay of the 3-bit LSTM cell operation 47× with less than 0.01% accuracy loss. Seyed Ahmad Mirsalari, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
ISCAS | 4 |
| 2020 | Message from Program Co-Chairs: PDP 2020abstractParallel, Distributed, and Network-Based Processing has undergone impressive change over recent years. New architectures and applications have rapidly become the central focus of the discipline. These changes are often a result of cross-fertilization of parallel and distributed technologies with other rapidly evolving technologies. This is the reason why the PDP conference continues to have a distinctive composition: a main track invites papers over a broad range of topics, and ten Special Sessions focus each on a particular sub-domain related to the Parallel, Distributed and Network-based Computing research fields. Each Special Session has its own Chair(s) and Program Committee and invites and selects its own papers, all under the umbrella of the overall conference structure. The growing number of interesting and significant research papers submitted to PDP demonstrates that the conference is becoming an ever more important international event in the field of parallel and distributed computing research. In particular, the Program Committee of this edition received 120 submissions from 31 countries. On average each paper received 3.5 reviews, with no paper receiving fewer than three reviews. Masoud Daneshtalab, Mats Brorsson |
PDP | 1 |
| 2020 | Preface from General Co-Chairs: PDP 2020abstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Masoud Daneshtalab, Francesco Leporati, Mikael Sjödin |
PDP | 1 |
| 2020 | Scalable Parallel Genetic Algorithm For Solving Large Integer Linear Programming Models Derived From Behavioral SynthesisabstractSolving Integer Linear Programming (ILP) models generally lies in the category of NP-hard problems. Therefore, as the size of ILP models grows, the efficiency of exact algorithms for solving the models reduced significantly and for large models it is not possible to have the result. Genetic Algorithm (GA) is a metaheuristic method capable of adjusting and redesigning parameters and operations according to the characteristics of ILP models. Still GA has huge search space for large models and parallelization is a suitable technique to tackle this problem. This paper presents a scalable parallel GA to solve large ILP models derived from behavioral synthesis of digital circuits. We show that although models have non-binary variables, only binary variables are sufficient for coding chromosomes. We also use ”unknown” values for some genes to decrease the likelihood of inconsistency in the encoded constraints. Our experiments verify the efficiency and scalability of the proposed algorithm on multicore platforms. The proposed method outperforms IBM ILOG CPLEX 12.6 and MI-LXPM algorithm where the ILP models include 550 to 2258 int / binary decision variables. Also, the results indicate that the saturation point of using parallel processing elements for solving the large ILP models is at least 60. Mohammad K. Fallah, Mina Mirhosseini, Mahmood Fazlali, Masoud Daneshtalab |
PDP | 4 |
| 2020 | Multi-level Binarized LSTM in EEG Classification for Wearable DevicesabstractLong Short-Term Memory (LSTM) is widely used in various sequential applications. Complex LSTMs could be hardly deployed on wearable and resourced-limited devices due to the huge amount of computations and memory requirements. Binary LSTMs are introduced to cope with this problem, however, they lead to significant accuracy loss in some applications such as EEG classification which is essential to be deployed in wearable devices. In this paper, we propose an efficient multi-level binarized LSTM which has significantly reduced computations whereas ensuring an accuracy pretty close to full precision LSTM. By deploying 5-level binarized weights and inputs, our method reduces area and delay of MAC operation about $31\times and 27\times$ in 65nm technology, respectively with less than 0.01% accuracy loss. In contrast to many compute-intensive deep-learning approaches, the proposed algorithm is lightweight, and therefore, brings performance efficiency with accurate LSTM-based EEG classification to realtime wearable devices. Najmeh Nazari, Seyed Ahmad Mirsalari, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
PDP | 5 |
| 2019 | TOT-Net: An Endeavor Toward Optimizing Ternary Neural NetworksabstractHigh computation demands and big memory resources are the major implementation challenges of Convolutional Neural Networks (CNNs) especially for low-power and resource-limited embedded devices. Many binarized neural networks are recently proposed to address these issues. Although they have significantly decreased computation and memory footprint, they have suffered from accuracy loss especially for large datasets. In this paper, we propose TOT-Net, a ternarized neural network with [-1, 0, 1] values for both weights and activation functions that has simultaneously achieved a higher level of accuracy and less computational load. In fact, first, TOT-Net introduces a simple bitwise logic for convolution computations to reduce the cost of multiply operations. To improve the accuracy, selecting proper activation function and learning rate are influential, but also difficult. As the second contribution, we propose a novel piece-wise activation function, and optimized learning rate for different datasets. Our findings first reveal that 0.01 is a preferable learning rate for the studied datasets. Third, by using an evolutionary optimization approach, we found novel piece-wise activation functions customized for TOT-Net. According to the experimental results, TOT-Net achieves 2.15%, 8.77%, and 5.7/5.52% better accuracy compared to XNOR-Net on CIFAR-10, CIFAR-100, and ImageNet top-5/top-1 datasets, respectively. Najmeh Nazari, Mohammad Loni, Mostafa E. Salehi, Masoud Daneshtalab, Mikael Sjödin |
DSD | 4 |
| 2019 | NeuroPower: Designing Energy Efficient Convolutional Neural Network Architecture for Embedded Systems
Mohammad Loni, Ali Zoljodi, Sima Sinaei, Masoud Daneshtalab, Mikael Sjödin |
ICANN (1) | 4 |
| 2019 | SoFA: A Spark-oriented Fog ArchitectureabstractFog computing offers a wide range of service levels including low bandwidth usage, low response time, support of heterogeneous applications, and high energy efficiency. Therefore, real-time embedded applications could potentially benefit from Fog infrastructure. However, providing high system utilization is an important challenge of Fog computing especially for processing embedded applications. In addition, although Fog computing extends cloud computing by providing more energy efficiency, it still suffers from remarkable energy consumption, which is a limitation for embedded systems. To overcome the above limitations, in this paper, we propose SoFA, a Spark-oriented Fog architecture that leverages Spark functionalities to provide higher system utilization, energy efficiency and scalability. Compared to the common Fog computing platforms where edge devices are only responsible for processing data received from their IoT nodes, SoFA leverages the remaining processing capacity of all other edge devices. To attain this purpose, SoFA provides a distributed processing paradigm by the help of Spark to utilize the whole processing capacity of all the available edge devices leading to increase energy efficiency and system utilization. In other words, SoFA proposes a near-sensor processing solution in which the edge devices act as the Fog nodes. In addition, SoFA provides scalability by taking advantage of Spark functionalities. According to the experimental results, SoFA is a power-efficient and scalable solution desirable for embedded platforms by providing up to 3.1x energy efficiency for the Word-Count benchmark compared to the common Fog processing platform. Neda Maleki, Mohammad Loni, Masoud Daneshtalab, Mauro Conti, Hossein Fotouhi |
IECON | 3 |
| 2019 | Work in Progress: Investigating the Effects of High Priority Traffic on the Best Effort Traffic in TSN NetworksabstractThis paper investigates the effects of various parameters of high priority traffic classes on the Best Effort (BE) traffic in the networks based on the IEEE Time Sensitive Networking (TSN) standards. In this regard, the paper discusses ongoing work and presents preliminary results using a TSN simulator. The results indicate that several parameters of the high priority traffic such as periods, offsets and preemption modes can have a significant impact on the quality of service (e.g., guaranteed message delivery and message delays) of the BE traffic. Bahar Houtan, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen |
RTSS | 3 |
| 2019 | An energy-efficient partition-based XYZ-planar routing algorithm for a wireless network-on-chip
Fahimeh Yazdanpanah, Raheel Afsharmazayejani, Mohammad Alaei, Amin Rezaei 0001, Masoud Daneshtalab |
J. Supercomput. | 5 |
| 2018 | A Customized Processing-in-Memory Architecture for Biological Sequence AlignmentabstractSequence alignment is the most widely used operation in bioinformatics. With the exponential growth of the biological sequence databases, searching a database to find the optimal alignment for a query sequence (that can be at the order of hundreds of millions of characters long) would require excessive processing power and memory bandwidth. Sequence alignment algorithms can potentially benefit from the processing power of massive parallel processors due their simple arithmetic operations, coupled with the inherent fine-grained and coarse-grained parallelism that they exhibit. However, the limited memory bandwidth in conventional computing systems prevents exploiting the maximum achievable speedup. In this paper, we propose a processing-in-memory architecture as a viable solution for the excessive memory bandwidth demand of bioinformatics applications. The design is composed of a set of simple and lightweight processing elements, customized to the sequence alignment algorithm, integrated at the logic layer of an emerging 3D DRAM architecture. Experimental results show that the proposed architecture results in up to 2.4x speedup and 41% reduction in power consumption, compared to a processor-side parallel implementation. Nasrin Akbari, Mehdi Modarressi, Masoud Daneshtalab, Mohammad Loni |
ASAP | 3 |
| 2018 | Using Optimization, Learning, and Drone Reflexes to Maximize Safety of Swarms of DronesabstractDespite the growing popularity of swarm-based applications of drones, there is still a lack of approaches to maximize the safety of swarms of drones by minimizing the risks of drone collisions. In this paper, we present an approach that uses optimization, learning, and automatic immediate responses (reflexes) of drones to ensure safe operations of swarms of drones. The proposed approach integrates a high-performance dynamic evolutionary algorithm and a reinforcement learning algorithm to generate safe and efficient drone routes and then augments the generated routes with dynamically computed drone reflexes to prevent collisions with unforeseen obstacles in the flying zone. We also present a parallel implementation of the proposed approach and evaluate it against two benchmarks. The results show that the proposed approach maximizes safety and generates highly efficient drone routes. Amin Majd, Adnan Ashraf, Elena Troubitsyna, Masoud Daneshtalab |
CEC | 4 |
| 2018 | ADONN: Adaptive Design of Optimized Deep Neural Networks for Embedded SystemsabstractNowadays, many modern applications, e.g. autonomous system, and cloud data services need to capture and process a big amount of raw data at runtime that ultimately necessitates a high-performance computing model. Deep Neural Network (DNN) has already revealed its learning capabilities in runtime data processing for modern applications. However, DNNs are becoming more deep sophisticated models for gaining higher accuracy which require a remarkable computing capacity. Considering high-performance cloud infrastructure as a supplier of required computational throughput is often not feasible. Instead, we intend to find a near-sensor processing solution which will lower the need for network bandwidth and increase privacy and power efficiency, as well as guaranteeing worst-case response-times. Toward this goal, we introduce ADONN framework, which aims to automatically design a highly robust DNN architecture for embedded devices as the closest processing unit to the sensors. ADONN adroitly searches the design space to find improved neural architectures. Our proposed framework takes advantage of a multi-objective evolutionary approach, which exploits a pruned design space inspired by a dense architecture. Unlike recent works that mainly have tried to generate highly accurate networks, ADONN also considers the network size factor as the second objective to build a highly optimized network fitting with limited computational resource budgets while delivers comparable accuracy level. In comparison with the best result on CIFAR-10 dataset, a generated network by ADONN presents up to 26.4 compression rate while loses only 4% accuracy. In addition, ADONN maps the generated DNN on the commodity programmable devices including ARM Processor, High-Performance CPU, GPU, and FPGA. Mohammad Loni, Masoud Daneshtalab, Mikael Sjödin |
DSD | 2 |
| 2018 | Reconfigurable Network-on-Chip for 3D Neural Network AcceleratorsabstractParallel hardware accelerators for large-scale neural networks typically consist of several processing nodes, arranged as a multi- or many-core system-on-chip, connected by a network-on-chip (NoC). Recent proposals also benefit from the emerging 3D memory-on-logic architectures to provide sufficient bandwidth for neural networks. Handling the heavy traffic between neurons and memory and also the multicast-based inter-neuron traffic, which often varies over time, is the most challenging design consideration for the networks-on-chip in such accelerators. To address these issues, a reconfigurable network-on-chip architecture for 3D memory-on-logic neural network accelerators is presented in this paper. The reconfigurable NoC can adapt its topology to the on-chip traffic patterns. It can be also configured as a tree-like structure to support multicast-based neuron-to-neuron and memory-to-neuron traffic of neural networks. The evaluation results show that the proposed architecture can better manage the multicast-based traffic of neural networks than some state-of-the-art topologies and considerably increase throughput and power efficiency. Arash Firuzan, Mehdi Modarressi, Masoud Daneshtalab, Midia Reshadi |
NOCS | 3 |
| 2018 | Integrating Learning, Optimization, and Prediction for Efficient Navigation of Swarms of DronesabstractSwarms of drones are increasingly been used in a variety of monitoring and surveillance, search and rescue, and photography and filming tasks. However, despite the growing popularity of swarm-based applications of drones, there is still a lack of approaches to generate efficient drone routes while minimizing the risks of drone collisions. In this paper, we present a novel approach that integrates learning, optimization, and prediction for generating efficient and safe routes for swarms of drones. The proposed approach comprises three main components: (1) a high-performance dynamic evolutionary algorithm for optimizing drone routes, (2) a reinforcement learning algorithm for incorporating the feedback and runtime data about the system state, and (3) a prediction approach to predict the movement of drones and moving obstacles in the flying zone. We also present a parallel implementation of the proposed approach and evaluate it against two benchmarks. The results demonstrate that the proposed approach allows to significantly reduce the route lengths and computation overhead while producing efficient and safe routes. Amin Majd, Adnan Ashraf, Elena Troubitsyna, Masoud Daneshtalab |
PDP | 4 |
| 2018 | Parallel imperialist competitive algorithmsabstractSummary The importance of optimization and NP‐problem solving cannot be overemphasized. The usefulness and popularity of evolutionary computing methods are also well established. There are various types of evolutionary methods; they are mostly sequential but some of them have parallel implementations as well. We propose a multi‐population method to parallelize the Imperialist Competitive Algorithm. The algorithm has been implemented with the Message Passing Interface on 2 computer platforms, and we have tested our method based on shared memory and message passing architectural models. An outstanding performance is obtained, demonstrating that the proposed method is very efficient concerning both speed and accuracy. In addition, compared with a set of existing well‐known parallel algorithms, our approach obtains more accurate results within a shorter time period. Amin Majd, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila, Shahriar Lotfi, Hannu Tenhunen |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | EbDa: A New Theory on Design and Verification of Deadlock-free Interconnection Networks
Masoumeh Ebrahimi, Masoud Daneshtalab |
ISCA | 2 |
| 2017 | Hierarchal Placement of Smart Mobile Access Points in Wireless Sensor Networks Using Fog ComputingabstractRecent advances in computing and sensor technologies have facilitated the emergence of increasingly sophisticated and complex cyber-physical systems and wireless sensor networks. Moreover, integration of cyber-physical systems and wireless sensor networks with other contemporary technologies, such as unmanned aerial vehicles (i.e. drones) and fog computing, enables the creation of completely new smart solutions. By building upon the concept of a Smart Mobile Access Point (SMAP), which is a key element for a smart network, we propose a novel hierarchical placement strategy for SMAPs to improve scalability of SMAP based monitoring systems. SMAPs predict communication behavior based on information collected from the network, and select the best approach to support the network at any given time. In order to improve the network performance, they can autonomously change their positions. Therefore, placement of SMAPs has an important role in such systems. Initial placement of SMAPs is an NP problem. We solve it using a parallel implementation of the genetic algorithm with an efficient evaluation phase. The adopted hierarchical placement approach is scalable, it enables construction of arbitrarily large SMAP based systems. Amin Majd, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2017 | Multi-objective Task Mapping Approach for Wireless NoC in Dark Silicon AgeabstractHybrid Wireless Network-on-Chip (HWNoC) provides high bandwidth, low latency and flexible topology configurations, making this emerging technology a scalable communication fabric for future Many-Core System-on-Chips (MCSoCs). On the other hand, dark silicon is dominating the chip footage of upcoming MCSoCs since Dennard scaling fails due to the voltage scaling problem that results in higher power densities. Moreover, congestion avoidance and hot-spot prevention are two important challenges of HWNoC-based MCSoCs in dark silicon age, Therefore, in this paper, a novel task mapping approach for HWNoC is introduced in order to first balance the usage of wireless links by avoiding congestion over wireless routers and second spread temperature across the whole chip by utilizing dark silicon. Simulation results show significant improvement in both congestion and temperature control of the system, compared to state-of-the-art works. Amin Rezaei 0001, Danella Zhao, Masoud Daneshtalab, Hai Zhou 0001 |
PDP | 3 |
| 2017 | Customizing Clos Network-on-Chip for Neural NetworksabstractLarge-scale neural network accelerators are often implemented as a many-core chip and rely on a network-on-chip to manage the huge amount of inter-neuron traffic. The baseline and different variations of the well-known mesh and tree topologies are the most popular topologies in prior many-core implementations of neural networks. However, the grid-like mesh and hierarchical tree topologies suffer from high diameter and low bisection bandwidth, respectively. In this paper, we present ClosNN, a customized Clos topology for Neural Networks. The inherent capability of Clos to support multicast and broadcast traffic in a simple and efficient way, as well as its adaptable bisection bandwidth, is the major motivation behind proposing a customized version of this topology as the communication infrastructure of large-scale neural network implementations. We compare ClosNN with some state-of-the-art NoC topologies adopted in recent neural network hardware accelerators and show that it offers lower average message hop count and higher throughput, which directly translates to faster neural information processing. Reza Hojabr, Mehdi Modarressi, Masoud Daneshtalab, Ali Yasoubi, Ahmad Khonsari |
IEEE Trans. Computers | 3 |
| 2016 | Shift sprinting: fine-grained temperature-aware NoC-based MCSoC architecture in dark silicon ageabstractReliability is a critical feature of chip integration and unreliability can lead to performance, cost, and time-to-market penalties. Moreover, upcoming Many-Core System-on-Chips (MCSoCs), notably future generations of mobile devices, will suffer from high power densities due to the dark silicon problem. Thus, in this paper, a novel NoC-based MCSoC architecture, called Shift Sprinting, is introduced in order to reliably utilize dark silicon under the power budget constraint. By employing the concept of distributional sprinting, our proposed architecture provides Quality of Service (QoS) to efficiently run real-time streaming applications in mobile devices. Simulation results show meaningful gain in performance and reliability of the system compared to state-of-the-art works. Amin Rezaei 0001, Danella Zhao, Masoud Daneshtalab, Hongyi Wu |
DAC | 3 |
| 2016 | Fault-tolerant 3-D network-on-chip design using dynamic link sharing
Seyyed Hossein Seyyedaghaei Rezaei, Mehdi Modarressi, Reza Yazdani, Masoud Daneshtalab |
DATE | 4 |
| 2016 | PICA: Multi-population Implementation of Parallel Imperialist Competitive AlgorithmsabstractThe importance of optimization and NP-problems solving cannot be over emphasized. The usefulness and popularity of evolutionary computing methods are also well established. There are various types of evolutionary methods that are mostly sequential, and some others have parallel implementation. We propose a method to parallelize Imperialist Competitive Algorithm (Multi-Population). The algorithm has been implemented with MPI on two platforms and have tested our algorithms on a shared-memory and message passing architecture. An outstanding performance is obtained, which indicates that the method is efficient concern to speed and accuracy. In the second step, the proposed algorithm is compared with a set of existing well known parallel algorithms and is indicated that it obtains more accurate solutions in a lower time. Amin Majd, Shahriar Lotfi, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila |
PDP | 4 |
| 2016 | Efficient Congestion-Aware Scheme for Wireless on-Chip NetworksabstractWireless NoC is becoming popular to be a promising future on-chip interconnection network as a result of high bandwidth, low latency and flexible topology configurations provided by this emerging technology. Nonetheless, congestion occurrence in wireless routers negatively affects the usability of high speed wireless links and considerably increases the network latency, therefore, in this paper, a congestion-aware platform (CAP-W) is introduced for wireless NoCs in order to reduce both internal and external congestions. The whole platform of CAP-W consists of an adaptive routing algorithm that balances utilization of wired and wireless networks, a dynamic task mapping approach that tries to minimize congestion probability, and a task migration strategy that considers dynamic variation of application behaviors. Simulation results show significant gain in congestion control over PEs of wireless NoC, compared to state-of-the-art works. Amin Rezaei 0001, Masoud Daneshtalab, Maurizio Palesi, Danella Zhao |
PDP | 2 |
| 2016 | A Three-Dimensional Networks-on-Chip Architecture with Dynamic Buffer Sharingabstract3D integration is a practical solution for overcoming the failure of Dennard scaling in future technology generations. This emerging technology stacks several die slices on top of each other on a single chip in order to provide higher-bandwidth and lower-latency than a 2D design due to extremely shorter inter-layer distances in the third dimension and. In this paper, we leverage the low-latency vertical links to address buffer management, one of the most important design and management issues in Network-on-Chip (NoC). To this end, we present VerBuS, an architecture for 3D routers with Vertical BUffer Sharing capability enabled by ultra-low latency vertical links of a 3D chip. VerBuS can share virtual channels (VC) between vertically stacked routers. This way, the buffering capacity of a highly loaded router is increased by using idle VCs of vertically adjacent routers. Experimental results show up to 20% improvement in NoC performance metrics over state-of-the-art 3D router designs. Seyyed Hossein Seyyedaghaei Rezaei, Mehdi Modarressi, Masoud Daneshtalab, Shervin Roshanisefat |
PDP | 3 |
| 2016 | Special issue on energy efficient methods and systems in the emerging cloud era
Maurizio Palesi, Mario Collotta, Masoud Daneshtalab, Pradip Bose |
J. Comput. Syst. Sci. | 3 |
| 2016 | Non-Blocking Testing for Network-on-ChipabstractTo achieve high reliability in on-chip networks, it is necessary to test the network as frequently as possible to detect physical failures before they lead to system-level failures. A main obstacle is that the circuit under test has to be isolated, resulting in network cuts and packet blockage which limit the testing frequency. To address this issue, we propose a comprehensive network-level approach which could test multiple routers simultaneously at high speed without blocking or dropping packets. We first introduce a reconfigurable router architecture allowing the cores to keep their connections with the network while the routers are under test. A deadlock-free and highly adaptive routing algorithm is proposed to support reconfigurations for testing. In addition, a testing sequence is defined to allow testing multiple routers to avoid dropping of packets. A procedure is proposed to control the behavior of the affected packets during the transition of a router from the normal to the testing mode and vice versa. This approach neither interrupts the execution of applications nor has a significant impact on the execution time. Experiments with the PARSEC benchmarks on an 8x8 NoC-based chip multiprocessors show only 3 percent execution time increase with four routers simultaneously under test. Letian Huang, Junshi Wang, Masoumeh Ebrahimi, Masoud Daneshtalab, Xiaofan Zhang 0004, Guangjun Li, Axel Jantsch |
IEEE Trans. Computers | 4 |
| 2016 | TransMap: Transformation Based Remapping and Parallelism for High Utilization and Energy Efficiency in CGRAsabstractIn the era of platforms hosting multiple applications with arbitrary inter application communication and computation patterns, compile time mapping decisions are neither optimal nor desirable. As a solution to this problem, recently proposed architectures offer run-time remapping. The run-time remapping techniques displace or parallelize/serialize an application to optimize different parameters (e.g., utilization and energy). To implement the dynamic remapping, reconfigurable architectures commonly store multiple (compile-time generated) implementations of an application. Each implementation represents a different platform location and/or degree of parallelism. The optimal implementation is selected at run-time. However, the compile-time binding either incurs excessive configuration memory overheads and/or is unable to map/parallelize an application even when sufficient resources are available. As a solution to this problem, we present Transformation based reMapping and parallelism (TransMap). TransMap stores only a single implementation and applies a series for transformations to the stored bitstream for remapping or parallelizing an application. Compared to state of the art, in addition to simple relocation in horizontal/vertical directions, TransMap also allows to rotate an application for mapping or parallelizing an application in resource constrained scenarios. By storing only a single implementation, TransMap offers significant reductions in configuration memory requirements (up to 73 percent for the tested applications), compared to state of the art compaction techniques. Simulation results reveal that the additional flexibility reduces the energy requirements by 33 percent and enhances the device utilization by 50 percent for the tested applications. Gate level analysis reveals that TransMap incurs negligible silicon (0.2 percent of the platform) and timing (6 additional cycles per application) penalty. Syed M. A. H. Jafri, Masoud Daneshtalab, Naeem Abbas, Guillermo Serrano Leon, Ahmed Hemani |
IEEE Trans. Computers | 2 |
| 2016 | On Fine-Grained Runtime Power Budgeting for Networks-on-Chip SystemsabstractPower budgeting is an essential aspect of networks-on-chip (NoC) to meet the power constraint for on-chip communications while assuring the best possible overall system performance. For simplicity and ease of implementation, existing NoC power budgeting schemes treat all the individual routers uniformly when allocating power to them. However, such homogeneous power budgeting schemes ignore the fact that the workloads of different NoC routers may vary significantly, and thus may provide excess power to routers with low workloads, whereas insufficient power to those with high workloads. In this paper, we formulate the NoC power budgeting problem in order to optimize the network performance over a power budget through per-router frequency scaling. We take into account of heterogeneous workloads across different routers as imposed by variations in traffic. Correspondingly, we propose a fine-grained solution using an agile algorithm with low time complexity. Frequency of each router is set individually according to its contribution to the average network latency while meeting the power budget. Experimental results have confirmed that with fairly low runtime and hardware overhead, the proposed scheme can help save up to$50$percent application execution time when compared with the latest proposed methods. Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab |
IEEE Trans. Computers | 6 |
| 2015 | Fine-grained runtime power budgeting for networks-on-chipabstractPower budgeting for NoC needs to be performed to meet limited power budget while assuring the best possible overall system performance. For simplicity and ease of implementation, existing NoC power budgeting schemes, irrespective of the fact that the packet arrival rates of different NoC routers may vary significantly, treat all the individual routers indiscriminately when allocating power to them. However, such homogeneous power allocation may provide excess power to routers with low packet arrival rates whereas insufficient power to those with high arrival rates. In this paper, we formulate the NoC power budgeting problem as to optimize the network performance over a power budget through per-router frequency scaling, taking into account of heterogeneous packet arrival rates across different routers as imposed by run time traffic dynamics. Correspondingly, we propose a fine-grained solution using an agile dynamic programming network with a linear time complexity. In essence, frequency of a router is set individually according to its contribution to the average network latency while meeting the power budget. Experimental results have confirmed that with fairly low runtime and hardware overhead, the proposed scheme can help save up to 50% application execution time when compared with the best existing methods. Xiaohang Wang 0001, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab |
ASP-DAC | 6 |
| 2015 | CuPAN - High Throughput On-chip Interconnection for Neural Networks
Ali Yasoubi, Reza Hojabr, Hengameh Takshi, Mehdi Modarressi, Masoud Daneshtalab |
ICONIP (3) | 5 |
| 2015 | FIST: A Framework to Interleave Spiking Neural Networks on CGRAsabstractCoarse Grained Reconfigurable Architectures (CGRAs) are emerging as enabling platforms to meet the high performance demanded by modern embedded applications. In many application domains (e.g. robotics and cognitive embedded systems), the CGRAs are required to simultaneously host processing (e.g. Audio/video acquisition) and estimation (e.g. audio/video/image recognition) tasks. Recent works have revealed that the efficiency and scalability of the estimation algorithms can be significantly improved by using neural networks. However, existing CGRAs commonly employ homogeneous processing resources for both the tasks. To realize the best of both the worlds (conventional processing and neural networks), we present FIST. FIST allows the processing elements and the network to dynamically morph into either conventional CGRA or a neural network, depending on the hosted application. We have chosen the DRRA as a vehicle to study the feasibility and overheads of our approach. Synthesis results reveal that the proposed enhancements incur negligible overheads (4.4% area and 9.1% power) compared to the original DRRA cell. Tuan Ngyen, Syed M. A. H. Jafri, Masoud Daneshtalab, Ahmed Hemani, Sergei Dytckov, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2015 | Dynamic Application Mapping Algorithm for Wireless Network-on-ChipabstractBecause of high bandwidth, low latency and flexible topology configurations provided by wireless NoC, this emerging technology is gaining momentum to be a promising future on-chip interconnection paradigm. However, congestion occurrence in wireless routers reduces the benefit of high speed wireless links and significantly increases the network latency, therefore, in this paper, a Dynamic Application Mapping Algorithm (DAMA) is introduced for wireless NoCs in order to reduce both internal and external congestion. DAMA has three key steps: finding the first node to map, choosing the first task to be mapped onto the first node, and allocation of the remaining tasks to the remaining nodes. Simulation results show significant gain in the mapping cost functions compared to state-of-the-art works. Amin Rezaei 0001, Masoud Daneshtalab, Danella Zhao, Farshad Safaei, Xiaohang Wang 0001, Masoumeh Ebrahimi |
PDP | 2 |
| 2015 | Parallel Implementation of Fuzzified Pattern Matching Algorithm on GPUabstractApproximate pattern discovery is one of the fundamental and challenging problems in computer science. Fast and high performance algorithms are highly demanded in many applications in bioinformatics and computational molecular biology, which are the domains that are mostly and directly benefit from any enhancement of pattern matching theoretical knowledge and solutions. This paper proposed an efficient GPU implementation of fuzzified Aho-Corasick algorithm using Levenshtein method and N-gram technique as a solution for approximate pattern matching problem. Shima Soroushnia, Masoud Daneshtalab, Tapio Pahikkala, Juha Plosila |
PDP | 2 |
| 2015 | On-chip parallel and network-based systems
Masoud Daneshtalab, Nader Bagherzadeh, Hamid Sarbazi-Azad |
Integr. | 1 |
| 2015 | An efficient runtime power allocation scheme for many-core systems inspired from auction theory
Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab |
Integr. | 6 |
| 2015 | Special Issue on Emerging Many-Core Systems for Exascale ComputingabstractNo abstract available. Masoud Daneshtalab, Farhad Mehdipour, Zhiyi Yu, Hannu Tenhunen |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2015 | In-order delivery approach for 2D and 3D NoCs
Masoud Daneshtalab, Masoumeh Ebrahimi, Sergei Dytckov, Juha Plosila |
J. Supercomput. | 1 |
| 2014 | Agile frequency scaling for adaptive power allocation in many-core systems powered by renewable energy sourcesabstractAs low-power electronics and miniaturization conspire to populate the world with emerging devices, one appealing approach is to power these multi-core/many-core-based devices with energy harvested from various environments. Of the most important issues concerning these devices is how to effectively allocate power budget among the cores competing for power, which is formulated as one specific type of power-performance optimization problem in this paper. We attempt to solve this problem by proposing an Adaptive Power Allocation Technique (APAT) that uses a dynamic programming network. Our goal here is to maximize the overall system performance, taking into account a unique yet challenging fact that, available power budget might have to undergo a significant change when a renewable energy source is scavenging. APAT has a linear time complexity and low hardware overhead. Experiments have confirmed that APAT can reduce 20 ~ 30% of execution time compared to other state-of-the-art power allocation algorithms. In addition, as APAT is quite insensitive to the changing rate of the power, lending itself well for power management in many-core systems powered by energy-harvesting sources. Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Terrence S. T. Mak |
ASP-DAC | 5 |
| 2014 | Adaptive power allocation for many-core systems inspired from multiagent auction modelabstractScaling of future many-core chips is hindered by the challenge imposed by ever-escalating power consumption. At its worst, an increasing fraction of the chips will have to be shut down, as power supply is inadequate to simultaneously switch all the transistors. This so-called dark silicon problem brings up a critical issue regarding how to achieve the maximum performance within a given limited power budget. This issue is further complicated by two facts. First, high variation in power budget calls for wide range power control capability, whereas most current frequency/voltage scaling techniques cannot effectively adjust power over such a wide range. Second, as the applications' behavior becomes more complicated, there is a pressing need for scalability and global coordination, rendering heuristic-based centralized or fully distributed control schemes inefficient. To address the aforementioned problems, in this paper, a power allocation method employing multiagent auction models is proposed, referred as Hierarchal MultiAgent based Power allocation (HiMAP). Tiles act the role of consumers to bid for power budget and the whole process is modeled by a combinatorial auction, whereas HiMAP finds the Walrasian equilibria. Experimental results have confirmed that HiMAP can reduce the execution time by as much as 45% compared to three competing methods. The runtime overhead and cost of HiMAP are also small, which makes it suitable for adaptive power allocation in many-core systems. Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi |
DATE | 6 |
| 2014 | Parameterized AES-Based Crypto Processor for FPGAsabstractIn this paper, we propose a parameterized crypto co-processor based on Advanced Encryption Standard (AES). This parameterized AES module is combined with a 32-bit general purpose 5-stage pipelined MIPS processor. The AES module used in this paper is fully pipelined. The processor fetches an instruction from the instruction memory and sends it to the decode stage. If the instruction is the crypto instruction it is pushed into the AES module during the decode stage. However if the instruction belongs to the MIPS processor, the remaining cycles will be completed on the MIPS processor. The parameterized AES module has different latencies on different rounds of AES according to the application requirements. The effects of different number of rounds on latency, memory, and area are studied and reported. Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen, Sergei Dytckov, Giovanni Beltrame |
DSD | 2 |
| 2014 | Efficient STDP Micro-Architecture for Silicon Spiking Neural NetworksabstractSpiking neural networks (SNNs) are the closest approach to biological neurons in comparison with conventional artificial neural networks (ANN). SNNs are composed of neurons and synapses which are interconnected with a complex pattern. As communication in such massively parallel computational systems is getting critical, the network-on-chip (NoC) becomes a promising solution to provide a scalable and robust interconnection fabric. However, using NoC for large-scale SNNs arises a trade-off between scalability, throughput, neuron/router ratio (cluster size), and area overhead. In this paper, we tackle the trade-off using a clustering approach and try to optimize the synaptic resource utilization. An optimal cluster size can provide the lowest area overhead and power consumption. For the learning purposes, a phenomenon known as spike-timing-dependent plasticity (STDP) is utilized. The micro-architectures of the network, clusters, and the computational neurons are also described. The presented approach suggests a promising solution of integrating NoCs and STDP-based SNNs for the optimal performance based on the underlying application. Sergei Dytckov, Masoud Daneshtalab, Masoumeh Ebrahimi, Hassan Anwar, Juha Plosila, Hannu Tenhunen |
DSD | 2 |
| 2014 | Morphable Compression Architecture for Efficient Configuration in CGRAsabstractToday, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads (up to 50% area of the overall platform). As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties (i.e. compression ratio and decoding time), and is therefore best suited for a particular class of applications (and situation). However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform. The proposed architecture allows each application to enjoy a separate compression/decompression hierarchy (consisting of various types and implementations of hardware/software decoders) tailored to its needs. Thereby, our solution offers minimal memory while meeting the required configuration deadlines. Simulation results, using different applications (FFT, Matrix multiplication, and WLAN), reveal that the choice of compression hierarchy has a significant impact on compression ratio (from configware replication to 52%) and configuration cycles (from 33 nsec to 1.5 secs) for the tested applications. Synthesis results reveal that introducing adaptivity incurs negligible additional overheads (1%) compared to the overall platform area. Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen |
DSD | 3 |
| 2014 | Customizable Compression Architecture for Efficient Configuration in CGRAsabstractToday, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads. As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties, and is therefore best suited for a particular class of applications. However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform. Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen |
FCCM | 3 |
| 2014 | TransPar: Transformation based dynamic Parallelism for low power CGRAsabstractCoarse Grained Reconfigurable Architectures (CGRAs) are emerging as enabling platforms to meet the high performance demanded by modern applications (e.g. 4G, CDMA, etc.). Recently proposed CGRAs offer runtime parallelism to reduce energy consumption (by lowering voltage/frequency). To implement the runtime parallelism, CGRAs commonly store multiple compile-time generated implementations of an application (with different degree of parallelism) and select the optimal version at runtime. However, the compile-time binding incurs excessive configuration memory overheads and/or is unable to parallelize an application even when sufficient resources are available. As a solution to this problem, we propose Transformation based dynamic Parallelism (TransPar). TransPar stores only a single implementation and applies a series for transformations to generate the bitstream for the parallel version. In addition, it also allows to displace and/or rotate an application to parallelize in resource constrained scenarios. By storing only a single implementation, TransPar offers significant reductions in configuration memory requirements (up to 73% for the tested applications), compared to state of the art compaction techniques. Simulation and synthesis results, using real applications, reveal that the additional flexibility allows up to 33% energy reduction compared to static memory based parallelism techniques. Gate level analysis reveals that TransPar incurs negligible silicon (0.2% of the platform) and timing (6 additional cycles per application) penalty. Syed M. A. H. Jafri, Guilermo Serrano, Masoud Daneshtalab, Naeem Abbas, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen |
FPL | 3 |
| 2014 | Highly adaptive and congestion-aware routing for 3D NoCsabstractIn this paper, we propose a novel highly adaptive and congestion aware routing algorithm 3D meshes which is equally applicable to 2D meshes as well. The proposed algorithm allows cyclic dependencies in channel dependency graph (CDG) providing higher degree of adaptiveness. The algorithm uses congestion-aware channel selection strategy that results balanced distribution of traffic flows across the network. A packet follows non-minimal paths only when minimal paths are congested at the neighboring channels. The deadlock avoidance methodology adopted by our algorithm remains cost-efficient as it uses one extra virtual channel along each of Y and Z dimensions to achieve deadlock freedom. Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Seok-Bum Ko, Mark Zwolinski |
ACM Great Lakes Symposium on VLSI | 4 |
| 2014 | A novel non-minimal/minimal turn model for highly adaptive routing in 2D NoCsabstractNetworks-on-Chip (NoCs) are emerging as a promising communication paradigm to overcome bottleneck of traditional bus-based interconnects for current micro-architectures (MCSoC and CMP). One of the current issues in NoC routing is the use of acyclic Channel Dependency Graph (CDG) for deadlock freedom. This requirement forces certain routing turns to be prohibited, thus, reducing the degree of adaptiveness. In this paper, we propose a novel non-minimal turn model which allows cycles in CDG provided that Extended Channel Dependency Graph (ECDG) remains acyclic. The proposed turn model reduces number of restrictions on routing turns, hence able to provide path diversity through additional minimal and non-minimal routes between source and destination. Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Pankaj Kumar Srivastava, Seok-Bum Ko, Mark Zwolinski |
NOCS | 4 |
| 2014 | Integration of AES on Heterogeneous Many-Core SystemabstractIncreasing in the transistor density in a single chip makes it possible for many-core systems to utilize design space for implementing complex embedded systems. In this paper, we propose an architecture for heterogeneous many-core system to integrate block cipher algorithm which is based on Advanced Encryption Standard (AES). In order to implement AES as a crypto-core along with heterogeneous many-core system two different approaches are proposed. In the first approach, the platform is composed of an AES module, a crypto-core, a network interface, and an internal memory which are managed through a controller. In the second approach, the platform is composed of a Direct Memory Access (DMA), network interface, an internal memory, and a microprocessor in which the AES module is integrated as a crypto-core. Both approaches have been analyzed and compared in terms of area overhead and performance. Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Marco Ramírez 0001, Juha Plosila, Hannu Tenhunen |
PDP | 2 |
| 2014 | A novel non-minimal turn model for highly adaptive routing in 2D NoCsabstractNetwork-on-Chip (NoC) is emerging as a promising communication paradigm to overcome bottleneck of traditional bus-based interconnects for future micro-architectures (MPSoC and CMP). One of current issue in NoC routing is the use of acyclic channel dependency graph (ACDG) for deadlock freedom prohibiting certain routing turns. Thus, ACDG reduces the degree of adaptiveness. In this paper, we propose a novel nonminimal turn model which allows cycles in channel dependency graph provided that extended channel dependency graph is acyclic. Proposed turn model reduces number of restrictions on routing turns (specially on 90-degree), hence able to provide additional minimal and non-minimal routes between source and destination. We also propose a non-minimal and congestion-aware adaptive routing algorithm based on proposed turn model to demonstrate advantages. From results, we can observe that proposed method improves the network performance by distributing the traffic load in the non-congested regions. Manoj Kumar 0001, Vijay Laxmi, Manoj Singh Gaur, Masoud Daneshtalab, Mark Zwolinski |
VLSI-SoC | 4 |
| 2014 | Path-Based Partitioning Methods for 3D Networks-on-Chip with Minimal Adaptive RoutingabstractCombining the benefits of 3D ICs and Networks-on-Chip (NoCs) schemes provides a significant performance gain in Chip Multiprocessors (CMPs) architectures. As multicast communication is commonly used in cache coherence protocols for CMPs and in various parallel applications, the performance of these systems can be significantly improved if multicast operations are supported at the hardware level. In this paper, we present several partitioning methods for the path-based multicast approach in 3D mesh-based NoCs, each with different levels of efficiency. In addition, we develop novel analytical models for unicast and multicast traffic to explore the efficiency of each approach. In order to distribute the unicast and multicast traffic more efficiently over the network, we propose the Minimal and Adaptive Routing (MAR) algorithm for the presented partitioning methods. The analytical and experimental results show that an advantageous method named Recursive Partitioning (RP) outperforms the other approaches. RP recursively partitions the network until all partitions contain a comparable number of switches and thus the multicast traffic is equally distributed among several subsets and the network latency is considerably decreased. The simulation results reveal that the RP method can achieve performance improvement across all workloads while performance can be further improved by utilizing the MAR algorithm. Nineteen percent average and 42 percent maximum latency reduction are obtained on SPLASH-2 and PARSEC benchmarks running on a 64-core CMP. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, José Flich, Hannu Tenhunen |
IEEE Trans. Computers | 2 |
| 2014 | Editorial: Special issue on design challenges for many-core processorsabstractNo abstract available. Masoud Daneshtalab, Maurizio Palesi, Juha Plosila |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2014 | On self-tuning networks-on-chip for dynamic network-flow dominance adaptationabstractModern network-on-chip (NoC) systems are required to handle complex runtime traffic patterns and unprecedented applications. Data traffics of these applications are difficult to fully comprehend at design time so as to optimize the network design. However, it has been discovered that the majority of dataflows in a network are dominated by less than 10% of the specific pathways. In this article, we introduce a method that is capable of identifying critical pathways in a network at runtime and can then dynamically reconfigure the network to optimize for network performance subject to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerging dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance dataflows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulations. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications. Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Peng Liu 0016, Masoud Daneshtalab, Maurizio Palesi, Terrence S. T. Mak |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2014 | Adaptive load balancing in learning-based approaches for many-core embedded systems
Fahimeh Farahnakian, Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila |
J. Supercomput. | 3 |
| 2013 | MD: Minimal path-based fault-tolerant routing in on-Chip NetworksabstractThe communication requirements of many-core embedded systems are convened by the emerging Network-on-Chip (NoC) paradigm. As on-chip communication reliability is a crucial factor in many-core systems, the NoC paradigm should address the reliability issues. Using fault-tolerant routing algorithms to reroute packets around faulty regions will increase the packet latency and create congestion around the faulty region. On the other hand, the performance of NoC is highly affected by the network congestion. Congestion in the network can increase the delay of packets to route from a source to a destination, so it should be avoided. In this paper, a minimal and defect-resilient (MD) routing algorithm is proposed in order to route packets adaptively through the shortest paths in the presence of a faulty link, as long as a path exists. To avoid congestion, output channels can be adaptively chosen whenever the distance from the current to destination node is greater than one hop along both directions. In addition, an analytical model is presented to evaluate MD for two-faulty cases. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Farhad Mehdipour |
ASP-DAC | 2 |
| 2013 | Smart hill climbing for agile dynamic mapping in many-core systemsabstractStochastic hill climbing algorithm is adapted to rapidly find the appropriate start node in the application mapping of network-based many-core systems. Due to highly dynamic and unpredictable workload of such systems, an agile run-time task allocation scheme is required. The scheme is desired to map the tasks of an incoming application at run-time onto an optimum contiguous area of the available nodes. Contiguous and un-fragmented area mapping is to settle the communicating tasks in close proximity. Hence, the power dissipation, the congestion between different applications, and the latency of the system will be significantly reduced. To find an optimum region, we first propose an approximate model that quickly estimates the available area around a given node. Then the stochastic hill climbing algorithm is used as a search heuristic to find a node that has the required number of available nodes around it. Presented agile climber takes the steps using an adapted version of hill climbing algorithm named Smart Hill Climbing, SHiC, which takes the runtime status of the system into account. Finally, the application mapping is performed starting from the selected first node. Experiments show significant gain in the mapping contiguousness which results in better network latency and power dissipation, compared to state-of-the-art works. Mohammad Fattah, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila |
DAC | 2 |
| 2013 | CARS: congestion-aware request scheduler for network interfaces in NoC-based manycore systemsabstractNetwork congestion is a critical issue of memory parallelism in network-based manycore systems where multiple memories can be accessed simultaneously. Therefore, a congestion-aware method is necessitated to deal with the network congestion. In this paper, we present a streamlined method in order to reduce the network congestion. The idea is to use the global congestion information as a metric in network interfaces to reduce the congestion level of highly congested areas. Network interfaces connected to memory modules are equipped with an adaptive scheduler using the global congestion information to reduce additional traffic to congested areas. Experimental results with synthetic test cases demonstrate that the on-chip network utilizing the proposed adaptive scheduler presents up to 23% improvement in average latency. Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen |
DATE | 1 |
| 2013 | Fault-tolerant routing algorithm for 3D NoC using Hamiltonian path strategyabstractWhile Networks-on-Chip (NoC) have been increasing in popularity with industry and academia, it is threatened by the decreasing reliability of aggressively scaled transistors. In this paper, we address the problem of faulty elements by the means of routing algorithms. Commonly, fault-tolerant algorithms are complex due to supporting different fault models while preventing deadlock. When moving from 2D to 3D network, the complexity increases significantly due to the possibility of creating cycles within and between layers. In this paper, we take advantages of the Hamiltonian path to tolerate faults in the network. The presented approach is not only very simple but also able to support almost all one-faulty unidirectional links in 2D and 3D NoCs. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila |
DATE | 2 |
| 2013 | Minimal-path fault-tolerant approach using connection-retaining structure in Networks-on-ChipabstractThere are many fault-tolerant approaches presented both in off-chip and on-chip networks. Regardless of all varieties, there has always been a common assumption between them. Most of all known fault-tolerant methods are based on rerouting packets around faults. Rerouting might take place through nonminimal paths which affect the performance significantly not only by taking longer paths but also by creating hotspot around a fault. In this paper, we present a fault-tolerant approach based on using the shortest paths. This method maintains the performance of Networks-on-Chip in the presence of faults. To avoid using non-minimal paths, the router architecture is slightly modified. In the new form of architecture, there is an ability to connect the horizontal and orthogonal links of a faulty router such that healthy routers are kept connected to each other. Based on this architecture, a fault-tolerant routing algorithm is presented which is obviously much simpler than traditional fault-tolerant routing algorithms. According to this algorithm, only the shortest paths are used by packets in the presence of fault. This results retains the performance of NoCs in faulty situations. This algorithm is highly reliable, for an instance, the reliability is more than 99.5% when there are six faulty routers in an 8×8 mesh network. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen |
NOCS | 2 |
| 2013 | On self-tuning networks-on-chip for dynamic network-flow dominance adaptationabstractModern networks-on-chip (NoC) systems are required to handle complex run-time traffic patterns and unprecedented applications. Data traffics of these applications are difficult to be fully comprehended at design-time so as to optimize the network design. However, it has been discovered that the majority data flows in a network are dominated by less than 10% of the specific pathways. In this paper, we introduce a method that is capable of identifying critical pathways in a network at run-time and, then, can dynamically reconfigure the network to optimize for the network performance subjected to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerged dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance data flows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulators. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications. Xiaohang Wang 0001, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi |
NOCS | 5 |
| 2013 | DyXYZ: Fully Adaptive Routing Algorithm for 3D NoCsabstractTraditional methods in 3D NoCs simply use a deterministic routing algorithm to deliver packets from a source to a destination node. However, deterministic methods are unable to distribute the traffic load over the network, which results in degrading the performance. In this paper, we present a fully adaptive routing algorithm for 3D NoCs, named DyXYZ. In DyXYZ, the congestion information at the input buffer of the neighboring routers is used as congestion metric to select among the output channels. This algorithm is proven to be deadlock free by using 4, 4, and 2 virtual channels along the X, Y, and Z dimensions, respectively. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen |
PDP | 3 |
| 2013 | High Performance Fault-Tolerant Routing Algorithm for NoC-Based Many-Core SystemsabstractNetworks-on-Chip (NoCs) has become a promising approach for the on-chip communication infrastructure of many-core Systems-on-Chip (SoCs). Faults may occur in the NoC both at the router and link level. There are many fault-tolerant approaches presented both in the off-chip and on-chip networks. Some approaches disable some healthy components in order to form a specific shape and others not. Regardless of all varieties, there has always been a common assumption among them. Most of all traditional fault-tolerant methods are based on rerouting packets around a faulty node or region. These approaches affect the performance significantly not only by taking longer paths but also by creating hotspot around a fault. The focus of this paper is to maintain the performance of NoC in the presence of faults. The presented method takes advantage of a fully adaptive routing algorithm using one and two virtual channels along the X and Y dimensions. This method is able to tolerate all cases of one-faulty node without losing the performance of NoC. According to the experimental results, this presented fault-tolerant routing algorithm is able to support up to six faulty nodes in the 8×8 mesh network by up to 98% reliability. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila |
PDP | 2 |
| 2013 | Cluster-based topologies for 3D Networks-on-Chip using advanced inter-layer bus architecture
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Comput. Syst. Sci. | 2 |
| 2013 | A systematic reordering mechanism for on-chip networks using efficient congestion-aware method
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Syst. Archit. | 1 |
| 2013 | Special issue on network-based many-core embedded systems
Masoud Daneshtalab, Pasi Liljeberg, Mehdi Modarressi, Leandro Soares Indrusiak |
J. Syst. Archit. | 1 |
| 2012 | CATRA- congestion aware trapezoid-based routing algorithm for on-chip networksabstractCongestion occurs frequently in Networks-on-Chip when the packets demands exceed the capacity of network resources. Congestion-aware routing algorithms can greatly improve the network performance by balancing the traffic load in adaptive routing. Commonly, these algorithms either rely on purely local congestion information or take into account the congestion conditions of several nodes even though their statuses might be out-dated for the source node, because of dynamically changing congestion conditions. In this paper, we propose a method to utilize both local and non-local network information to determine the optimal path to forward a packet. The non-local information is gathered from the nodes that not only are more likely to be chosen as intermediate nodes in the routing path but also provide up-to-date information to a given node. Moreover, to collect and deliver the non-local information, a distributed propagation system is presented. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DATE | 2 |
| 2012 | MAFA: Adaptive Fault-Tolerant Routing Algorithm for Networks-on-ChipabstractWhile Networks-on-Chip have been increasing in popularity with industry and academia, it is threatened by the decreasing reliability of aggressively scaled transistors. This level of failure has architectural level ramifications, as it may cause an entire on-chip network to fail. Traditional fault-tolerant routing algorithms can overcome the faulty links or routers by rerouting packets around faulty regions. These approaches increase the packet latency and create congestion around the faulty region. In this paper, we present a novel fault-tolerant method that is able to route packets through shortest paths in the presence of faulty links, as long as a path exists. Although the same idea can be applied to a network with any number of virtual channels, we utilize two virtual channels to tolerate all one and two faulty links. Finally, the method is extended to support multiple faulty links by fully utilizing all allowable turns in the network. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen |
DSD | 2 |
| 2012 | NoC-AXI interface for FPGA-based MPSoC platformsabstractStreaming applications are a keystone in several emerging multimedia services like DVB-IPTV, VoD and on-line gaming. Due to the high computing requirements and real-time constraints inherent to this kind of applications multi-processor system-on-chip (MPSoCs) have been proposed as a solution. In addition, the FPGA technology has become popular among systems-on-chip (SoCs) designers due to its low development cost and short time to market. Here we present a FPGA-based MPSoC platform for streaming applications where the important component of this platform is the AXI interface. Marco Ramírez 0001, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg |
FPL | 2 |
| 2012 | CoNA: Dynamic application mapping for congestion reduction in many-core systemsabstractIncreasing the number of processors in a single chip toward network-based many-core systems requires a run-time task allocation algorithm. We propose an efficient mapping algorithm that assigns communicating tasks of incoming applications onto resources of a many-core system utilizing Network-on-Chip paradigm. In our contiguous neighborhood allocation (CoNA) algorithm, we target at the reduction of both internal and external congestion due to detrimental impact of congestion on the network performance. We approach the goal by keeping the mapped region contiguous and placing the communicating tasks in a close neighborhood. A completely synthesizable simulation environment where none of the system objects are assumed to be ideal is provided. Experiments show at least 40% gain in different mapping cost functions, as well as 16% reduction in average network latency compared to existing algorithms. Mohammad Fattah, Marco Ramírez 0001, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila |
ICCD | 3 |
| 2012 | HARAQ: Congestion-Aware Learning Model for Highly Adaptive Routing Algorithm in On-Chip NetworksabstractThe occurrence of congestion in on-chip networks can severely degrade the performance due to increased message latency. In mesh topology, minimal methods can propagate messages over two directions at each switch. When shortest paths are congested, sending more messages through them can deteriorate the congestion condition considerably. In this paper, we present an adaptive routing algorithm for on-chip networks that provide a wide range of alternative paths between each pair of source and destination switches. Initially, the algorithm determines all permitted turns in the network including 180-degree turns on a single channel without creating cycles. The implementation of the algorithm provides the best usage of all allowable turns to route messages more adaptively in the network. On top of that, for selecting a less congested path, an optimized and scalable learning method is utilized. The learning method is based on local and global congestion information and can estimate the latency from each output channel to the destination region. Masoumeh Ebrahimi, Masoud Daneshtalab, Fahimeh Farahnakian, Juha Plosila, Pasi Liljeberg, Maurizio Palesi, Hannu Tenhunen |
NOCS | 2 |
| 2012 | LEAR - A Low-Weight and Highly Adaptive Routing Method for Distributing Congestions in On-chip NetworksabstractCongestion-aware routing algorithms can improve network throughput by avoiding packets to be routed through congested areas. In this paper, we propose a minimal/non-minimal routing algorithm to alleviate congestion in the network by making use of all available paths between sources and destinations. The simplicity of the proposed algorithm provides a cost and power efficient solution for Networks-on-Chip while the high degree of adaptive ness, achieved by using an additional virtual channel along the Y dimension, leads to an increased performance. In this method, different restrictions are imposed on the use of each virtual channel, so that the prohibited turns in one virtual channel are permitted in the other one. By fully exploiting of the eligible turns in the network, a large number of output channels can be provided by the proposed method. Based on this method, a packet is routed along the non-minimal path when the neighboring routers in the minimal path are congested. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 2 |
| 2012 | Memory-Efficient On-Chip Network With Adaptive InterfacesabstractTo achieve higher memory bandwidth in network-based multiprocessor architectures, multiple dynamic random access memories can be accessed simultaneously. In such architectures, not only resource utilization and latency are the critical issues but also a reordering mechanism is required to deliver the response transactions of concurrent memory accesses in-order. In this paper, we present a memory-efficient on-chip network architecture to cope with these issues efficiently. Each node of the network is equipped with a novel network interface (NI) to deal with out-of-order delivery, and a priority-based router to decrease the network latency. The proposed NI exploits a streamlined reordering mechanism to handle the in-order delivery and utilizes the advance extensible interface transaction-based protocol to maintain compatibility with existing intellectual property cores. To improve the memory utilization and reduce the memory latency, an optimized memory controller is integrated in the presented NI. Experimental results with synthetic test cases demonstrate that the proposed on-chip network architecture provides significant improvements in average network latency (16%), average memory access latency (19%), and average memory utilization (22%). Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2011 | Formal Modeling of Multicast Communication in 3D NoCsabstractA reliable approach to designing systems is by applying formal methods, based on logics and set theory. In formal methods refinement based, we develop the system models stepwise, from an abstract level to a concrete one by gradually adding details. Each detail-adding level is proved to still validate the properties of the more abstract level. Due to the high complexity and the high reliability requirements of 3D NoCs, formal methods provide promising solutions for modeling and verifying their communication schemes. In this paper, we present a general model for specifying the 3D NoC multicast communication scheme. We then refine our model to two communication schemes, : unicast and multicast, via the XYZ routing algorithm in order to put forward the correct-by-construction concrete models. Maryam Kamali, Luigia Petre, Kaisa Sere, Masoud Daneshtalab |
DSD | 4 |
| 2011 | Exploring partitioning methods for 3D Networks-on-Chip utilizing adaptive routing modelabstractThree-Dimensional (3D) integration is a solution to the interconnect bottleneck in Two-Dimensional (2D) MultiProcessor System on Chip (MPSoC). 3D IC design improves performance and decreases power consumption by replacing long horizontal interconnects with shorter vertical ones. As the multicast communication is utilized commonly in various parallel applications, the performance can be significantly improved by supporting of multicast operations at the hardware level. In this paper, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method named Recursive Partitioning (RP) in which the network is recursively partitioned until all partitions contain comparable number of nodes. By this approach, the multicast traffic is distributed among several subsets and the network latency is considerably decreased. We also present Minimal Adaptive Routing (MAR) algorithm for the unicast and multicast traffic in 3D-mesh Networks-on-Chip (NoCs). The idea behind the MAR algorithm is utilizing the Hamiltonian path to provide a set of alternative paths. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 2 |
| 2011 | Agent-based on-chip network using efficient selection methodabstractCongestion in on-chip networks may cause many drawbacks in multiprocessor systems including throughput reduction, increase in latency, and additional power consumption. Furthermore, conventional congestion control methods, employed for on-chip networks, cannot efficiently collect congestion information and distribute them over the on-chip network. In this paper, we present a novel structure for on-chip networks, named Agent-based Network-on-Chip (ANoC), to diagnose the congested areas. In addition to the presented structure, an efficient Congestion-Aware Selection (CAS) method is proposed to reduce overall network latency. CAS is capable of selecting an appropriate output channel to route packets along a less congested path. 29% average and 35% maximum latency reduction are achieved on SPLASH-2 and PARSEC benchmarks running on a 36-core Chip Multi-Processor. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
VLSI-SoC | 2 |
| 2011 | A generic adaptive path-based routing method for MPSoCs
Masoud Daneshtalab, Masoumeh Ebrahimi, Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen |
J. Syst. Archit. | 1 |
| 2010 | Partitioning methods for unicast/multicast traffic in 3D NoC architectureabstractAs the scale of integration grows, the interconnection problem becomes one of the major design considerations of Multi Processor System on Chip (MPSoC). In recent years, many researchers have conducted studies on 3D IC designs stacking multiple layers on top of each other. In order to decrease the transmission delay of unicast/multicast messages in a network based multicore system, the network is divided into several partitions. In this paper, we first introduce a novel idea of balanced partitioning that allows the network to be partitioned effectively. Then, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method based on the idea of balanced partitioning to provide a high degree of parallelism with a considerable reduction of packet delay in unicast/multicast traffic. Simulations are provided to evaluate and compare the performance of proposed methods. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
DDECS | 2 |
| 2010 | Power-aware NoC router using central forecasting-based dynamic virtual channel allocationabstractIn this paper, we propose a high performance central dynamic virtual channel allocation mechanism for on-chip routers. This central management unit devotes each input port a number of virtual channels (VC) among a shared VC bank based on a traffic forecasting technique. The forecasting technique exploits the link and VC utilizations in predicting the traffic. Based on the predicted traffic, for each input port, the number of active virtual channels may be increased, decreased, or kept unchanged. The clock-gating power management technique is used to activate/deactivate the VCs. Simulation results using uniform and Negative Exponential Distribution (NED) traffic profiles show that a considerable power savings in the virtual channels and overall router power consumption may be achieved especially in low traffic loads. The area overhead of the technique is negligible. Amir-Mohammad Rahmani, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
ISCAS | 2 |
| 2010 | A Low-Latency and Memory-Efficient On-chip NetworkabstractUsing multiple SDRAMs in MPSoCs and NoCs to increase memory parallelism is very common nowadays. In-order delivery, resource utilization, and latency are the most critical issues in such architectures. In this paper, we present a novel network interface architecture to cope with these issues efficiently. The proposed network interface exploits a resourceful reordering mechanism to handle the in-order delivery and to increase the resource utilization. A brilliant memory controller is efficiently integrated into this network interface to improve the memory utilization and reduce both memory and network latencies. In addition, to bring compatibility with existing IP cores the proposed network interface utilizes AXI transaction based protocol. Experimental results with synthetic test cases demonstrate that the proposed architecture gives significant improvements in average network latency (12%), average memory access latency (19%), and average memory utilization (22%). Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 1 |
| 2010 | A High-Performance Network Interface Architecture for NoCs Using Reorder Buffer SharingabstractIncreasing memory parallelism in MPSoCs to provide higher memory bandwidth is achieved by accessing multiple memories simultaneously. Inasmuch as the response transactions of concurrent memory accesses must be in-order, a reordering mechanism is required. To our knowledge the resource utilization of conventional reordering mechanisms is low. In this paper, we present a novel network interface architecture for on-chip networks to increase the resource utilization and to improve overall performance. Also, based on the proposed architecture, a hybrid network interface is presented to integrate both memory and processor in a tile. The proposed architecture exploits AXI transaction based protocol to be compatible with existing IP cores. Experimental results with synthetic test cases demonstrate that the proposed architecture outperforms the conventional architecture in terms of latency. Also, the cost of the presented architecture is evaluated with UMC 0.09 ¿ m technology. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 2 |
| 2010 | HAMUM - A Novel Routing Protocol for Unicast and Multicast Traffic in MPSoCsabstractMany parallel applications in MPSoCs take advantage of multicast communication. Several multicast schemes such as path-based, tree-based, and unicast-based have been proposed in interconnection networks. Path-based multicast scheme has been proven to be more efficient than the other schemes in on-chip interconnection network. A new adaptive routing model based on Hamiltonian path for both the multicast and unicast traffics, called Hamiltonian Adaptive Multicast and Unicast Model (HAMUM), is presented. Results obtained in both multicast and mixed traffic models show that the proposed adaptive algorithm for multicast aspect has lower latency and power dissipation compared to previously proposed path-based multicasting algorithms with less than 0.5% hardware overhead. Additionally, for the unicast aspect the proposed adaptive model outperforms the other unicast turn models. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
PDP | 2 |
| 2010 | EDXY - A low cost congestion-aware routing algorithm for network-on-chips
Pejman Lotfi-Kamran, Amir-Mohammad Rahmani, Masoud Daneshtalab, Ali Afzali-Kusha, Zainalabedin Navabi |
J. Syst. Archit. | 3 |
| 2009 | An efficent dynamic multicast routing protocol for distributing traffic in NOCsabstractNowadays, in MPSoCs and NoCs, multicast protocol is significantly used for many parallel applications such as cache coherency in distributed shared-memory architectures, clock synchronization, replication, or barrier synchronization. Among several multicast schemes proposed in on chip interconnection networks, path-based multicast scheme has been proven to be more efficient than the tree-based, and unicast-based. In this paper a low distance path-based multicast scheme is proposed. The proposed method takes advantage of the network partitioning, and utilizing of an efficient destination ordering algorithm. The results in performance, and power consumption show that the proposed method outstands the previous on chip path-based multicasting algorithms. Masoumeh Ebrahimi, Masoud Daneshtalab, Mohammad Hossein Neishaburi, Siamak Mohammadi, Ali Afzali-Kusha, Juha Plosila, Hannu Tenhunen |
DATE | 2 |
| 2009 | An Adaptive Unicast/Multicast Routing Algorithm for MPSoCsabstractSeveral parallel applications in MPSoCs take advantage of multicast communication. Path-based multicast scheme has been proven to be more efficient than the others multicast schemes in on-chip interconnection network. We present a new adaptive path based model for both the multicast and unicast wormhole routing protocols. The proposed model under mixed traffic models has lower latency than the previous path-based methods with negligible hardware overhead. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
DSD | 2 |
| 2008 | BARP-A Dynamic Routing Protocol for Balanced Distribution of Traffic in NoCsabstractA novel routing algorithm, named balanced adaptive routing protocol (BARP), is proposed for NoCs to provide adaptive routing and ensure deadlock-free and livelock-free routing at the same time. By evenly distributing input packets of a router among all its shortest path output ports, a novel adaptive routing protocol for avoiding congestion condition emerges. It is observed that BARP can achieve better performance compared to static XY routing, odd- even routing and dynamic XY routing. Pejman Lotfi-Kamran, Masoud Daneshtalab, Caro Lucas, Zainalabedin Navabi |
DATE | 2 |
| 2007 | Improving Robustness of Real-Time Operating Systems (RTOS) Services Related to Soft-ErrorsabstractNowadays, more critical applications that have stringent real-time constraint are placed and run in an environment with real-time operating system (RTOS). The provided services of RTOSs are subject to faults that affect both functional and timing of Tasks which are running based on RTOS. In this paper, we try to evaluate and analyze robustness of services due to soft-errors in two proposed architecture of RTOS which are (SW-RTOS and HW/SW- RTOS). According to experimental result we finally propose an architecture which provides more robust services in term of soft-error. Real-Time Operating System (RTOS) users desire predictable response time at an affordable cost, due to this demand Hardware/Software Real-Time Operating Systems (HW/SW-RTOS) appeared. This paper analyzes the impact of soft-errors in real-time systems running applications under purely Software RTOS versus HW/SW-RTOS. The proposed model is used to evaluate robustness of services like scheduling, synchronization time management and memory management and inter process communication in Software based RTOS and HW/SW-RTOS. Experimental results show HW/SW-RTOS provide more robust services in term of soft-err or against purely software based RTOS. Mohammad Hossein Neishaburi, Masoud Daneshtalab, Mohammad Reza Kakoee, Saeed Safari |
AICCSA | 2 |
| 2007 | System Level Voltage Scheduling Technique Using UML-RT ModelabstractIn this paper, we present optimized methodology for Intra-task voltage scheduling. Our proposed method gets data flow and control flow of application that represents coloration between different parts of the application at the early stage of design using UML-RT model and decides to schedule processor's voltage. By applying this technique on JPEG encoder system experimental results show reduction in energy consumption by 18-54 % over common Intra-DVS algorithm. Mohammad Hossein Neishaburi, Masoud Daneshtalab, Majid Nabi, Siamak Mohammadi |
AICCSA | 2 |
| 2007 | On-Chip Verification of NoCs Using Assertion ProcessorsabstractNoC verification has become an increasingly difficult task due to the growing complexity of these systems. In this paper, we propose a methodology based on assertions for on-chip verification of NoCs. We have a local assertion processor (LAP) in each core to manage the outputs of assertions inside it. To route assertions' outputs toward this processor, we offer a boundary scan chain mechanism. Moreover, after detecting error, each LAP dispatches a packet called error packet to a global assertion processor (GAP) which receives error packets from all cores and performs necessary actions regarding to errors and their severities. Finally, in order to evaluate our method, we apply it on an NoC structure and show the experimental results. Mohammad Reza Kakoee, Mohammad Hossein Neishaburi, Masoud Daneshtalab, Saeed Safari, Zainalabedin Navabi |
DSD | 3 |
| 2006 | NoC Hot Spot minimization Using AntNet Dynamic Routing AlgorithmabstractIn this paper, a routing model for minimizing hot spots in the network on chip (NOC) is presented. The model makes use of AntNet routing algorithm which is based on Ant colony. Using this algorithm, which we call AntNet routing algorithm, heavy packet traffics are distributed on the chip minimizing the occurrence of hot spots. To evaluate the efficiency of the scheme, the proposed algorithm was compared to the XY, Odd- Even, and DyAD routing models. The simulation results show that in realistic (Transpose) traffic as well as in heavy packet traffic, the proposed model has less average delay and peak power compared to the other routing models. In addition, the maximum temperature in the proposed algorithm is less than those of the other routing algorithms. Masoud Daneshtalab, Ashkan Sobhani, Ali Afzali-Kusha, Omid Fatemi, Zainalabedin Navabi |
ASAP | 1 |