EDBT 2026 Demo / reviewers in the wild / expert
Mehdi Modarressi
dblp:82/5248
· DBLP profile ↗
42ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0002-4117-7609ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 12 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 first-authorArtificial intelligence and machine learning · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ChronoSpike: A low-cost, deterministic circuit-switched NoC for neuromorphic accelerators
Mohammad Rasoul Roshanshah, Saeed Safari, Mehdi Modarressi |
Neurocomputing | 3 |
| 2025 | A novel deep learning-based approach for video quality enhancement
Parham Zilouchian Moghaddam, Mehdi Modarressi, MohammadAmin Sadeghi |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | NeuroPIM: Felxible Neural Accelerator for Processing-in-Memory ArchitecturesabstractThe performance of microprocessors under many modern workloads is mainly limited by the off-chip memory bandwidth. The emerging process-in-memory paradigm present a unique opportunity to reduce data movement overheads by moving computation closer to memory. State-of-the-art processing-in-memory proposals stack a logic layer on top of one or multiple memory layers in a 3D fashion and leverage the logic layer to build near-memory processing units. Such processing units are either application-specific accelerators or general-purpose cores. In this paper, we present NeuroPIM, a new processing-in-memory architecture that uses a neural network as the memory-side general-purpose accelerator. This design is mainly motivated by the observation that in many real-world applications, some program regions, or even the entire program, can be replaced by a neural network that is learned to approximate the program’s output. NeuroPIM benefits from both the flexibility of general-purpose processors and superior performance of application-specific accelerators. Experimental results show that NeuroPIM provides up to 41% speedup over a processor-side neural network accelerator and up to 8x speedup over a general-purpose processor. Ali Monavari Bidgoli, Sepideh Fattahi, Seyyed Hossein Seyyedaghaei Rezaei, Mehdi Modarressi, Masoud Daneshtalab |
DDECS | 4 |
| 2023 | NCOD: Near-Optimum Video Compression for Object DetectionabstractWith the emergence of technologies like smart cities, Internet of things (IoT), and 5G, the amount of produced visual data at the edges and remote nodes has exploded. Since for a considerable portion of the captured video the target is a machine learning task, rather than a human audience, transmission of videos in such applications requires efficient video compression tailored for machine vision. However, existing compression solutions are optimized for human vision. This paper presents a methodology to optimize an existing video compression standard, HEVC, for a machine vision task, Object Detection (OD). To this end, (1) a dataset of compressed videos, including several compression-ratios and their corresponding OD performance is collected to enable modeling, (2) A trade-off point (knee-point) between bitrate and OD performance is defined, that finds the point after which no major improvements will be achieved, (3) a set of features were extracted and studied to model this point, via a practical machine learning method. The resulting solution can predict the knee-point with$\boldsymbol{\text{MAE}=1.28}$, resulting in a ∆Recall of only 0.012 and bitrate reduction of 86.56%, compared to OD with very high-quality video. Ardavan Elahi, Ali Falahati, Farhad Pakdaman, Mehdi Modarressi, Moncef Gabbouj |
ISCAS | 4 |
| 2022 | Multi-Precision Deep Neural Network Acceleration on FPGAsabstractQuantization is a promising approach to reduce the computational load of neural networks. The minimum bit-width that preserves the original accuracy varies significantly across different neural networks and even across different layers of a single neural network. Most existing designs over-provision neural network accelerators with sufficient bit-width to preserve the required accuracy across a wide range of neural networks. In this paper, we present mpDNN, a multi-precision multiplier with dynamically adjustable bit-width for deep neural network acceleration. The design supports run-time splitting an arithmetic operator into multiple independent operators with smaller bit-width, effectively increasing throughput when lower precision is required. The proposed architecture is designed for FPGAs, in that the multipliers and bit-width adjustment mechanism are optimized for the LUT-based structure of FPGAs. Experimental results show that by enabling run-time precision adjustment, mpDNN can offer 3-15x improvement in throughput. Negar Neda, Salim Ullah, Azam Ghanbari, Hoda Mahdiani, Mehdi Modarressi, Akash Kumar 0001 |
ASP-DAC | 5 |
| 2022 | Reconfigurable Network-on-Chip based Convolutional Neural Network Accelerator
Arash Firuzan, Mehdi Modarressi, Midia Reshadi, Ahmad Khademzadeh |
J. Syst. Archit. | 2 |
| 2022 | Energy-efficient acceleration of convolutional neural networks using computation reuse
Azam Ghanbari, Mehdi Modarressi |
J. Syst. Archit. | 2 |
| 2021 | Network-on-ReRAM for Scalable Processing-in-Memory Architecture DesignabstractThe non-volatile metal-oxide resistive random access memory (ReRAM) is an emerging alternative for the current memory technologies. The unique capability of ReRAM to perform analog and digital arithmetic and logic operations has enabled this technology to incorporate both computation and memory capabilities on the same unit. Due to this interesting property, there is a growing trend in recent years to implement emerging data-intensive applications on ReRAM structures. A typical ReRAM-based processing-in-memory architecture may consist tens to hundreds of ReRAM units (mats) that can either store or process data. To support such large-scale ReRAM structure, this paper proposes a scalable network-on-ReRAM architecture. The proposed network employs a novel associative router architecture, designed based on the ReRAM-based content-addressable memories. With the in-memory packet processing capability, this router yields higher throughput and resource utilization levels than a conventional router. This router is technology compatible with ReRAM and as our evaluations show, employing it to build a network-on-ReRAM makes the emerging ReRAM-based processing-in-memory architectures more scalable and performance-efficient. Bita Dabiri, Mehdi Modarressi, Masoud Daneshtalab |
DSD | 2 |
| 2021 | RoCo-NAS: Robust and Compact Neural Architecture SearchabstractDeep model compression has been studied widely, and state-of-the-art methods can now achieve high compression ratios with minimum accuracy loss. Recent advances in adversarial attacks reveal the inherent vulnerability of deep neural networks to slightly perturbed images called adversarial examples. Since then, extensive efforts have been performed to enhance deep networks' robustness via specialized loss functions and learning algorithms. Previous works suggest that network size and robustness against adversarial examples contradict on most occasions. In this paper, we investigate how to optimize compactness and robustness to adversarial attacks of neural network architectures while maintaining the accuracy using multi-objective neural architecture search. We propose the use of previously generated adversarial examples as an objective to evaluate the robustness of our models in addition to the number of floating-point operations to assess model complexity i.e. compactness. Experiments on some recent neural architecture search algorithms show that due to their limited search space they fail to find robust and compact architectures. By creating a novel neural architecture search (RoCo-NAS), we were able to evolve an architecture that is up to 7% more accurate against adversarial samples than its more complex architecture counterpart. Thus, the results show inherently robust architectures regardless of their size. This opens up a new range of possibilities for the exploration and design of deep neural networks using automatic architecture search. Vahid Geraeinejad, Sima Sinaei, Mehdi Modarressi, Masoud Daneshtalab |
IJCNN | 3 |
| 2018 | A Customized Processing-in-Memory Architecture for Biological Sequence AlignmentabstractSequence alignment is the most widely used operation in bioinformatics. With the exponential growth of the biological sequence databases, searching a database to find the optimal alignment for a query sequence (that can be at the order of hundreds of millions of characters long) would require excessive processing power and memory bandwidth. Sequence alignment algorithms can potentially benefit from the processing power of massive parallel processors due their simple arithmetic operations, coupled with the inherent fine-grained and coarse-grained parallelism that they exhibit. However, the limited memory bandwidth in conventional computing systems prevents exploiting the maximum achievable speedup. In this paper, we propose a processing-in-memory architecture as a viable solution for the excessive memory bandwidth demand of bioinformatics applications. The design is composed of a set of simple and lightweight processing elements, customized to the sequence alignment algorithm, integrated at the logic layer of an emerging 3D DRAM architecture. Experimental results show that the proposed architecture results in up to 2.4x speedup and 41% reduction in power consumption, compared to a processor-side parallel implementation. Nasrin Akbari, Mehdi Modarressi, Masoud Daneshtalab, Mohammad Loni |
ASAP | 2 |
| 2018 | Reconfigurable Network-on-Chip for 3D Neural Network AcceleratorsabstractParallel hardware accelerators for large-scale neural networks typically consist of several processing nodes, arranged as a multi- or many-core system-on-chip, connected by a network-on-chip (NoC). Recent proposals also benefit from the emerging 3D memory-on-logic architectures to provide sufficient bandwidth for neural networks. Handling the heavy traffic between neurons and memory and also the multicast-based inter-neuron traffic, which often varies over time, is the most challenging design consideration for the networks-on-chip in such accelerators. To address these issues, a reconfigurable network-on-chip architecture for 3D memory-on-logic neural network accelerators is presented in this paper. The reconfigurable NoC can adapt its topology to the on-chip traffic patterns. It can be also configured as a tree-like structure to support multicast-based neuron-to-neuron and memory-to-neuron traffic of neural networks. The evaluation results show that the proposed architecture can better manage the multicast-based traffic of neural networks than some state-of-the-art topologies and considerably increase throughput and power efficiency. Arash Firuzan, Mehdi Modarressi, Masoud Daneshtalab, Midia Reshadi |
NOCS | 2 |
| 2018 | Fast Data Delivery for Many-Core ProcessorsabstractServer workloads operate on large volumes of data. As a result, processors executing these workloads encounter frequent L1-D misses. In a many-core processor, an L1-D miss causes a request packet to be sent to an LLC slice and a response packet to be sent back to the L1-D, which results in high overhead. While prior work targeted response packets, this work focuses on accelerating the request packets. Unlike aggressive OoO cores, simpler cores used in many-core processors cannot hide the latency of L1-D request packets. We observe that LLC slices that serve L1-D misses are strongly temporally correlated. Taking advantage of this observation, we design a simple and accurate predictor. Upon the occurrence of an L1-D miss, the predictor identifies the LLC slice that will serve the next L1-D miss and a circuit will be set up for the upcoming miss request to accelerate its transmission. When the upcoming miss occurs, the resulting request can use the already established circuit for transmission to the LLC slice. We show that our proposal outperforms data prefetching mechanisms in a many-core processor due to (1) higher prediction accuracy and (2) not wasting valuable off-chip bandwidth, while requiring significantly less overhead. Using full-system simulation, we show that our proposal accelerates serving data misses by 22 percent and leads to 10 percent performance improvement over the state-of-the-art network-on-chip. Mohammad Bakhshalipour, Pejman Lotfi-Kamran, Abbas Mazloumi, Farid Samandi, Mahmood Naderan-Tahan, Mehdi Modarressi, Hamid Sarbazi-Azad |
IEEE Trans. Computers | 6 |
| 2017 | Near-Ideal Networks-on-Chip for ServersabstractServer workloads benefit from execution on many-core processors due to their massive request-level parallelism. A key characteristic of server workloads is the large instruction footprints. While a shared last-level cache (LLC) captures the footprints, it necessitates a low-latency network-on-chip (NOC) to minimize the core stall time on accesses serviced by the LLC. As strict quality-of-service requirements preclude the use of lean cores in server processors, we observe that even state-of-the-art single-cycle multi-hop NOCs are far from ideal because they impose significant NOC-induced delays on the LLC access latency, and diminish performance. Most of the NOC delay is due to per-hop resource allocation. In this paper, we take advantage of proactive resource allocation (PRA) to eliminate per-hop resource allocation time in single-cycle multi-hop networks to reach a near-ideal network for servers. PRA is undertaken during (1) the time interval in which it is known that LLC has the requested data, but the data is not yet ready, and (2) the time interval in which a packet is stalled in a router because the required resources are dedicated to another packet. Through detailed evaluation targeting a 64-core processor and a set of server workloads, we show that our proposal improves system performance by 12% over the state-of-the-art single-cycle multi-hop mesh NOC. Pejman Lotfi-Kamran, Mehdi Modarressi, Hamid Sarbazi-Azad |
HPCA | 2 |
| 2017 | Customizing Clos Network-on-Chip for Neural NetworksabstractLarge-scale neural network accelerators are often implemented as a many-core chip and rely on a network-on-chip to manage the huge amount of inter-neuron traffic. The baseline and different variations of the well-known mesh and tree topologies are the most popular topologies in prior many-core implementations of neural networks. However, the grid-like mesh and hierarchical tree topologies suffer from high diameter and low bisection bandwidth, respectively. In this paper, we present ClosNN, a customized Clos topology for Neural Networks. The inherent capability of Clos to support multicast and broadcast traffic in a simple and efficient way, as well as its adaptable bisection bandwidth, is the major motivation behind proposing a customized version of this topology as the communication infrastructure of large-scale neural network implementations. We compare ClosNN with some state-of-the-art NoC topologies adopted in recent neural network hardware accelerators and show that it offers lower average message hop count and higher throughput, which directly translates to faster neural information processing. Reza Hojabr, Mehdi Modarressi, Masoud Daneshtalab, Ali Yasoubi, Ahmad Khonsari |
IEEE Trans. Computers | 2 |
| 2016 | Fault-tolerant 3-D network-on-chip design using dynamic link sharing
Seyyed Hossein Seyyedaghaei Rezaei, Mehdi Modarressi, Reza Yazdani, Masoud Daneshtalab |
DATE | 2 |
| 2016 | Low-Power Online ECG Analysis Using Neural NetworksabstractDigital healthcare devices have attracted many attentions in recent years. In particular, the cardiovascular system is always an important target of such digital healthcare devices. In this paper, we present a neural network-based low-power online ECG analysis and monitoring system. Electrocardiogram (ECG) is a representation of the heart's electrical activity recorded from electrodes on the body surface and serves as an important source of assessing heart's function. The proposed low-power design is based on the observation that a considerable amount of computations in neural network-based ECG monitoring is repetitive and can be eliminated to save power. We identify two sources of computation redundancy: the input data (which have relatively little change, as long as the heart functions properly) and neuron weights. By exploiting computation reuse to eliminate redundant computation the proposed architecture can achieve high power-efficiency for online ECG analysis which is crucial for an implant or wearable healthcare device. Experimental results on some MIT-BIH database ECG waveforms show that the proposed mechanism can reduce ECG analysis power consumption by 45%. Mehdi Modarressi, Ali Yasoubi, Maryam Modarressi |
DSD | 1 |
| 2016 | A Three-Dimensional Networks-on-Chip Architecture with Dynamic Buffer Sharingabstract3D integration is a practical solution for overcoming the failure of Dennard scaling in future technology generations. This emerging technology stacks several die slices on top of each other on a single chip in order to provide higher-bandwidth and lower-latency than a 2D design due to extremely shorter inter-layer distances in the third dimension and. In this paper, we leverage the low-latency vertical links to address buffer management, one of the most important design and management issues in Network-on-Chip (NoC). To this end, we present VerBuS, an architecture for 3D routers with Vertical BUffer Sharing capability enabled by ultra-low latency vertical links of a 3D chip. VerBuS can share virtual channels (VC) between vertically stacked routers. This way, the buffering capacity of a highly loaded router is increased by using idle VCs of vertically adjacent routers. Experimental results show up to 20% improvement in NoC performance metrics over state-of-the-art 3D router designs. Seyyed Hossein Seyyedaghaei Rezaei, Mehdi Modarressi, Masoud Daneshtalab, Shervin Roshanisefat |
PDP | 2 |
| 2016 | An Efficient Hybrid-Switched Network-on-Chip for Chip MultiprocessorsabstractChip multiprocessors (CMPs) require a low-latency interconnect fabric network-on-chip (NoC) to minimize processor stall time on instruction and data accesses that are serviced by the last-level cache (LLC). While packet-switched mesh interconnects sacrifice performance of many-core processors due to NoC-induced delays, existing circuit-switched interconnects do not offer lower network delays as they cannot hide the time it takes to set up a circuit. To address this problem, this work introduces CIMA-a hybrid circuit-switched and packet-switched mesh-based interconnection network that affords low LLC access delays at a small area cost. CIMA uses virtual cut-through (VCT) switching for short request packets, and benefits from circuit switching for longer, delay sensitive response packets. While a request is being served by the LLC, CIMA attempts to set up a circuit for the corresponding response packet. By the time the request packet is served and the response gets ready, a circuit has already been prepared, and as a result, the response packet experiences short delay in the network. A detailed evaluation targeting a 64-core CMP running scale-out workloads reveals that CIMA improves system performance by 21 percent over the state-of-the-art hybrid circuit-packet-switched network. Pejman Lotfi-Kamran, Mehdi Modarressi, Hamid Sarbazi-Azad |
IEEE Trans. Computers | 2 |
| 2015 | A hybrid packet/circuit-switched router to accelerate memory access in NoC-based chip multiprocessors
Abbas Mazloumi, Mehdi Modarressi |
DATE | 2 |
| 2015 | An energy-efficient virtual channel power-gating mechanism for on-chip networks
Amirhossein Mirhosseini, Mohammad Sadrosadati, Ali Fakhrzadehgan, Mehdi Modarressi, Hamid Sarbazi-Azad |
DATE | 4 |
| 2015 | CuPAN - High Throughput On-chip Interconnection for Neural Networks
Ali Yasoubi, Reza Hojabr, Hengameh Takshi, Mehdi Modarressi, Masoud Daneshtalab |
ICONIP (3) | 4 |
| 2015 | Improving the performance of packet-switched networks-on-chip by SDM-based adaptive shortcut paths
Mehdi Modarressi, Nasibeh Teimouri, Hamid Sarbazi-Azad |
Integr. | 1 |
| 2015 | Leveraging dark silicon to optimize networks-on-chip topology
Mehdi Modarressi, Hamid Sarbazi-Azad |
J. Supercomput. | 1 |
| 2015 | Integrated circuit-packet switching NoC with efficient circuit setup mechanism
Farhad Pakdaman, Abbas Mazloumi, Mehdi Modarressi |
J. Supercomput. | 3 |
| 2014 | A reconfigurable network-on-chip architecture for heterogeneous CMPs in the dark-silicon eraabstractCore specialization is a promising solution to the dark silicon challenge. This approach trades off the cheaper silicon area with energy-efficiency by integrating a selection of many diverse application-specific cores into a single billion-transistor multicore chip. Each application then activates the subset of cores that best matches its processing requirements. These cores act as a customized application-specific CMP for the application. Such an arrangement of cores requires some special on-chip inter-core communication treatment to efficiently connect active cores. In this paper, we propose a reconfigurable network-on-chip that leverages the routers of the dark portion of the chip to customize the topology for the powered cores at any time. To this end, routers of the dark parts of the chip are used as bypass switches that can directly connect distant active nodes in the network. Our experimental results show considerable reduction in energy consumption and latency of on-chip communication. Mehdi Modarressi, Hamid Sarbazi-Azad |
ASAP | 1 |
| 2013 | Power and Performance Efficient Partial Circuits in Packet-Switched Networks-on-ChipabstractIn this paper, we propose a hybrid packet-circuit switching for networks-on-chip to benefit from the advantages of both switching mechanisms. Integrating circuit and packet switching into a single NoC is achieved by partitioning the link bandwidth and router data-path and control-path elements into two parts and allocating each part to one of the switching methods. In this NoC, during injection in the source node, packets are initially forwarded on the packet-switched sub-network, but keep requesting a circuit towards the destination node. The circuit-switched part, at each cycle, collects the circuit construction requests, performs arbitration among the conflicting requests, and constructs circuits over the unallocated circuit-switched sub-network links. Unlike traditional circuit-switching, the circuit end point in this NoC is not necessarily the packet destination, rather the circuits can be terminated in any intermediate node between the packet source and destination nodes. At that node, the packet may either travel over another circuit (in case of successful circuit request) or continue its path over the packet-switched part. Therefore, packets may switch between the two sub-networks several times during their life-time in the network. Circuit construction is handled by a low-latency and low-cost setup network. To keep the complexity of the circuit construction low, the circuits are restricted to span within a neighborhood of d hops of the requesting node. The experimental results show considerable improvement in energy and latency over a traditional packet-switched NoC. Nasibeh Teimouri, Mehdi Modarressi, Hamid Sarbazi-Azad |
PDP | 2 |
| 2013 | Special issue on network-based many-core embedded systems
Masoud Daneshtalab, Pasi Liljeberg, Mehdi Modarressi, Leandro Soares Indrusiak |
J. Syst. Archit. | 3 |
| 2013 | Using task migration to improve non-contiguous processor allocation in NoC-based CMPs
Mehdi Modarressi, Marjan Asadinia, Hamid Sarbazi-Azad |
J. Syst. Archit. | 1 |
| 2012 | Reconfigurable Cluster-Based Networks-on-Chip for Application-Specific MPSoCsabstractIn this paper, we propose a reconfigurable NoC in which a customized topology for a given application can be implemented. In this NoC, the nodes are grouped into some clusters interconnected by a reconfigurable communication infrastructure. The nodes inside a cluster are connected by a fixed topology. From the traffic management perspective, this structure benefits from the interesting characteristics of the mesh topology (efficient handling of local traffic where each node communicates with its neighbors), while avoids its drawbacks (the lack of short paths between remotely located nodes). We then present a design flow that maps the frequently communicating tasks of a given application into the same cluster and exploits the reconfigurable connections to set up appropriate inter-cluster links. The evaluation results show the power- and performance-efficiency of the proposed NoC. Mehdi Modarressi, Hamid Sarbazi-Azad |
ASAP | 1 |
| 2012 | A Game Theoretical Thermal - Aware Run - Time Task Synchronization Method for Multiprocessor Systems - on - ChipabstractThis paper presents a distributed run-time task synchronization method for multicore processors aiming to reduce the average power consumption of the chip and satisfy a given thermal constraint, while imposing no performance overhead. Being built on the game theory concepts, this is achieved by dynamically changing the frequency of each individual core based on its current workload iteratively until converging to an optimal point. In this work we target two thermal constraints: keeping (1) the core peak temperature and, (2) thermal gradient across the cores below a predefined threshold. The results show that the proposed framework can find the appropriate frequency for each core based on the optimization target, within an acceptable time. Yashar Asgarieh, Mohammad Hassan Khabbazian, Mehdi Modarressi, Hamid Sarbazi-Azad |
DSD | 3 |
| 2012 | The 2D digraph-based NoCs: attractive alternatives to the 2D mesh NoCs
Reza Sabbaghi-Nadooshan, Mehdi Modarressi, Hamid Sarbazi-Azad |
J. Supercomput. | 2 |
| 2011 | Supporting non-contiguous processor allocation in mesh-based CMPs using virtual point-to-point linksabstractIn this paper, we propose a processor allocation mechanism for run-time assignment of a set of communicating tasks of input applications onto the processing nodes of a Chip Multiprocessor (CMP), when the arrival order and execution lifetime of the input applications are not known a priori. This mechanism targets the on-chip communication and aims to reduce the power and latency of the NoC employed as the communication infrastructure. In this work, we benefit from the advantages of non-contiguous processor allocation mechanisms, by allowing the tasks of the input application mapped onto disjoint regions (sub-meshes) and then virtually connecting them by bypassing the router pipeline stages of the inter-region routers. The experimental results show considerable improvement over one of the best existing allocation mechanisms. Marjan Asadinia, Mehdi Modarressi, Arash Tavakkol, Hamid Sarbazi-Azad |
DATE | 2 |
| 2011 | A reconfigurable fault-tolerant routing algorithm to optimize the network-on-chip performance and latency in presence of intermittent and permanent faultsabstractAs the semiconductor industry advances to the deep sub-micron and nano technology points, the on-chip components are more prone to the defects during manufacturing and faults during system operation. Consequently, fault tolerant techniques are essential to improve the yield of modern complex chips. We propose a fault-tolerant routing algorithm that keeps the negative effect of faulty components on the NoC power and performance as low as possible. Targeting intermittent faults, we achieve fault tolerance by employing a simple and fast mechanism composed of two processes: NoC monitoring and route adaption. Experimental results show the effectiveness of the proposed technique, in that it offers lower average message latency and power consumption and a higher reliability, compared to some related work. Reyhaneh Jabbarvand Behrouz, Mehdi Modarressi, Hamid Sarbazi-Azad |
ICCD | 2 |
| 2011 | A Distributed Task Migration Scheme for Mesh-Based Chip-MultiprocessorsabstractA task migration scheme for homogeneous chip multiprocessors (CMP) is presented in this paper. The proposed migration mechanism focuses on the communication sub-system and aims to reduce the total power consumption and latency of the network-on-chip (NoC). In this work, starting from an initial mapping, the tasks migrate to new cores in such a way that the distance between the end-point nodes of high-volume communication flows is reduced. Finding the new place for a task is done in a distributed manner by applying an iterative local search that relies on the local information of each task about its communication demand. The task migration procedure also includes a pre-migration step that aims to produce a high quality (i.e. closer to the optimum point) starting point for the main distributed algorithm. The experimental results under some synthetic and realistic CMP workloads show that this method can effectively adapt the mapping of the tasks to the on-chip communication pattern and improve the power consumption and performance of the on-chip networks. Hossein Yaghoubi, Mehdi Modarressi, Hamid Sarbazi-Azad |
PDCAT | 2 |
| 2011 | Application-Aware Topology Reconfiguration for On-Chip NetworksabstractIn this paper, we present a reconfigurable architecture for networks-on-chip (NoC) on which arbitrary application-specific topologies can be implemented. When a new application starts, the proposed NoC tailors its topology to the application traffic pattern by changing the inter-router connections to some predefined configuration corresponding to the application. It addresses one of the main drawbacks of the existing application-specific NoC optimization methods, i.e., optimization of NoCs based on the traffic pattern of a single application. Supporting multiple applications is a critical feature of an NoC when several different applications are integrated into a single modern and complex multicore system-on-chip or chip multiprocessor. The proposed reconfigurable NoC architecture supports multiple applications by appropriately configuring itself to a topology that matches the traffic pattern of the currently running application. This paper first introduces the proposed reconfigurable topology and then addresses the problems of core to network mapping and topology exploration. Further on, we evaluate the impact of different architectural attributes on the performance of the proposed NoC. Evaluations consider network latency, power consumption, and area complexity. Mehdi Modarressi, Arash Tavakkol, Hamid Sarbazi-Azad |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | An efficient dynamically reconfigurable on-chip network architectureabstractIn this paper, we present a reconfigurable architecture for NoCs on which arbitrary application-specific topologies can be implemented. The proposed NoC can dynamically tailor its topology to the traffic pattern of different applications at run-time. The run-time topology construction mechanism involves monitoring the network traffic and changing the inter-node connections in order to reduce the number of intermediate routers between the source and destination nodes of heavy communication flows. This mechanism should also preserve the NoC connectivity. In this paper, we first introduce the proposed reconfigurable topology and then address the problem of run-time topology reconfiguration. Experimental results show that this architecture effectively improves the NoC power and performance over the existing conventional architectures. Mehdi Modarressi, Hamid Sarbazi-Azad, Arash Tavakkol |
DAC | 1 |
| 2010 | Virtual Point-to-Point Connections for NoCsabstractIn this paper, we aim to improve the performance and power metrics of packet-switched network-on-chips (NoCs) and benefits from the scalability and resource utilization advantages of NoCs and superior communication performance of point-to-point dedicated links. The proposed method sets up the virtual point-to-point (VIP) connections over one virtual channel (which bypasses the entire router pipeline) at each physical channel of the NoC. We present two schemes for constructing such VIP circuits. In the first scheme, the circuits are constructed for an application based on its task-graph at design time. The second scheme addresses constructing the connections at run-time using a light-weight setup network. It involves monitoring the NoC traffic in order to detect heavy communication flows and setting up a VIP connection for them using a run-time circuit construction mechanism. The proposed mechanism is compared to a traditional packet-switched NoC and some modern switching mechanisms and the results show a significant reduction in the network power and latency over the other considered NoCs. Mehdi Modarressi, Arash Tavakkol, Hamid Sarbazi-Azad |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | A hybrid packet-circuit switched on-chip network based on SDMabstractIn this paper, we propose a novel on-chip communication scheme by dividing the resources of a traditional packet-switched network-on-chip between a packet-switched and a circuit-switched sub-network. The former directs packets according to the traditional packet-switching mechanism, while the latter forwards packets over circuits which are directly established between two non-adjacent nodes by bypassing the intermediate routers. A packet may switch between the sub-networks several times to reach its destination. The circuits are set up using a low-latency and low-cost setup-network. The network resources are split between the two sub-networks using Spatial-Division Multiplexing (SDM). The work aims to improve the power and performance metrics of Network-on-Chip (NoC) architectures and benefits from the power and scalability advantage of packet-switched NoCs and superior communication performance of circuit-switching. The evaluation results show a significant reduction in power and latency over a traditional packet-switched NoC. Mehdi Modarressi, Hamid Sarbazi-Azad, Mohammad Arjomand |
DATE | 1 |
| 2009 | Performance and power efficient on-chip communication using adaptive virtual point-to-point connectionsabstractIn this paper, we propose a packet-switched network-on-chip (NoC) architecture which can provide a number of low-power, low-latency virtual point-to-point connections for communication flows. The work aims to improve the power and performance metrics of packet-switched NoC architectures and benefits from the power and resource utilization advantages of NoCs and superior communication performance of point-to-point dedicated links. The virtual point-to-point connections are set up by bypassing the entire router pipeline stages of the intermediate nodes. This work addresses constructing the virtual point-to-point connections at run-time using a light-weight setup network. It involves monitoring the NoC traffic in order to detect heavy communication flows and setting up a virtual point-to-point connection for them using a run-time circuit construction mechanism. The evaluation results show a significant reduction in power and latency over a traditional packet-switched NoC. Mehdi Modarressi, Hamid Sarbazi-Azad, Arash Tavakkol |
NOCS | 1 |
| 2008 | The 2D DBM: An attractive alternative to the simple 2D mesh topology for on-chip networksabstractDuring the recent years, 2D mesh network-onchip has attracted much attention due to its suitability for VLSI implementation. The 2-dimensional de Bruijn topology for network-on-chip is introduced in this paper as an attractive alternative to the popular simple 2D mesh NoC. Its cost is equal to that of the simple 2D mesh but it has a logarithmic diameter. We compare the proposed network and the popular mesh network in terms of power consumption and network performance. Compared to the equal sized simple mesh NoC, the proposed de Bruijn-based network has better performance while consuming less energy. Reza Sabbaghi-Nadooshan, Mehdi Modarressi, Hamid Sarbazi-Azad |
ICCD | 2 |
| 2008 | A novel high-performance and low-power mesh-based NoCabstractIn this paper, a 2D shuffle-exchange based mesh topology, or 2D SEM (shuffle-exchange mesh) for short, is presented for network-on-chips. The proposed two-dimensional topology applies the conventional well-known shuffle-exchange structure in each row and each column of the network. Compared to an equal sized mesh which is the most common topology in on-chip networks, the proposed shuffle-exchange based mesh network has smaller diameter but for an equal cost. Simulation results show that the 2D SEM effectively reduces the power consumption and improves performance metrics of the on-chip networks with regard to the conventional mesh topology. Reza Sabbaghi-Nadooshan, Mehdi Modarressi, Hamid Sarbazi-Azad |
IPDPS | 2 |
| 2007 | Power-aware mapping for reconfigurable NoC architecturesabstractA core mapping method for reconfigurable network-on-chip (NoC) architectures is presented in this paper. In most of the existing methods, mapping is carried out based on the traffic characteristics of a single application. However, several different applications are implemented and integrated in the modern complex system-on-chips which should be considered by mapping methods. In the proposed method, the reconfiguration (which is achieved by embedding programmable switches between routers of a mesh-based NoC) allows us to dynamically change the network topology in order to adapt it with the running application and optimize the power and performance metrics. The presented network architecture can be configured as an application- specific topology, while it still holds the benefits of the regular NoC topologies such as modularity and predictable electrical properties. The experimental results show that this method can effectively adapt the NoC to the running application and improve the power consumption and performance of the system. Mehdi Modarressi, Hamid Sarbazi-Azad |
ICCD | 1 |