Yaoyao Ye

dblp:55/7960 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-0022-228XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 BuffeRS: A Buffer Reservation Scheduling Strategy for Router Bypassing in NoCs and Multichiplet Networks
abstract
Network-on-Chip (NoC) is a crucial communication infrastructure for multi-processor Systems-on-Chip (SoCs). The bypass flow control mechanism is an advanced NoC technique to reduce packet delay, by enabling bufferless packet forwarding through direct path establishment. However, existing works either used inefficient contention-free path allocation methods or complex transmission protocols to ensure destination buffer availability, which significantly limited the efficiency of router bypassing. In this work, we proposeBuffeRS, a buffer reservation scheduling strategy for enhancing the performance of router bypassing by efficiently providing contention-free path allocation and destination buffer availability for bypass packets. It facilitates a dual-mode packet delivery scheme via conventional buffered routing paths and dynamically configured bypass paths, while ensuring buffer availability through coordinated time-division activation of router groups. For chiplet-based systems, we utilizeBuffeRSon chiplet routers to enhance the intra-chiplet bypass transmission efficiency, meanwhile utilizingBuffeRSon boundary routers to bypass inter-chiplet packets and subsequently resolve inter-chiplet deadlocks. Experimental evaluations have demonstrated thatBuffeRSachieves substantial performance gains with a small hardware overhead for both monolithic SoCs and chiplet-based systems.
Xiaoma Wu, Yaoyao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 A Photonic Interconnect Enabling Cross-Layer Scheduling for Chiplet-Based DNN Accelerators
abstract
Deep neural networks (DNNs) demand high-performance and energy-efficient hardware acceleration. Chiplet-based DNN accelerators offer a promising solution for scaling up DNN acceleration by integrating multiple chiplets. However, efficient interconnection of these chiplets remains a significant challenge. Traditional metallic interconnect suffers from limitations in both bandwidth and energy efficiency. Photonic interconnect, offering ultra-high bandwidth and energy efficiency, is emerging as a new interconnection technology for chiplet-based DNN accelerators. However, existing photonic interconnect designs lack the flexibility to implement advanced cross-layer scheduling, a crucial approach for enhancing computational efficiency in contemporary chiplet-based accelerators. To address this issue, we propose PICS, an efficient Photonic Interconnect design enabling flexible Cross-layer Scheduling for chiplet-based DNN accelerators. PICS incorporates a novel photonic interconnect architecture that facilitates efficient data transmission and flexible scheduling alongside a customized fine-grained scheduling and mapping framework. Compared to the state-of-the-art accelerator employing exclusively metallic interconnects with the state-of-the-art cross-layer scheduling framework, PICS demonstrates an average reduction of 9.5% and 31.9% in latency and energy consumption. Compared to the state-of-the-art accelerator employing photonic interconnects with the layer-sequential scheduling framework, PICS achieves an average reduction of 47.2% and 26.0% in latency and energy consumption.
Yaoyao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 A Buffer Reservation Scheduling Strategy for Enhancing Performance of NoC Router Bypassing
abstract
Network-on-Chip (NoC) is one of the key technologies for augmenting performance and energy efficiency of many-core processors, which addresses the limitations of conventional bus architecture. Flow control in NoC manages the allocation of buffers and links, as well as determines resource assignment among packets. Bypass flow control further optimizes this process by permitting certain packets to bypass specific router pipelines to diminish router latency. Nonetheless, bypass flow control necessitates the assurance of conflict-free bypass paths and the availability of buffers at the destinations. In this work, we propose a Buffer Reservation Scheduling strategy (BuffeRS) aimed at enhancing NoC performance by increasing secure bypass packet transmissions within a time period. BuffeRS enables packets to reach their destinations via regular transmission or dynamically generated private bypass paths. By designating specific roles to router groups within corresponding time slots, BuffeRS ensures that destination buffers are reserved before packet arrival. We further delineate router microarchitecture design to implement BuffeRS in NoC. Simulation results demonstrate an average performance enhancement by 17.9%~43% under synthetic traffic and by 6.23% under PARSEC benchmarks in full-system simulations, as compared to contemporary state-of-the-art works.
Yaoyao Ye
ASP-DAC2
2025 An Efficient Branch-and-Bound Routing Optimization Method for Optical NoCs
abstract
Silicon photonics-based optical networks-on-chip (ONoCs) are emerging as a power-efficient on-chip communication architecture for the next generation of chip multiprocessors. However, the thermal sensitivity of photonic devices presents power consumption challenges. Existing routing schemes optimized for optical power loss tend to avoid passing through high-temperature nodes, which in turn leads to contention at low-temperature nodes. It remains a crucial challenge to develop an adaptive routing algorithm that strikes a balance between the power consumption optimization and performance optimization. In this work, we first propose an efficient branch-and-bound routing (BBR) optimization method for ONoCs. To the best of our knowledge, it is the first time that the branch-and-bound (BB) method is adopted to solve the routing optimization problem in ONoCs. We further developed three variants of the BBR optimization method: 1) BBTR; 2) BBCR; and 3) 3BOR. Among them, 3BOR employs a bi-objective bounding function to optimize both optical power loss and network performance, while enhancing algorithmic efficiency. Experimental results demonstrate that, compared to the state-of-the-art heuristic contention-aware thermal-reliable routing algorithm, 3BOR reduces the thermal-induced optical power loss by 15.6% while reducing the algorithm running time by 82.2%.
Yaoyao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 IPDR: An Inter-Chiplet Priority-Driven Deadlock Resolution for 2-D/2.5-D Multichiplet Systems
Yaoyao Ye, Jianfei Jiang 0001, Weiguang Sheng, Ningyi Xu, Yong Lian 0001, Guanghui He 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2024 An Efficient Branch-and-Bound Routing Algorithm for Optical NoCs
abstract
Silicon photonics based optical networks-on-chip (ONoCs) are emerging as a power-efficient on-chip communication architecture for the next generation of chip multiprocessors. However, the thermal sensitivity of photonic devices presents power consumption challenges. Existing routing schemes optimized for optical power loss tend to avoid passing through high-temperature nodes, which in turn leads to contention at low-temperature nodes. It remains a crucial challenge to develop an adaptive routing algorithm that strikes a balance between the power consumption optimization and performance optimization. In this paper, we firstly propose an efficient branch-and-bound routing (BBR) algorithm for ONoCs. Secondly, we obtain 3BOR, a variant of the BBR algorithm with a bi-objective bounding function to optimize both the optical power loss and network performance. To the best of our knowledge, it is the first time that the branch-and-bound method is adopted to solve the routing optimization problem in ONoCs. Experimental results demonstrate that the proposed branch-and-bound routing method outperforms the state-of-the-art heuristic routing algorithm in terms of bi-objective optimization effect as well as algorithm running time. In detail, the 3BOR reduces thermal-induced optical power loss by 14.8% while enhancing the saturation injection rate by 6.5% as compared to the state-of-the-art heuristic contention-aware thermal-reliable routing algorithm.
Yihao Liu 0006, Yaoyao Ye
ASPDAC2
2024 MEIN: A Multicast-Efficient Interconnect Network for Multi-Chiplet DNN Accelerators
abstract
Multi-chiplet DNN accelerator is a promising solution to balancing performance and cost. However, the limited communication bandwidth between chiplets exacerbates the performance bottleneck of the interconnect network. Besides, in DNN dataflows, the same weight or activation is often shared by multiple processing elements (PEs). One-to-many dataflows, also known as multicast, are widespread. Existing works lack specific optimizations for DNN dataflows, thereby yielding suboptimal multicast efficiency. To overcome these challenges, we propose MEIN, a Multicast-Efficient Interconnect Network for multi-chiplet DNN accelerators. Firstly, we introduce a highly efficient routing algorithm tailored for DNN dataflows. It optimizes the multicast tree structure to reduce path latency, while also improving the path selection mechanism to minimize link contention. Secondly, we propose a lightweight router microarchitecture that enhances the hardware resource utilization by simplifying multicast ports. Based on the gem5 simulator, our evaluation demonstrates that MEIN achieves the latency reduction by 19.6%-81.9% as compared to the state-of-the-art related works.
Xuyan Wang, Yaoyao Ye, Guanghui He 0002
ISCAS6
2024 HPPI: A High-Performance Photonic Interconnect Design for Chiplet-Based DNN Accelerators
abstract
In pursuit of higher inference accuracy, the complexity and parameter size of recent deep neural networks (DNNs) have increased significantly. Due to the increasing demand for computing power, the chiplet-based accelerator has been an important computing platform that can handle these DNN models more efficiently. In widely used DNN models, the feature sizes and the number of channels vary greatly among different convolutional layers. Existing chiplet-based accelerators typically adopt consistent optimization strategy for all of the convolutional layers regardless of their sizes, which would limit the inference performance. In this work, we carry out communication-aware customized optimization for convolutional layers with different sizes. First, we propose a reconfigurable high-performance photonic interconnect (HPPI) architecture to facilitate the communication in chiplet-based DNN accelerators. Second, we propose a customized dataflow as the mapping framework and provide four communication patterns of the photonic interconnect with different ways of spatial mapping. Third, we propose a lightweight back propagation neural network to efficiently select the optimal communication pattern for each convolutional layer. The proposed photonic interconnect can be switched between the four communication patterns to enable communication-aware customized optimization for each convolutional layer in the DNN model. As compared to Simba (a representative chiplet-based accelerator with electronic interconnect), HPPI reduces the execution time by 72.23% on average, while saving the energy consumption by 25.49% on average. As compared to ASCEND (a state-of-the-art chiplet-based accelerator with photonic interconnect), HPPI reduces the execution time by 34.04% on average, while saving the energy consumption by 10.34% on average.
Guanglong Li, Yaoyao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 INDM: Chiplet-Based Interconnect Network and Dataflow Mapping for DNN Accelerators
abstract
Chiplet-based deep neural network (DNN) accelerator is a promising solution to balance the performance and manufacturing cost. However, different from monolithic chips, interconnect network design and architectural partitioning for multiple chiplets would result in a huge design space and make it difficult to keep scalability and high hardware utilization. Moreover, how to efficiently map DNN workloads onto multiple DRAM dies and compute dies is another major challenge. To alleviate the above issues, in this work, we propose INDM, a chiplet-based interconnect network and dataflow mapping co-optimization for DNN accelerators. First, we propose an efficient hierarchical interconnect network composed of a multiring on-die network and a cluster-based interdie network, to facilitate the data reuse and traffic pattern in DNN workloads. Second, architectural partitioning and topology exploration for chiplet-based DNN accelerators are proposed to find the optimal architecture configurations. Third, an interdie communication-aware dataflow mapping is proposed to minimize traffic congestion during DNN layer switching. We implement the proposed chiplet-based interconnect network design and dataflow mapping algorithm for a set of popular DNN models, including VGG-16, ResNet-18, DarkNet-19, ResNet-50, and ResNet-101. Experimental results show that as compared with the state-of-the-art related work, such as NN-Baton and SIMBA, our work achieves 26.00%–73.81% energy-delay-product (EDP) reduction and 26.93%–79.78% latency reduction.
Xi Fan, Yaoyao Ye, Xuyan Wang, Guojie Xiong, Xianglun Leng, Ningyi Xu, Yong Lian 0001, Guanghui He 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 M2M: A Fine-Grained Mapping Framework to Accelerate Multiple DNNs on a Multi-Chiplet Architecture
abstract
With the advancement of artificial intelligence, the collaboration of multiple deep neural networks (DNNs) has been crucial to existing embedded systems and cloud systems, especially for automatic driving applications as well as augmented and virtual reality (AR/VR) applications. To trade off between cost and performance, chiplet-based DNN accelerators have emerged as a promising solution for accelerating DNN workloads. However, most existing mapping methods for multiple DNNs target for the monolithic chip, which fail to solve the problems faced by the emerging multi-chiplet architecture, such as the problems of distributed memory access, complex heterogeneous interconnect network, and the scaling-up of computing resources. In this work, we propose M2M, a fine-grained mapping framework for accelerating multiple DNNs on a multi-chiplet architecture. It includes a temporal and spatial task scheduling for reconfigurable dataflow accelerators and a communication-aware task mapping in a heterogeneous interconnect network. To enhance communication efficiency and reduce the overall latency, we further propose a fine-tuned quality-of-service (QoS) policy for network-on-package (NoP) links. To the best of our knowledge, this is the first fine-grained mapping framework for multiple DNNs on a multi-chiplet architecture. We implemented the proposed fine-grained mapping framework using genetic algorithm and simulated annealing algorithm. Experimental results show that our work achieves 7.18%–61.09% latency reduction under vision, language, and mixed workloads when compared with the state-of-the-art related work.
Xuyan Wang, Yaoyao Ye, Dongxu Lyu, Guojie Xiong, Ningyi Xu, Yong Lian 0001, Guanghui He 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2021 Ant Colony Optimization-Based Thermal-Aware Adaptive Routing Mechanism for Optical NoCs
abstract
Optical networks-on-chip (NoC) based on silicon photonics has been proposed as an emerging on-chip communication infrastructure of chip multiprocessors. However, due to thermal sensitivity of optical devices under on-chip temperature variations, significant thermal-induced optical power loss would offset the benefit of optical NoCs in power efficiency. In this work, we propose a thermal-aware adaptive routing scheme based on ant colony optimization (ACO) to alleviate the thermal issue. The proposed ACO-based routing scheme applies the ACO method to formulate and optimize the routing decisions with the objective of reducing optical power loss under temperature variations. The traditional implementation of the ACO-based routing scheme requires a table in each node to keep and update pheromone, and the table size increases linearly with the number of nodes in the network. To avoid the table overhead, we further propose an approximate ACO-based routing (AACO) scheme based on linear regression. A case study on an$8\times 8$mesh-based optical NoC under a series of synthetic traffic patterns and real applications shows that the proposed routing schemes are able to select near-optimal paths under varying on-chip temperature variations. We further verify the scalability of the proposed routing schemes in a larger network.
Jing Wang 0097, Yaoyao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 A Table-Free Approximate Q-Learning-Based Thermal-Aware Adaptive Routing for Optical NoCs
abstract
Optical networks-on-chips (NoCs) based on silicon photonics have been proposed as an emerging communication architecture for many-core chip multiprocessors. However, the thermal sensitivity of silicon photonics is one of the major challenges. Q-learning-based adaptive routing has been proposed in related work to mitigate the thermal issue. However, table overhead of the traditional table-based Q-routing would scale up quickly with the increase of network size. In this article, we propose a table-free approximate Q-learning-based thermal-aware adaptive routing to find optimal low-loss paths in the presence of on-chip temperature variations. The simulation results show that the proposed table-free approximate Q-learning-based adaptive routing can converge faster and it can achieve similar optimization effect as compared to the best optimization effect of the traditional table-based Q-routing. The performance gap between the proposed approximation method and the traditional table-based Q-routing expands when the network size increases.
Wenfei Zhang, Yaoyao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Toward a High-Performance and Low-Loss Clos-Benes-Based Optical Network-on-Chip Architecture
abstract
As chip multiprocessors (CMPs) keep growing in capability, on-chip communication efficiency is crucial to the overall performance. However, on-chip networks based on electronic switches suffer from excessive power consumption and limited performance. In order to take advantages of optical interconnect, we propose an optimized design toward a high-performance and low-loss Clos-Benes-based hierarchical optical network-on-chip (NoC) for large-scale CMPs. We propose several key techniques, including a loss-aware adaptive (LAA) routing for intraswitch Benes network, a priority-based round-Robin virtual output queue selection and a Q-learning-based heuristic routing for interswitch Clos network, and a local transfer link (LTL) technique to improve the traffic locality. A case study on a 256-core CMP under uniform traffic shows that the network throughput is increased by 346.7%, 61%, and 12.9%, respectively, than the mesh, fat-tree, and the traditional generic Clos-Benes optical NoC. On average of a set of real applications, the application end-to-end delay is reduced by 47.6%, 28.2%, and 19.4%, respectively, than the mesh, fattree, and the traditional generic Clos-Benes network. Meanwhile, the average optical power loss is decreased by 11.8% and 49.9%, respectively, as compared to the mesh and fattree. As compared to a baseline Clos-Benes network, the use of LAA routing together with the LTL could reduce the average optical power loss by 28.4%.
Renjie Yao, Yaoyao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Thermal-Aware Design and Simulation Approach for Optical NoCs
abstract
For chip multiprocessors, one major challenge is to bridge the increasing speed gap between processor and the global on-chip interconnect delay. By integrating optical interconnects in network-on-chip (NoC) architectures, optical NoCs can overcome the power and bandwidth bottleneck of traditional electrical on-chip networks. However, while considering the thermal sensitivity of silicon photonic devices used in optical NoCs, optical interconnects may not have advantages in power efficiency as compared with their electrical counterparts. To tackle this problem, in this article, we propose a thermal-aware design and simulation approach for optical NoCs. Key techniques include thermal-sensitive optical power loss models from device level to network level, a thermal-aware adaptive routing mechanism, and a thermal-aware simulation platform. The thermal-aware simulation platform enables optical NoC simulation together with on-chip temperature simulation as well as optical thermal effect modeling. With the proposed thermal-aware simulation platform, we conducted a case study of an 8 x 8 mesh-based optical NoC under a set of synthetic traffic patterns as well as real applications at typical temperature scenarios. By comparing and analyzing different temperature distributions, we can conclude that it can achieves a better optimization effect for the temperature distributions where the hot spots are scattered across the chip.
Yaoyao Ye, Wenfei Zhang, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Hardware-Software Collaborative Thermal Sensing in Optical Network-on-Chip-based Manycore Systems
abstract
Continuous technology scaling in manycore systems leads to severe overheating issues. To guarantee system reliability, it is critical to accurately yet efficiently monitor runtime temperature distribution for effective chip thermal management. As an emerging communication architecture for new-generation manycore systems, optical network-on-chip (ONoC) satisfies the communication bandwidth and latency requirements with low power dissipation. Moreover, observation shows that it can be leveraged for runtime thermal sensing. In this article, we propose a brand-new on-chip thermal sensing approach for ONoC-based manycore systems by utilizing the intrinsic thermal sensitivity of optical devices and the inter-processor communications in ONoCs. It requires no extra hardware but utilizes existing optical devices in ONoCs and combines them with lightweight software computation in a hardware-software collaborative manner. The effectiveness of the our approach is validated both at the device level and the system level through professional photonic simulations. Evaluation results based on synthetic communication traces and realistic benchmarks show that our approach achieves an average temperature inaccuracy of only 0.6648 K compared to ground-truth values and is scalable to be applied for large-size ONoCs.
Mengquan Li, Weichen Liu 0001, Nan Guan, Yiyuan Xie, Yaoyao Ye
ACM Trans. Embed. Comput. Syst.5
2019 A thermal-sensitive design of a 3D torus-based optical NoC architecture
Yaoyao Ye
Integr.1
2018 User Experience-Enhanced and Energy-Efficient Task Scheduling on Heterogeneous Multi-Core Mobile Systems
abstract
Heterogeneous Multi-Core Mobile Systems has been widely used to improve performance. However, it faces with the challenge of tradeoff between energy saving and user experience. ARM big. LITTLE architecture, a heterogeneous computing architecture, is a power-optimization technology. In most big. LITTLE devices, however, it still cannot achieve excellent user experience and higher energy saving. In this paper, we propose an improved task scheduling (UCES-GTS) by introducing the concept of user-centric task on big. LITTLE mobile device. In order to enhance user experience, the response time of user-centric tasks is shortened with reducing slack time of them properly. We then present a detailed algorithm to compute appropriate frequency and allocate the CPU resources to each task. The experimental evaluation results show that our improved global task scheduling model can achieve 17 % and 8 % energy saving average compared with the clustered switching scheduling and the original global task scheduling respectively. And the response time of user-centric tasks can decrease 27 % average, which means excellent user experience.
Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye
ICPADS7
2018 Fine-Grained Task-Level Parallel and Low Power H.264 Decoding in Multi-Core Systems
abstract
In the past few years, the extinction of Moore's Law makes people reconsider the solutions for dealing with the low computing resource utilization of applications on multicore processor systems. However, making good use of computing resources in multi-core processors systems is not easy due to the differences between single-core and multi-core architecture. Nowadays short video apps like Instagram and Tik Tok have successfully caught people's eyes by fascinating short videos, typically just 10 to 30 seconds long, uploaded by the users of apps. And almost all of these videos are recorded by their mobile devices, which are typically HD (High Definition) or FHD (Full High Definition) videos, which prefer to be encoded/decoded by H.264/AVC rather then HEVC (High Efficiency Video Coding) on mobile devices in view of the energy consumption and decoding speed. How to dive the huge potential of the computing resource on multi-core mobile devices to speed up decoding these videos while consuming low energy, is a big challenge. In our previous work [1], a relatively simple parallel framework was proposed to implement a parallel H.264/ AV C decoder. This work further proposes a more detailed systematic task-level parallel framework, together with an energy saving strategy based on this framework, to research a new H.264/AVC decoder on multi-core processor systems. The proposed parallel method is composed of a set of rules to guide parallel software programming (PSPR) and a software parallelization framework (SPF). The PSPR is applied in pre-processing steps to address the potential issues limiting the inherent parallelism, and the SPF is applied to parallelize the original serial programs. After the parallelization is successfully deployed, DVFS technique would be applied to decrease the power dissipation based on the SPF. Results show that proposed solutions make a significant improvement in decoding speed of 32% at 720p, 27% at 1080p and 29% at 2160p, and in energy savings of 25% at 720p, 25% at 1080p and 23% at 2160p on a four-core workstation running Linux, compared to the original serial H.264/ AV C decoder. The results demonstrate our methods are effective and scalable, served as a reference for future parallel software development.
Wenyang Liu, Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye
ICPADS7
2018 A Learning-Based Thermal-Sensitive Power Optimization Approach for Optical NoCs
abstract
Optical networks-on-chip (NoCs) based on silicon photonics have been proposed as emerging on-chip communication architectures for chip multiprocessors with large core counts. However, due to the thermal sensitivity of optical devices used in optical NoCs, on-chip temperature variations cause significant thermal-induced optical power loss, which would counteract the power advantages of optical NoCs. To tackle this problem, in this work, we propose a learning-based thermal-sensitive power optimization approach for mesh- or torus-based optical NoCs in the presence of temperature variations. The key techniques proposed include an initial device-setting and thermal-tuning mechanism that is a device-level optimization technique, and a learning-based thermal-sensitive adaptive routing algorithm that is a network-level optimization technique. Simulation results of an 8x8 mesh-based optical NoC show that the proposed initial device-setting and thermal-tuning mechanism confines the worst-case thermal-induced optical energy consumption to be on the order of tens of pJ/bit, by avoiding significant thermal-induced optical power loss caused by temperature-dependent wavelength shifts. Besides, it shows that the learning-based thermal-sensitive adaptive routing algorithm is able to find an optimal path with the minimum estimated thermal-induced optical power consumption for each communication pair. The proposed routing has a greater space for optimization, especially for applications with more long-distance traffic.
Yaoyao Ye
ACM J. Emerg. Technol. Comput. Syst.2
2018 Thermal-Sensor-Based Occupancy Detection for Smart Buildings Using Machine-Learning Methods
abstract
In this article, we propose a novel approach to detect the occupancy behavior of a building through the temperature and/or possible heat source information. The new method can be used for energy reduction and security monitoring for emerging smart buildings. Our work is based on a building simulation program, EnergyPlus, from the Department of Energy. EnergyPlus can model various time-series inputs to a building such as ambient temperature; heating, ventilation, and air-conditioning (HVAC) inputs; power consumption of electronic equipment; lighting; and number of occupants in a room, sampled each hour, and produce resulting temperature traces of zones (rooms). Two machine-learning-based approaches for detecting human occupancy of a smart building are applied herein, namely support vector regression (SVR) and recurrent neural network (RNN). Experimental results with SVR show that the four-feature model provides accurate detection rates, giving a 0.638 average error and 5.32% error rate, and the five-feature model delivers a 0.317 average error and 2.64% error rate. This indicates that SVR is a viable option for occupancy detection. In the RNN method, Elman’s RNN can estimate occupancy information of each room of a building with high accuracy. It has local feedback in each layer and, for a five-zone building, it is very accurate for occupancy behavior estimation. The error level, in terms of number of people, can be as low as 0.0056 on average and 0.288 at maximum, considering ambient, room temperatures, and HVAC powers as detectable information. Without knowing HVAC powers, the estimation error can still be 0.044 on average, and only 0.71% estimated points have errors greater than 0.5. Our article further shows that both methods deliver similar accuracy in the occupancy detection. But the SVR model is more stable for adding or removing features of the system, while the RNN method can deliver more accuracy when the features used in the model do not change a lot.
Hengyang Zhao, Qi Hua, Haibao Chen, Yaoyao Ye, Hai Wang 0002, Sheldon X.-D. Tan, Esteban Tlelo-Cuautle
ACM Trans. Design Autom. Electr. Syst.4
2017 Thermal-sensitive design and power optimization for a 3D torus-based optical NoC
abstract
In order to overcome limitations of traditional electronic interconnects in terms of power efficiency and bandwidth density, optical networks-on-chip (NoCs) based on 3D integrated silicon photonics have been proposed as an emerging on-chip communication architecture for multiprocessor systems-on-chip (MPSoCs) with large core counts. However, due to thermo-optic effects, wavelength-selective silicon photonic devices such as microresonators, which are widely used in optical NoCs, suffer from temperature-dependent wavelength shifts. As a result, on-chip temperature variations cause significant thermal-induced optical power loss which may counteract the power advantages of optical NoCs. To tackle this problem, in this work, we present a thermal-sensitive design and power optimization approach for a 3D torus-based optical NoC architecture. Based on an optical thermal modeling platform which models the thermal effect in optical NoCs from a system-level perspective, a thermal-sensitive routing algorithm is proposed for the 3D torus-based optical NoC to optimize its power consumption in the presence of on-chip temperature variations. Simulation results show that in an 8×8×2 3D torus-based optical NoC under a set of real applications, as compared with a matched 3D mesh-based optical NoC with traditional dimension order routing, the power consumption is reduced by 25% if thermal tuning for microresonators is not utilized, by 19% if thermal tuning is utilized for microresonators, and by 17% if athermal microresonators are used.
Kang Yao, Yaoyao Ye, Sudeep Pasricha, Jiang Xu 0001
ICCAD2
2015 Alleviate chip I/O pin constraints for multicore processors through optical interconnects
abstract
Chip I/O pins are an increasingly limited resource and significantly affect the performance, power and cost of multicore processors. Optical interconnects promise low power and high bandwidth, and are potential alternatives to electrical interconnects. This work systematically developed a set of analytical models for electrical and optical interconnects to study their structures, receiver sensitivities, crosstalk noises, and attenuations. We verified the models by published implementation results. The analytical models quantitatively identified the advantages of optical interconnects in terms of bandwidth, energy consumption, and transmission distance. We showed that optical interconnects can significantly reduce chip pin counts. For example, compared to electrical interconnects, optical interconnects can save at least 92% signal pins when connecting chips more than 25 cm (10 inches) apart.
Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Xuan Wang 0001, Zhe Wang 0003, Luan H. K. Duong, Haoran Li 0002, Rafael Kioji Vivas Maeda, Xiaowen Wu, Yaoyao Ye, Qinfen Hao
ASP-DAC11
2015 Efficient SAT-based application mapping and scheduling on multiprocessor systems for throughput maximization
abstract
Multiprocessor systems are becoming ubiquitous in today's embedded systems design. In this paper, we address the problem of mapping an application represented by a Homogeneous Synchronous Dataflow (HSDF) graph onto a real-time multiprocessor platform with the objective of maximizing total throughput. We propose that the optimal solution to the problem is composed of three components: actor-to-processor mapping, retiming, and actor ordering on each processor. The entire problem is systematically modeled into a SAT problem and solved by a modern SAT solver formally such that the optimal solution can be guaranteed. In order to explore the vast solution space more efficiently, we develop a specific HSDF theory solver based on the special characteristics of the timed HSDF, and integrate it into the general search framework of the SAT solver. The enhanced optimization framework implemented in branch and bound is able to conduct early branch pruning in the search space, and the scalability is thus greatly improved. Extensive performance evaluation on synthetic examples and a case study on the realistic H.264 Video Decoder shows that our technique provides as much as 76.9% throughput improvement, and it is scalable to industry-sized applications.
Weichen Liu 0001, Zonghua Gu 0001, Yaoyao Ye
CASES3
2015 Crosstalk Noise in WDM-Based Optical Networks-on-Chip: A Formal Study and Comparison
abstract
Optical networks-on-chip (ONoCs) using wavelength-division multiplexing (WDM) technology have progressively attracted more and more attention for their use in tackling the high-power consumption and low bandwidth issues in growing metallic interconnection networks in multiprocessor systems-on-chip. However, the basic optical devices employed to construct WDM-based ONoCs are imperfect and suffer from inevitable power loss and crosstalk noise. Furthermore, when employing WDM, optical signals of various wavelengths can interfere with each other through different optical switching elements within the network, creating crosstalk noise. As a result, the crosstalk noise in large-scale WDM-based ONoCs accumulates and causes severe performance degradation, restricts the network scalability, and considerably attenuates the signal-to-noise ratio (SNR). In this paper, we systematically study and compare the worst case as well as the average crosstalk noise and SNR in three well-known optical interconnect architectures, mesh-based, folded-torus-based, and fat-tree-based ONoCs using WDM. The analytical models for the worst case and the average crosstalk noise and SNR in the different architectures are presented. Furthermore, the proposed analytical models are integrated into a newly developed crosstalk noise and loss analysis platform (CLAP) to analyze the crosstalk noise and SNR in WDM-based ONoCs of any network size using an arbitrary optical router. Utilizing CLAP, we compare the worst case as well as the average crosstalk noise and SNR in different WDM-based ONoC architectures. Furthermore, we indicate how the SNR changes in respect to variations in the number of optical wavelengths in use, the free-spectral range, and the microresonators$\boldsymbol {Q}$factor. The analyses’ results demonstrate that the crosstalk noise is of critical concern to WDM-based ONoCs: in the worst case, the crosstalk noise power exceeds the signal power in all three WDM-based ONoC architectures, even when the number of processor cores is small, e.g., 64.
Mahdi Nikdast, Jiang Xu 0001, Luan H. K. Duong, Xiaowen Wu, Xuan Wang 0001, Zhehui Wang, Zhe Wang 0003, Peng Yang 0003, Yaoyao Ye, Qinfen Hao
IEEE Trans. Very Large Scale Integr. Syst.9
2015 Actively Alleviate Power Gating-Induced Power/Ground Noise Using Parasitic Capacitance of On-Chip Memories in MPSoC
abstract
By integrating multiple processing units (PUs) and memories on a single chip, multiprocessor system-on-chip (MPSoC) can provide higher performance per energy and lower cost per function to applications with growing complexity. On the other hand, shrinking feature sizes and reducing power supply voltages also make MPSoCs more susceptible to various reliability threats, such as power/ground (P/G) noises. Power gating is an effective technique to minimize leakage power. However, it also introduces significant P/G noises in MPSoCs. With significant area, power and performance overheads, traditional methods rely on reinforced circuits or fixed protection strategies to reduce P/G noises caused by power gating. In this paper, we propose a systematic approach to actively alleviating P/G noises using the parasitic capacitance of on-chip memories through sensor network on-chip (SENoC). We use the parasitic capacitance of on-chip memories as dynamic decoupling capacitance to suppress P/G noises and develop a detailed HSPICE model for related study. SENoC is developed to not only monitor and report P/G noises, but also coordinate PUs and memories to alleviate such transient threats at run time. Extensive evaluations show that compared with traditional method, our approach saves 12.6%–62.8% energy consumption and achieves 14.3%–69.8% performance improvement for different applications and MPSoCs with different scales. We implement the circuit details of our approach and show its low area and energy consumption overheads.
Xuan Wang 0001, Jiang Xu 0001, Wei Zhang 0012, Xiaowen Wu, Yaoyao Ye, Zhehui Wang, Mahdi Nikdast, Zhe Wang 0003
IEEE Trans. Very Large Scale Integr. Syst.5
2015 An Inter/Intra-Chip Optical Network for Manycore Processors
abstract
Manycore processor system is becoming an attractive platform for applications seeking both high performance and high energy efficiency. However, huge communication demands among cores, large power density, and low process yield will be three significant limitations for the scalability of future manycore processors. Breaking a large chip into multiple smaller ones can alleviate the problems of power density and yield, but would worsen the problem of communication efficiency due to the limited off-chip bandwidth. In response, we propose an inter/intra-chip optical network, which will not only fulfill the intra-chip communication requirements but also address the inter-chip communication, by exploiting the advantages of optical links with high bandwidth and energy efficiency. The network is composed of an inter-chip subnetwork and multiple intra-chip subnetworks, and the subnetworks closely coordinate with each other to balance the traffic. The proposed network effectively explores the distinctive properties of optical signals and photonic devices, and dynamically partitions each data channel into multiple sections. Each section can be utilized independently to boost performance as well as reduce energy consumption. Simulation results show that our network can achieve higher throughput with lower power consumption than alternative designs under most of synthetic traffics and real applications.
Xiaowen Wu, Jiang Xu 0001, Yaoyao Ye, Xuan Wang 0001, Mahdi Nikdast, Zhehui Wang, Zhe Wang 0003
IEEE Trans. Very Large Scale Integr. Syst.3
2014 CLAP: a crosstalk and loss analysis platform for optical interconnects
abstract
Basic photonic devices in inter- and intra-chip optical networks suffer from inevitable power loss and crosstalk noise. Incoherent crosstalk introduces quick power fluctuations, while coherent crosstalk varies the optical power of the optical signal in optical interconnection networks (OINs). As a result, the accumulative crosstalk in large scale OINs considerably hurts the signal-to-noise ratio (SNR) and imposes high power penalties. In this work, we aim at studying the worst-case incoherent and coherent crosstalk in OINs at the system level. The proposed analytical models are integrated into a newly developed crosstalk and loss analysis platform, called CLAP, to facilitate the SNR analyses in arbitrary OINs.
Mahdi Nikdast, Luan H. K. Duong, Jiang Xu 0001, Sébastien Le Beux, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Yaoyao Ye
NOCS8
2014 On-chip sensor networks for soft-error tolerant real-time multiprocessor systems-on-chip
abstract
As transistor density continues to increase with the advent of nanotechnology, reliability issues raised by the more frequent appearance of soft errors are becoming critical for future embedded multiprocessor systems design. State-of-the-art techniques for soft error protections targeting multiprocessor systems result either high chip cost and area overhead or high performance degradation and energy consumption, and do not fulfill the increasing requirements for high performance and dependability. In this article we present a systematic approach, that is, the Sensor Networks-on-Chip (SENoC), to collaboratively and efficiently manage on-chip applications and overcome reliability threats to Multiprocessor Systems-on-Chip (MPSoC). A hardware-software collaborative approach is proposed to solve soft error problems: a hardware-based on-chip sensor network is built for soft error detection, and a software-based recovery mechanism is applied for soft error correction. A two-step scheduling scheme is presented for reliable application and chip management, combining an off-line static optimization stage for application performance maximization and an online lightweight dynamic adjustment stage to handle runtime variations and exceptions. This strategy introduces only trivial overhead on hardware design and much lower overhead on software control and execution, and hence performance degradation and energy consumption is greatly reduced. We build a cycle-accurate simulator using SystemC, and verify the effectiveness of our technique by comparing performance with related techniques on several real-world applications.
Weichen Liu 0001, Xuan Wang 0001, Jiang Xu 0001, Wei Zhang 0012, Yaoyao Ye, Xiaowen Wu, Mahdi Nikdast, Zhehui Wang
ACM J. Emerg. Technol. Comput. Syst.5
2014 SUOR: Sectioned Undirectional Optical Ring for Chip Multiprocessor
abstract
Chip multiprocessor (CMP) is becoming an attractive platform for applications seeking both high performance and high energy efficiency. In large-scale CMPs, the communication efficiency among cores is crucial for the overall system performance and energy consumption. In this article, we propose a ring-based optical network-on-chip, called SUOR, to fulfill the communication requirement of CMPs. SUOR effectively explores the distinctive properties of optical signals and photonic devices, and dynamically partitions each data channel into multiple sections. Each section can be utilized independently to boost performance as well as reduce energy consumption. We develop a set of distributed control protocols and algorithms for SUOR, but physically allocate the corresponding cluster agents close to each other to benefit from the strengths of optical interconnects at long distances as well as electrical interconnects at short distances. Simulation results show that SUOR outperforms the alternative optical networks under a wide range of traffic patterns. For example, compared with MWSR design, SUOR achieves 2.58× throughput as well as saves 64% energy consumption on average in a 256-core CMP. Compared with MWMR design, SUOR achieves 1.52× throughput and reduces 73% energy consumption on average.
Xiaowen Wu, Jiang Xu 0001, Yaoyao Ye, Zhehui Wang, Mahdi Nikdast, Xuan Wang 0001
ACM J. Emerg. Technol. Comput. Syst.3
2014 Floorplan Optimization of Fat-Tree-Based Networks-on-Chip for Chip Multiprocessors
abstract
Chip multiprocessor (CMP) is becoming increasingly popular in the processor industry. Efficient network-on-chip (NoC) that has similar performance to the processor cores is important in CMP design. Fat-tree-based on-chip network has many advantages over traditional mesh or torus-based networks in terms of throughput, power efficiency, and latency. It has a bright future in the development of CMP. However, the floorplan design of the fat-tree-based NoC is very challenging because of the complexity of topology. There are a large number of crossings and long interconnects, which cause severe performance degradation in the network. In electronic NoCs, the parasitic capacitance and inductance will be significant. In optical ones, large crosstalk noise and power loss will be introduced. The novel contribution of this paper is to propose a method to optimize the fat-tree floorplan, which can effectively reduce the number of crossings and minimize the interconnect length. Two types of floorplans are proposed, which could be applied to fat-tree-based networks of arbitrary size. Compared with the traditional one, our floorplans could reduce more than 87% of the crossings. Since the traversal distance for signals is related to the aspect ratio of the processor cores, we also present a method to calculate the optimum aspect ratio of the processor cores to minimize the traversal distance.
Zhehui Wang, Jiang Xu 0001, Xiaowen Wu, Yaoyao Ye, Wei Zhang 0012, Mahdi Nikdast, Xuan Wang 0001, Zhe Wang 0003
IEEE Trans. Computers4
2014 Systematic Analysis of Crosstalk Noise in Folded-Torus-Based Optical Networks-on-Chip
abstract
Photonic devices are widely used in optical networks-on-chip (ONoCs) and suffer from crosstalk noise. The accumulative crosstalk noise in large scale ONoCs diminishes the signal-to-noise ratio (SNR), causes severe performance degradation, and constrains the network scalability. For the first time, this paper systematically analyzes and models the worst-case crosstalk noise and SNR in folded-torus-based ONoCs. Formal analytical models for the worst-case crosstalk noise and SNR are presented. The crosstalk noise analysis is hierarchically performed at the basic photonic device level, then at the optical router level, and finally at the network level. We consider a general 5$\,\times\,$5 optical router model to enable crosstalk noise and SNR analyses in folded-torus-based ONoCs using an arbitrary 5$\,\times\,$5 optical router. Using the general optical router model, the worst-case SNR link candidates, which restrict the network scalability, are found. Also, we present a novel crosstalk noise and loss analysis platform, called CLAP, which can analyze the crosstalk noise and SNR of arbitrary ONoCs. Case studies of optimized crossbar and Crux optical routers using recent photonic device parameters are presented. Moreover, we compare the worst-case crosstalk noise and SNR in folded-torus-based and mesh-based ONoCs using optimized crossbar and Crux optical routers. The quantitative simulation results show the critical behavior of crosstalk noise in large scale ONoCs. For example, in folded-torus-based ONoCs using the Crux optical router, the noise power exceeds the signal power for network sizes larger than 12$\,\times\,$12; when the network size is 20$\,\times\,$20 and the injection signal power equals 0 dBm, the signal power and noise power are${-}{\rm 9.4}~{\rm dBm}$and${-}{\rm 6.1}~{\rm dBm}$, respectively.
Mahdi Nikdast, Jiang Xu 0001, Xiaowen Wu, Wei Zhang 0012, Yaoyao Ye, Xuan Wang 0001, Zhehui Wang, Zhe Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2014 System-Level Modeling and Analysis of Thermal Effects in WDM-Based Optical Networks-on-Chip
abstract
Multiprocessor systems-on-chip show a trend toward integration of tens and hundreds of processor cores on a single chip. With the development of silicon photonics for short-haul optical communication, wavelength division multiplexing (WDM)-based optical networks-on-chip (ONoCs) are emerging on-chip communication architectures that can potentially offer high bandwidth and power efficiency. Thermal sensitivity of photonic devices is one of the main concerns about the on-chip optical interconnects. We systematically modeled thermal effects in optical links in WDM-based ONoCs. Based on the proposed thermal models, we developed OTemp, an optical thermal effect modeling platform for optical links in both WDM-based ONoCs and single-wavelength ONoCs. OTemp can be used to simulate the power consumption as well as optical power loss for optical links under temperature variations. We use case studies to quantitatively analyze the worst-case power consumption for one wavelength in an eight-wavelength WDM-based optical link under different configurations of low-temperature-dependence techniques. Results show that the worst-case power consumption increases dramatically with on-chip temperature variations. Thermal-based adjustment and optimal device settings can help reduce power consumption under temperature variations. Assume that off-chip vertical-cavity surface-emitting lasers are used as the laser source with WDM channel spacing of 1 nm, if we use thermal-based adjustment with guard rings for channel remapping, the worst-case total power consumption is 6.7 pJ/bit under the maximum temperature variation of 60 °C; larger channel spacing would result in a larger worst-case power consumption in this case. If we use thermal-based adjustment without channel remapping, the worst-case total power consumption is around 9.8 pJ/bit under the maximum temperature variation of 60 °C; in this case, the worst-case power consumption would benefit from a larger channel spacing.
Yaoyao Ye, Zhehui Wang, Peng Yang 0003, Jiang Xu 0001, Xiaowen Wu, Xuan Wang 0001, Mahdi Nikdast, Zhe Wang 0003, Luan H. K. Duong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 UNION: A Unified Inter/Intrachip Optical Network for Chip Multiprocessors
abstract
As modern computing systems become increasingly complex, communication efficiency among and inside chips has become as important as the computation speeds of individual processing cores. Traditionally, to maximize design flexibility, interchip and intrachip communication architectures are separately designed under different constraints. Jointly designing communication architectures for both interchip and intrachip communication could, however, potentially yield better solutions. In this paper, we present a unified inter/intrachip optical network, called UNION, for chip multiprocessors (CMPs). UNION is based on recent progresses in nanophotonic technologies. It connects not only cores on a single CMP, but also multiple CMPs in a system. UNION employs a hierarchical optical network to separate interchip communication traffic from intrachip communication traffic. It fully utilizes a single optical network to transmit both payload and control packets. The network controller on each CMP not only manages intrachip communications, but also collaborates with each other to facilitate interchip communications. We compared UNION with a matched electrical counterpart in 45-nm process. Simulation results for eight real CMP applications show that on average UNION improves CMP performance by 3× while reducing 88% of network energy consumption.
Xiaowen Wu, Yaoyao Ye, Jiang Xu 0001, Wei Zhang 0012, Weichen Liu 0001, Mahdi Nikdast, Xuan Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Active power-gating-induced power/ground noise alleviation using parasitic capacitance of on-chip memories
abstract
By integrating multiple processing units and memories on a single chip, multiprocessor system-on-chip (MPSoC) can provide higher performance per energy and lower cost per function to applications with growing complexity. In order to maintain the power budget, power gating technique is widely used to reduce the leakage power. However, it will introduce significant power/ground (P/G) noises, and threat the reliability of MPSoCs. With significant area, power and performance overheads, traditional methods rely on reinforced circuits or fixed protection strategies to reduce P/G noises caused by power gating. In this paper, we propose a systematic approach to actively alleviating P/G noises using the parasitic capacitance of on-chip memories through sensor network on-chip (SENoC). We utilize the parasitic capacitance of on-chip memories as dynamic decoupling capacitance to suppress P/G noises and develop a detailed Hspice model for related study. SENoC is developed to not only monitor and report P/G noises but also coordinate processing units and memories to alleviate such transient threats at run time. Extensive evaluations show that compared with traditional methods, our approach saves 11.7% to 62.2% energy consumption and achieves 13.3% to 69.3% performance improvement for different applications and MPSoCs with different scales. We implement the circuit details of our approach and show its low area and energy consumption overheads.
Xuan Wang 0001, Jiang Xu 0001, Wei Zhang 0012, Xiaowen Wu, Yaoyao Ye, Zhehui Wang, Mahdi Nikdast, Zhe Wang 0003
DATE5
2013 System-level analysis of mesh-based hybrid optical-electronic network-on-chip
abstract
Network-on-chip (NoC) can improve the performance, power efficiency, and scalability of multiprocessor system-on-chip (MPSoC). Optical NoCs, which are based on CMOS-compatible optical waveguides and microresonators, have significant bandwidth and power advantages over metallic interconnects. We propose a low-cost mesh-based hybrid optical-electronic NoC, HOME, with non-blocking 5×5, 4×4 and 3×3 optical switching fabrics. We systematically analyzed the key characteristics of HOME for a 64-core MPSoC in 45nm under different traffic conditions. Besides, we quantitatively analyzed the thermal effects in the 64-core HOME under temperature variations.
Yaoyao Ye, Xiaowen Wu, Jiang Xu 0001, Mahdi Nikdast, Zhehui Wang, Xuan Wang 0001, Zhe Wang 0003
ISCAS1
2013 3-D Mesh-Based Optical Network-on-Chip for Multiprocessor System-on-Chip
abstract
Optical networks-on-chip (ONoCs) are emerging communication architectures that can potentially offer ultrahigh communication bandwidth and low latency to multiprocessor systems-on-chip (MPSoCs). In addition to ONoC architectures, 3-D integrated technologies offer an opportunity to continue performance improvements with higher integration densities. In this paper, we present a 3-D mesh-based ONoC for MPSoCs, and new low-cost nonblocking 4$\,\times\,$4, 5$\,\times\,$5, 6$\,\times\,$6, and 7$\,\times\,$7 optical routers for dimension-order routing in the 3-D mesh-based ONoC. Besides, we propose an optimized floorplan for the 3-D mesh-based ONoC. The floorplan follows the regular 3-D mesh topology but implements all optical routers in a single optical layer. The floorplan is optimized to minimize the number of extra waveguide crossings caused when merging the 3-D ONoC to one optical layer. Based on a set of real applications and uniform traffic pattern, we develop a SystemC-based cycle-accurate NoC simulator and compare the 3-D mesh-based ONoC with the matched 2-D mesh-based ONoC and 2-D electronic NoC for performance and energy efficiency. Additionally, we quantitatively analyze thermal effects on the 3-D 8$\,\times\,$8$\,\times\,$2 mesh-based ONoC.
Yaoyao Ye, Jiang Xu 0001, Baihan Huang, Xiaowen Wu, Wei Zhang 0012, Xuan Wang 0001, Mahdi Nikdast, Zhehui Wang, Weichen Liu 0001, Zhe Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2013 Formal Worst-Case Analysis of Crosstalk Noise in Mesh-Based Optical Networks-on-Chip
abstract
Crosstalk noise is an intrinsic characteristic as well as a potential issue of photonic devices. In large scale optical networks-on-chips (ONoCs), crosstalk noise could cause severe performance degradation and prevent ONoC from communicating properly. The novel contribution of this paper is the systematical modeling and analysis of the crosstalk noise and the signal-to-noise ratio (SNR) of optical routers and mesh-based ONoCs using a formal method. Formal analytical models for the worst-case crosstalk noise and minimum SNR in mesh-based ONoCs are presented. The crosstalk analysis is performed at device, router, and network levels. A general 5$\,\times\,$5 optical router model is proposed for router level analysis. The minimum SNR optical link candidates, which constrain the scalability of mesh-based ONoCs, are identified. It is also shown that symmetric mesh-based ONoCs have the best SNR performance. The presented formal analyses can be easily applied to other optical routers and mesh-based ONoCs. Finally, we present case studies of mesh-based ONoCs using the optimized crossbar and Crux optical routers to evaluate the proposed formal method. We find that crosstalk noise can significantly limit the scalability of mesh-based ONoCs. For example, when the mesh-based ONoC size, using optimized crossbar, is larger than 8$\,\times\,$8, the optical signal power is smaller than the crosstalk noise power; when the network size is 16$\,\times\,$16 and the input power is 0 dBm, in the worst-case, the signal power is${-}{\rm 24.9}~{\rm dBm}$and the crosstalk noise power is${-}{\rm 11}~{\rm dBm}$.
Yiyuan Xie, Mahdi Nikdast, Jiang Xu 0001, Xiaowen Wu, Wei Zhang 0012, Yaoyao Ye, Xuan Wang 0001, Zhehui Wang, Weichen Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2013 System-Level Modeling and Analysis of Thermal Effects in Optical Networks-on-Chip
abstract
The performance of multiprocessor systems, such as chip multiprocessors (CMPs), is determined not only by individual processor performance, but also by how efficiently the processors collaborate with one another. It is the communication architecture that determines the collaboration efficiency on the hardware side. Optical networks-on-chip (ONoCs) are emerging communication architectures that can potentially offer ultra-high communication bandwidth and low latency to multiprocessor systems. Thermal sensitivity is an intrinsic characteristic of photonic devices used by ONoCs as well as a potential issue. This paper systematically modeled and quantitatively analyzed the thermal effects in ONoCs. We used an 8$\times$8 mesh-based ONoC as a case study and evaluated the impacts of thermal effects in the average power efficiency for real MPSoC applications. We revealed three important factors regarding ONoC power efficiency under temperature variations, and proposed several techniques to reduce the temperature sensitivity of ONoCs. These techniques include the optimal initial setting of microresonator resonant wavelength, increasing the 3-dB bandwidth of optical switching elements by parallel coupling multiple microresonators, and the use of passive-routing optical router Crux to minimize the number of switching stages in mesh-based ONoCs. We gave a mathematical analysis of periodically parallel coupling of multiple microresonators and show that the 3-dB bandwidth of optical switching elements can be widened nearly linearly with the ring number. Evaluation results for different real MPSoC applications show that, on the basis of thermal tuning, the optimal device setting improves the average power efficiency by 54% to 1.2 pJ/bit when chip temperature reaches 85$^{\circ}$C. The findings in this paper can help support the further development of this emerging technology.
Yaoyao Ye, Jiang Xu 0001, Xiaowen Wu, Wei Zhang 0012, Xuan Wang 0001, Mahdi Nikdast, Zhehui Wang, Weichen Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2012 A Torus-Based Hierarchical Optical-Electronic Network-on-Chip for Multiprocessor System-on-Chip
abstract
Networks-on-chip (NoCs) are emerging as a key on-chip communication architecture for multiprocessor systems-on-chip (MPSoCs). Optical communication technologies are introduced to NoCs in order to empower ultra-high bandwidth with low power consumption. However, in existing optical NoCs, communication locality is poorly supported, and the importance of floorplanning is overlooked. These significantly limit the power efficiency and performance of optical NoCs. In this work, we address these issues and propose a torus-based hierarchical hybrid optical-electronic NoC, called THOE. THOE takes advantage of both electrical and optical routers and interconnects in a hierarchical manner. It employs several new techniques including floorplan optimization, an adaptive power control mechanism, low-latency control protocols, and hybrid optical-electrical routers with a low-power optical switching fabric. Both of the unfolded and folded torus topologies are explored for THOE. Based on a set of real MPSoC applications, we compared THOE with a typical torus-based optical NoC as well as a torus-based electronic NoC in 45nm on a 256-core MPSoC, using a SystemC-based cycle-accurate NoC simulator. Compared with the matched electronic torus-based NoC, THOE achieves 2.46X performance and 1.51X network switching capacity utilization, with 84% less energy consumption. Compared with the optical torus-based NoC, THOE achieves 4.71X performance and 3.05X network switching capacity utilization, while reducing 99% of energy consumption. Besides real MPSoC applications, a uniform traffic pattern is also used to show the average packet delay and network throughput of THOE. Regarding hardware cost, THOE reduces 75% of laser sources and half of optical receivers compared with the optical torus-based NoC.
Yaoyao Ye, Jiang Xu 0001, Xiaowen Wu, Wei Zhang 0012, Weichen Liu 0001, Mahdi Nikdast
ACM J. Emerg. Technol. Comput. Syst.1
2011 Satisfiability Modulo Graph Theory for Task Mapping and Scheduling on Multiprocessor Systems
abstract
Task graph scheduling on multiprocessor systems is a representative multiprocessor scheduling problem. A solution to this problem consists of the mapping of tasks to processors and the scheduling of tasks on each processor. Optimal solution can be obtained by exploring the entire design space of all possible mapping and scheduling choices. Since the problem is NP-hard, scalability becomes the main concern in solving the problem optimally. In this paper, a SAT-based optimization framework is proposed to address this problem, in which SAT solver is enhanced by integrating with a scheduling analysis tool in a branch and bound manner to prune the solution space efficiently. Performance evaluation results show that our technique has average performance improvement in more than an order of magnitude compared to state-of-the-art techniques. We further build a cycle-accurate network-on-chip simulator based on SystemC to verify the effectiveness of the proposed technique on realistic multiprocessor systems.
Weichen Liu 0001, Zonghua Gu 0001, Jiang Xu 0001, Xiaowen Wu, Yaoyao Ye
IEEE Trans. Parallel Distributed Syst.5
2010 Crosstalk noise and bit error rate analysis for optical network-on-chip
abstract
Crosstalk noise is an intrinsic characteristic of photonic devices used by optical networks-on-chip (ONoCs) as well as a potential issue. For the first time, this paper analyzed and modeled the crosstalk noise, signal-to-noise ratio (SNR), and bit error rate (BER) of optical routers and ONoCs. The analytical models for crosstalk noise, minimum SNR, and maximum BER in meshbased ONoCs are presented. An automated crosstalk analyzer for optical routers is developed. We find that crosstalk noise significantly limits the scalability of ONoCs. For example, due to crosstalk noise, the maximum BER is 10-3 on the 8x8 mesh-based ONoC using an optimized crossbar-based optical router. To achieve the BER of 10-9 for reliable transmissions, the maximum ONoC size is 6x6. A novel compact high-SNR optical router is proposed to improve the maximum ONoC size to 8x8.
Yiyuan Xie, Mahdi Nikdast, Jiang Xu 0001, Wei Zhang 0012, Qi Li 0013, Xiaowen Wu, Yaoyao Ye, Xuan Wang 0001, Weichen Liu 0001
DAC7