Zhehui Wang

dblp:06/9981 · DBLP profile ↗
← Back
50ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0002-7139-724XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 42 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Is quantum optimization ready? An effort towards neural network compression using adiabatic quantum computing
Zhehui Wang, Benjamin Chen Ming Choong, Tian Huang, Daniel Gerlinghoff, Rick Siow Mong Goh, Cheng Liu 0008, Tao Luo 0014
Future Gener. Comput. Syst.1
2025 Coflex: Enhancing HW-NAS with Sparse Gaussian Processes for Efficient and Scalable DNN Accelerator Design
abstract
Hardware-Aware Neural Architecture Search (HW-NAS) is an efficient approach to automatically co-optimizing neural network performance and hardware energy efficiency, making it particularly useful for the development of Deep Neural Network accelerators on the edge. However, the extensive search space and high computational cost pose significant challenges to its practical adoption. To address these limitations, we propose Coflex, a novel HW-NAS framework that integrates the Sparse Gaussian Process (SGP) with multi-objective Bayesian optimization. By leveraging sparse inducing points, Coflex reduces the GP kernel complexity from cubic to near-linear with respect to the number of training samples, without compromising optimization performance. This enables scalable approximation of large-scale search space, substantially decreasing computational overhead while preserving high predictive accuracy. We evaluate the efficacy of Coflex across various benchmarks, focusing on accelerator-specific architecture. Our experimental results show that Coflex outperforms state-of-the-art methods in terms of network accuracy and Energy-Delay-Product, while achieving a computational speed-up ranging from 1.9× to 9.5×.
Yinhui Ma, Tomomasa Yamasaki, Zhehui Wang, Tao Luo 0014, Bo Wang 0020
ICCAD3
2025 Enabling Energy-Efficient Deployment of Large Language Models on Memristor Crossbar: A Synergy of Large and Small
abstract
Large language models (LLMs) have garnered substantial attention due to their promising applications in diverse domains. Nevertheless, the increasing size of LLMs comes with a significant surge in the computational requirements for training and deployment. Memristor crossbars have emerged as a promising solution, which demonstrated a small footprint and remarkably high energy efficiency in computer vision (CV) models. Memristors possess higher density compared to conventional memory technologies, making them highly suitable for effectively managing the extreme model size associated with LLMs. However, deploying LLMs on memristor crossbars faces three major challenges. First, the size of LLMs increases rapidly, already surpassing the capabilities of state-of-the-art memristor chips. Second, LLMs often incorporate multi-head attention blocks, which involve non-weight stationary multiplications that traditional memristor crossbars cannot support. Third, while memristor crossbars excel at performing linear operations, they are not capable of executing complex nonlinear operations in LLM such as softmax and layer normalization. To address these challenges, we present a novel architecture for the memristor crossbar that enables the deployment of state-of-the-art LLM on a single chip or package, eliminating the energy and time inefficiencies associated with off-chip communication. Our testing on BERT showed negligible accuracy loss. Compared to traditional memristor crossbars, our architecture achieves enhancements of up to in area overhead and in energy consumption. Compared to modern TPU/GPU systems, our architecture demonstrates at least a reduction in the area-delay product and a significant 69% energy consumption reduction.
Zhehui Wang, Tao Luo 0014, Cheng Liu 0008, Weichen Liu 0001, Rick Siow Mong Goh, Weng-Fai Wong
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 RBFleX-NAS: Training-Free Neural Architecture Search Using Radial Basis Function Kernel and Hyperparameter Detection
abstract
Neural architecture search (NAS) is an automated technique to design optimal neural network architectures for a specific workload. Conventionally, evaluating candidate networks in NAS involves extensive training, which requires significant time and computational resources. To address this, training-free NAS has been proposed to expedite network evaluation with minimal search time. However, state-of-the-art training-free NAS algorithms struggle to precisely distinguish well-performing networks from poorly performing networks, resulting in inaccurate performance predictions and consequently suboptimal top-one network accuracy. Moreover, they are less effective in activation function exploration. To tackle the challenges, this article proposes RBFleX-NAS, a novel training-free NAS framework that accounts for both activation outputs and input features of the last layer with a radial basis function (RBF) kernel. We also present a detection algorithm to identify optimal hyperparameters using the obtained activation outputs and input feature maps. We verify the efficacy of RBFleX-NAS over a variety of NAS benchmarks. RBFleX-NAS significantly outperforms state-of-the-art training-free NAS methods in terms of top-one accuracy, achieving this with short search time in NAS-Bench-201 and NAS-Bench-SSS. In addition, it demonstrates a higher Kendall correlation compared to layer-based training-free NAS algorithms. Furthermore, we propose the neural network activation function benchmark (NAFBee), a new activation design space that extends the activation type to encompass various commonly used functions. In this extended design space, RBFleX-NAS demonstrates its superiority by accurately identifying the best-performing network during activation function search, providing a significant advantage over other NAS algorithms.
Tomomasa Yamasaki, Zhehui Wang, Tao Luo 0014, Niangjun Chen, Bo Wang 0020
IEEE Trans. Neural Networks Learn. Syst.2
2024 IMI: In-memory Multi-job Inference Acceleration for Large Language Models
abstract
Large Language Models (LLMs) are increasingly used in various applications but are computationally complex and energy-consuming due to the high volume of off-chip memory accesses. Processing-in-Memory (PIM) has emerged as a potential solution for efficient inference. However, existing PIM accelerators designed for deep neural networks (DNNs) aren’t suitable for LLMs because of differences in operations, input sizes, and job completion times. This leads to performance issues like head-of-line blocking where an earlier job can monopolize resources at the expense of later jobs, and low resource utilization. To improve efficiency, a time-multiplex solution and job colocation accelerator could be beneficial. However, facilitating multi-job execution with in-memory acceleration is challenging due to limitations in memristor architecture, inefficiency of non-stationary weight programming, difficulty in dynamic partitioning of hardware resources, and complexity in dynamic job scheduling. This work proposes a PIM-based LLM accelerator to enable the concurrent LLM inference job execution. The experiment shows that IMI can significantly improve resource utilization as well as the rate of satisfying the service level requirements of jobs.
Bin Gao 0013, Zhehui Wang, Zhuomin He, Tao Luo 0014, Weng-Fai Wong, Zhi Zhou 0006
ICPP2
2024 Efficient Spiking Neural Networks With Radix Encoding
abstract
Spiking neural networks (SNNs) have advantages in latency and energy efficiency over traditional artificial neural networks (ANNs) due to their event-driven computation mechanism and the replacement of energy-consuming weight multiplication with addition. However, to achieve high accuracy, it usually requires long spike trains to ensure accuracy, usually more than 1000 time steps. This offsets the computation efficiency brought by SNNs because a longer spike train means a larger number of operations and larger latency. In this article, we propose a radix-encoded SNN, which has ultrashort spike trains. Specifically, it is able to use less than six time steps to achieve even higher accuracy than its traditional counterpart. We also develop a method to fit our radix encoding technique into the ANN-to-SNN conversion approach so that we can train radix-encoded SNNs more efficiently on mature platforms and hardware. Experiments show that our radix encoding can achieve 25× improvement in latency and 1.7% improvement in accuracy compared to the state-of-the-art method using the VGG-16 network on the CIFAR-10 dataset.
Zhehui Wang, Xiaozhe Gu, Rick Siow Mong Goh, Joey Tianyi Zhou, Tao Luo 0014
IEEE Trans. Neural Networks Learn. Syst.1
2024 EDCompress: Energy-Aware Model Compression for Dataflows
abstract
Edge devices demand low energy consumption, cost, and small form factor. To efficiently deploy convolutional neural network (CNN) models on the edge device, energy-aware model compression becomes extremely important. However, existing work did not study this problem well because of the lack of considering the diversity of dataflow types in hardware architectures. In this article, we propose EDCompress (EDC), an energy-aware model compression method for various dataflows. It can effectively reduce the energy consumption of various edge devices, with different dataflow types. Considering the very nature of model compression procedures, we recast the optimization process to a multistep problem and solve it by reinforcement learning algorithms. We also propose a multidimensional multistep (MDMS) optimization method, which shows higher compressing capability than the traditional multistep method. Experiments show that EDC could improve 20x, 17x, and 26x energy efficiency in VGG-16, MobileNet, and LeNet-5 networks, respectively, with negligible loss of accuracy. EDC could also indicate the optimal dataflow type for specific neural networks in terms of energy consumption, which can guide the deployment of CNN on hardware.
Zhehui Wang, Tao Luo 0014, Rick Siow Mong Goh, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2024 Optimizing for In-Memory Deep Learning With Emerging Memory Technology
abstract
In-memory deep learning executes neural network models where they are stored, thus avoiding long-distance communication between memory and computation units, resulting in considerable savings in energy and time. In-memory deep learning has already demonstrated orders of magnitude higher performance density and energy efficiency. The use of emerging memory technology (EMT) promises to increase density, energy, and performance even further. However, EMT is intrinsically unstable, resulting in random data read fluctuations. This can translate to nonnegligible accuracy loss, potentially nullifying the gains. In this article, we propose three optimization techniques that can mathematically overcome the instability problem of EMT. They can improve the accuracy of the in-memory deep learning model while maximizing its energy efficiency. Experiments show that our solution can fully recover most models' state-of-the-art (SOTA) accuracy and achieves at least an order of magnitude higher energy efficiency than the SOTA.
Zhehui Wang, Tao Luo 0014, Rick Siow Mong Goh, Wei Zhang 0012, Weng-Fai Wong
IEEE Trans. Neural Networks Learn. Syst.1
2023 MA-BERT: Towards Matrix Arithmetic-only BERT Inference by Eliminating Complex Non-Linear Functions
Neo Wei Ming, Zhehui Wang, Cheng Liu 0008, Rick Siow Mong Goh, Tao Luo 0014
ICLR2
2023 PATCorrect: Non-autoregressive Phoneme-augmented Transformer for ASR Error Correction
Ziji Zhang 0003, Zhehui Wang, Rajesh Kamma, Sharanya Eswaran, Narayanan Sadagopan
INTERSPEECH2
2022 A Resource-efficient Spiking Neural Network Accelerator Supporting Emerging Neural Encoding
abstract
Spiking neural networks (SNNs) recently gained momentum due to their low-power multiplication-free computing and the closer resemblance of biological processes in the nervous system of humans. However, SNNs require very long spike trains (up to 1000) to reach an accuracy similar to their artificial neural network (ANN) counterparts for large models, which offsets efficiency and inhibits its application to low-power systems for real-world use cases. To alleviate this problem, emerging neural encoding schemes are proposed to shorten the spike train while maintaining the high accuracy. However, current accelerators for SNN cannot well support the emerging encoding schemes. In this work, we present a novel hardware architecture that can efficiently support SNN with emerging neural encoding. Our implementation features energy and area efficient processing units with increased parallelism and reduced memory accesses. We verified the accelerator on FPGA and achieve 25% and 90% improvement over previous work in power consumption and latency, respectively. At the same time, high area efficiency allows us to scale for large neural network models. To the best of our knowledge, this is the first work to deploy the large neural network model VGG on physical FPGA-based neuromorphic hardware.
Daniel Gerlinghoff, Zhehui Wang, Xiaozhe Gu, Rick Siow Mong Goh, Tao Luo 0014
DATE2
2022 E3NE: An End-to-End Framework for Accelerating Spiking Neural Networks With Emerging Neural Encoding on FPGAs
abstract
Compiler frameworks are crucial for the widespread use of FPGA-based deep learning accelerators. They allow researchers and developers, who are not familiar with hardware engineering, to harness the performance attained by domain-specific logic. There exists a variety of frameworks for conventional artificial neural networks. However, not much research effort has been put into the creation of frameworks optimized for spiking neural networks (SNNs). This new generation of neural networks becomes increasingly interesting for the deployment of AI on edge devices, which have tight power and resource constraints. Our end-to-end framework E3NE automates the generation of efficient SNN inference logic for FPGAs. Based on a PyTorch model and user parameters, it applies various optimizations and assesses trade-offs inherent to spike-based accelerators. Multiple levels of parallelism and the use of an emerging neural encoding scheme result in an efficiency superior to previous SNN hardware implementations. For a similar model, E3NE uses less than 50% of hardware resources and 20% less power, while reducing the latency by an order of magnitude. Furthermore, scalability and generality allowed the deployment of the large-scale SNN models AlexNet and VGG.
Daniel Gerlinghoff, Zhehui Wang, Xiaozhe Gu, Rick Siow Mong Goh, Tao Luo 0014
IEEE Trans. Parallel Distributed Syst.2
2021 Reduce Loss and Crosstalk in Integrated Silicon-Photonic Multistage Switching Fabrics Through Multichip Partition
abstract
With the increasing popularity of data-intensive applications in data centers, the switching fabric in the internode network becomes significant. Silicon-photonic switching fabrics have a bright future in data centers, which offer high bandwidth, high energy efficiency, and low latency. However, integrating a high radix multistage switching fabric in a single chip faces challenges. A large number of waveguide crossings on the silicon photonic die causes massive power loss and introduces a tremendous amount of crosstalk noise. In this article, we propose a chip partition optimization platform (POP), which can decrease the number of waveguide crossings and shorten the on-chip traversal distance of optical signals. Our algorithms can effectively reduce the power loss and crosstalk noise in silicon-photonic multistage switching fabrics, and help to improve the signal integrity. For example, compared with the common design, POP can achieve 33-dB improvement on average power loss, 42-dB improvement on the worst-case power loss, and 39-dB improvement on the worst-case signal to noise ratio, in a$1024\times1024$butterfly based silicon-photonic switching fabric.
Zhehui Wang, Jiang Xu 0001, Jun Feng 0008, Shixi Chen, Xuanqi Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis
abstract
We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. We demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.
Jiayi Xu 0001, Hanqi Guo 0001, Han-Wei Shen, Mukund Raj, Xueyun Wang, Xueqiao Xu, Zhehui Wang, Tom Peterka
IEEE Trans. Vis. Comput. Graph.7
2020 Modeling and Analysis of Optical Modulators Based on Free-Carrier Plasma Dispersion Effect
abstract
Silicon photonic networks are revolutionizing computing systems by improving the energy efficiency, bandwidth, and latency of data movements. Optical modulators, such as microresonators (MRs) and Mach–Zehnder interferometers (MZIs), are the basic building blocks of silicon photonic networks. This paper proposes a SPICE-compatible electro-optical co-simulation model, basic optical switch integration model (BOSIM), to systematically study optical modulators using PN, PIN, and metal–insulator–silicon (MIS) capacitor device technologies. BOSIM holistically models both transient and steady state properties, such as switching speed, power, transmission spectrum, area, and carrier distribution. BOSIM is validated by the measured data from eight research groups and companies. Compared to MRs, BOSIM shows MZIs are fast, with a high extinction ratio and large bandwidth but in the sacrifice of loss, energy, and area. Using a PIN diode over a PN diode can save area, but retain the loss and energy, while an MIS capacitor has the shock response of carrier distribution in a narrow range and is marginalized gradually. For instance, an MZI can achieve a$2.5 {\times }$bit rate,$6.06{\times }$extinction ratio,$71.04 {\times }\,\,3$-dB bandwidth, but costs at least$1.93 {\times }$passing loss,$1.46 {\times }$energy consumption, and$16.67 {\times }$area, compared with MR.
Xuanqi Chen, Yi-Shing Chang, Jiang Xu 0001, Jun Feng 0008, Peng Yang 0003, Zhehui Wang, Luan H. K. Duong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2020 CAMON: Low-Cost Silicon Photonic Chiplet for Manycore Processors
abstract
While many new applications prefer manycore processor with a large number of cores, the exploding communications among multiple cores, caches, and off-chip memories is posing a fundamental challenge on manycore designs. Silicon photonics-based interconnection network promises high bandwidth, low latency, and high energy efficiency, and can potentially meet the communication requirements of manycore processors. In this paper, we propose CAMON, a small low-cost silicon photonic chiplet integrated into the manycore processor package. CAMON chiplet can effectively alleviate the communication bottlenecks of manycore processors and improve the energy efficiency of data movement, especially for large-scale systems. We develop a distributed arbitration system, a low-power low-latency optical interface, and an off-chip laser preactivation mechanism for CAMON. The experimental results show that compared with the electrical network, CAMON can improve the full-system performance per energy by 4.6×, speedup the manycore processor by 2.7×, and save the area of the processor die by 3%, in a 512-core system.
Zhehui Wang, Jiang Xu 0001, Yi-Shing Chang, Jun Feng 0008, Xuanqi Chen, Shixi Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 A Cross-Layer Optimization Framework for Integrated Optical Switches in Data Centers
abstract
The advancement of silicon photonics promises integrated optical switches to provide high-bandwidth, low-latency, and low-power communications in data centers. An optical switch’s loss limits its scale and affects the energy efficiency of the switch system. In this paper, we present cross-layer optical switch optimization (CLOSO), a cross-layer optimization (CLO) framework, based on not only photonic device models at the physical layer but also optical switch models at the fabric layer. With the proposed framework, optimal losses of optical switches can be evaluated efficiently, and the corresponding losses and design parameters of photonic devices can be obtained. Using CLOSO, we optimize four categories of integrated optical switches, Crossbar, PILOSS, DRAGON, and FODON, and compare them regarding their optimal worst-case loss with variation of the switch scale and data rate of signals. Furthermore, system-level evaluations of the optimized optical switches are performed, demonstrating a significant improvement of energy efficiency from the CLO. For instance, CLOSO helps to reduce the energy consumption of a 64-port DRAGON and FODON to as low as 6 pJ/bit and that of a 128-port DRAGON and FODON to as low as 10 pJ/bit. The investigation of 128-port switches also shows the necessity of adaptive power control on lasers for high-radix integrated optical switches. Through quantitative analyses and comparisons, CLOSO shows the capability of facilitating initial design exploration of optical switches and paves the way to fair evaluations and comparisons of switch systems in data centers.
Peng Yang 0003, Yi-Shing Chang, Jiang Xu 0001, Xuanqi Chen, Zhehui Wang, Jun Feng 0008
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Multidomain Inter/Intrachip Silicon Photonic Networks for Energy-Efficient Rack-Scale Computing Systems
abstract
Rack-scale computing systems are promising to undertake the emerging large-scale applications by distributing massive tasks to processing cores. The communication and coordination efficiency of these tasks and resources directly affect the system performance and energy consumption. Silicon photonic interconnects are expected to address the communication and system power consumption challenges imposed on rack-scale systems. However, the control for optical interconnects can cause server performance degradation if not properly designed, especially for the complicated and time-consuming multidomain networks. In this paper, we study the optical interconnects for rack-scale computing systems and propose a new communication flow and control scheme for the efficient coordination of distributed resources. Particularly, we first propose a forward propagation strategy that parallels the path reservation process with the distributed tasks connection setup. Second, we develop a pre-emptive chain feedback (PCF) scheme to optimize multidomain path reservation. The PCF scheme pre-emptively allocates network resources with the help of multicell reservation window and quickly releases resources with a feedback mechanism. This solution increases the network resources utilization and task coordination efficiency while minimizing path reservation overheads. Comparing to the baseline InfiniBand network fabric and handshake scheme, PCF can improve network throughput greatly under uniform and hotspot traffic patterns. Realistic benchmark results show that the PCF scheme on average reduces 52% and 60% energy consumption per unit system performance than InfiniBand and the handshake scheme for a 256-node rack system.
Peng Yang 0003, Zhehui Wang, Jiang Xu 0001, Yi-Shing Chang, Xuanqi Chen, Rafael Kioji Vivas Maeda, Jun Feng 0008
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Chip-Specific Power Delivery and Consumption Co-Management for Process-Variation-Aware Manycore Systems Using Reinforcement Learning
abstract
Energy efficiency has become a critical design metric for high-performance systems. Various power management techniques have been proposed for the processor cores such as dynamic voltage and frequency scaling (DVFS), whereas few solutions consider the power losses suffered on the power delivery system (PDS), despite the fact that they have a significant impact on the overall energy efficiency of the system. With the explosive growth of system complexity and highly dynamic workloads variations, it is also challenging to find the optimal power management policies which can effectively match the power delivery with the power consumption. In addition, process variations (PVs) add heterogeneity to systems and make traditional power management methods less effective. To tackle the above problems, we propose a reinforcement-learning-based Chip-Specific Power co-Management (CSPM) scheme for PV-aware manycore systems. Both PDS and processor cores are jointly adjusted by distributed agents with modular Q-learning to improve the overall energy efficiency of the system. System characteristics are naturally included in the learning process to obtain chip-specific policies. Experimental results show that when applied to PV-aware manycore systems with a hybrid PDS constructed by both on- and off-chip voltage regulators, the proposed method achieves a 60.1% reduction of the overall energy delay product (EDP) of the system, on average, compared to a traditional DVFS approach.
Haoran Li 0002, Zhongyuan Tian, Jiang Xu 0001, Rafael Kioji Vivas Maeda, Zhehui Wang
IEEE Trans. Very Large Scale Integr. Syst.5
2019 Systematic Exploration of High-Radix Integrated Silicon Photonic Switches for Datacenters
abstract
High-radix integrated silicon photonic switches promise ultrahigh bandwidth communications required by next generation data centers. To holistically explore the characteristics of high-radix integrated optical switches, this work systematically studies the latency, throughput and energy consumption, with detailed models and various system configurations. Three categories of space switches, blocking, rearrangeable non-blocking and strictly non-blocking switches, are investigated, together with one of the widely used wavelength switches, arrayed waveguide grating router (AWGR). The work paves the ways to automatically optimize high-radix integrated silicon photonic switches.
Jun Feng 0008, Xuanqi Chen, Zhehui Wang, Shixi Chen, Jiang Xu 0001
ICCAD4
2019 Crosstalk Noise Reduction Through Adaptive Power Control in Inter/Intra-Chip Optical Networks
abstract
In recent years, optical interconnection networks have been proposed in order to achieve the ultrahigh bandwidth and low latency requirements for inter/intra-chip communication. In these optical interconection networks, series of basic optical elements are employed. Via these series of optical elements, the intrinsic crosstalk noise is generated. With a large scale of these optical elements, the signal-to-noise ratio (SNR) of an optical interconnect can be reduced by this crosstalk noise. In this paper, we utilize the adaptive power control (APC) to enhance the SNR under the crosstalk noise constraints. APC has been known to save energy and reduce power consumption. We apply this technique in one of the inter/intra-chip optical interconnect called I2CON. A new cluster design, namely the Beam cluster, is also introduced. Results have demonstrated that the APC can help to reduce crosstalk noise, hence, the overall SNR is improved. Comparison results have also indicated the further improvement of SNR in I2CON using Beam cluster when APC is applied.
Luan H. K. Duong, Peng Yang 0003, Yi-Shing Chang, Jiang Xu 0001, Zhehui Wang, Xuanqi Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2018 RSON: An inter/intra-chip silicon photonic network for rack-scale computing systems
abstract
The increasing demand for more computational power from scientific computing, big data processing, and machine learning is pushing the development of HPC (high-performance computing) systems. As the basic HPC building blocks, modularized server racks with a large number of multicore nodes are facing performance and energy efficiency challenges. This paper proposes RSON, an optical network for rack-scale computing systems. RSON connects processor cores, caches, local memories, and remote memories through a novel inter/intra-chip silicon photonic network architecture. We develop a low-latency scalable channel partition and low-power dynamic path priority control scheme for RSON. Experimental results show that RSON can help rack-scale computing systems achieve up to 6.8X higher performance under the same energy consumption than state-of-the-art systems under the latest APEX (application performance at extreme scale) benchmarks.
Peng Yang 0003, Zhengbin Pang, Zhehui Wang, Xuanqi Chen, Luan H. K. Duong, Jiang Xu 0001
DATE4
2017 Modular reinforcement learning for self-adaptive energy efficiency optimization in multicore system
abstract
Energy-efficiency is becoming increasingly important to modern computing systems with multi-/many-core architectures. Dynamic Voltage and Frequency Scaling (DVFS), as an effective low-power technique, has been widely applied to improve energy-efficiency in commercial multi-core systems. However, due to the large number of cores and growing complexity of emerging applications, it is difficult to efficiently find a globally optimized voltage/frequency assignment at runtime. In order to improve the energy-efficiency for the overall multicore system, we propose an online DVFS control strategy based on core-level Modular Reinforcement Learning (MRL) to adaptively select appropriate operating frequencies for each individual core. Instead of focusing solely on the local core conditions, MRL is able to make comprehensive decisions by considering the running-states of multiple cores without incurring exponential memory cost which is necessary in traditional Monolithic Reinforcement Learning (RL). Experimental results on various realistic applications and different system scales show that the proposed approach improves up to 28% energy-efficiency compared to the recent individual-RL approach.
Zhe Wang 0003, Zhongyuan Tian, Jiang Xu 0001, Rafael Kioji Vivas Maeda, Haoran Li 0002, Peng Yang 0003, Zhehui Wang, Luan H. K. Duong, Xuanqi Chen
ASP-DAC7
2017 MOCA: an Inter/Intra-Chip Optical Network for Memory
abstract
The memory wall problem is due to the imbalanced developments and separation of processors and memories. It is becoming acute as more and more processor cores are integrated into a single chip and demand higher memory bandwidth through limited chip pins. Optical memory interconnection network (OMIN) promises high bandwidth, bandwidth density, and energy efficiency, and can potentially alleviate the memory wall problem. In this paper, we propose an optical inter/intra-chip processor-memory communication architecture, called MOCA. Experimental results and analysis show that MOCA can significantly improve system performance and energy efficiency. For example, comparing to Hybrid Memory Cube (HMC), MOCA can speedup application execution time by 2.6x, reduce communication latency by 75%, and improve energy efficiency by 3.4x for 256-core processors in 7 nm technology.
Zhehui Wang, Zhengbin Pang, Peng Yang 0003, Jiang Xu 0001, Xuanqi Chen, Rafael Kioji Vivas Maeda, Luan H. K. Duong, Haoran Li 0002, Zhe Wang 0003
DAC1
2017 Energy-Efficient Power Delivery System Paradigms for Many-Core Processors
abstract
The design of power delivery system plays a crucial role in guaranteeing the proper functionality of many-core processor systems. The power loss suffered on power delivery has become a salient part of total power consumption, and the energy efficiency of a highly dynamic system has been significantly challenged. Being able to achieve a fast response time and multiple voltage domain control, on-chip voltage regulators (VRs) have become popular choices to enable fine-grain power management, which also enlarge the design space of power delivery systems. This paper analytically studies different power delivery system paradigms and power management schemes in terms of energy efficiency, area overhead, and power pin occupation. The analysis shows that compared to the conventional paradigm with off-chip VRs, hybrid paradigms with both on-chip and off-chip VRs are able to maintain high efficiency in a larger range of workloads, though they suffer from low efficiency at light workload. Employed with the quantized power management scheme, the hybrid paradigm can improve the system energy efficiency at light workload by a maximum of 136% compared to the traditional load balanced scheme. Besides this, the in-package (iP) hybrid paradigm further shows its advantage in reducing the physical overheads. The results reveal that at 120 W workload, it occupies only a 10.94% total footprint area or 39.07% power pins of that of the off-chip paradigm. We conclude that the iP hybrid paradigm achieves the best tradeoffs between efficiency, physical overhead, and realization of fine-grain power management.
Haoran Li 0002, Xuan Wang 0001, Jiang Xu 0001, Zhe Wang 0003, Rafael Kioji Vivas Maeda, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2016 Inter/intra-chip optical interconnection network: opportunities, challenges, and implementations
abstract
Recent advances in photonics technologies have made optical interconnection network an attractive option for computing systems from high-performance computers and data centers to automobiles and cellphones. Optical interconnection network promises ultra-high bandwidth, low latency, and great energy efficiency to alleviate the inter-rack, intra-rack, intraboard, and intra-chip communication bottlenecks in multiprocessor systems. Silicon-based photonics technologies piggyback onto developed silicon fabrication processes to provide viable and cost-effective solutions. Both industry and academia have invested significant efforts to develop and commercialize optical interconnection network technologies. This paper reviews the latest progresses and provides insights into the challenges and future developments.
Peng Yang 0003, Shigeru Nakamura, Kenichiro Yashiki, Zhehui Wang, Luan H. K. Duong, Xuanqi Chen, Yuichi Nakamura 0002, Jiang Xu 0001
NOCS4
2016 Coherent and Incoherent Crosstalk Noise Analyses in Interchip/Intrachip Optical Interconnection Networks
abstract
Recently, interchip/intrachip optical interconnection networks have been proposed for ultrahigh-bandwidth and low-latency communications. These networks employ the microresonators (MRs) to modulate, direct, or detect the optical signal. However, utilized MRs suffer from intrinsic crosstalk noise and signal power loss, degrading the network efficiency via the signal-to-noise ratio (SNR). The amount of crosstalk noise and signal power loss may differ from network to network. Hence, there exists a need to systematically analyze the effect of the crosstalk noise and the power loss issues. In this paper, we have developed the analytical models considering both coherent and incoherent crosstalk for both the interchip and intrachip optical networks. The interchip/intrachip optical interconnection networks—the$\text{I}^{2}$CON—are analyzed as a case study. The quantitative results on the individual networks have demonstrated that the architectural design determines the impact of crosstalk on the SNR. We have also demonstrated that the optical interconnection networks with interchip/intrachip interconnects result in better bit error rate (BER) compared with that of only intrachip interconnect. Our analyses of the worst case can be utilized as a platform to compare the realistic performance among different optical interconnection networks via the degradation of SNR/BER and data bandwidth.
Luan H. K. Duong, Zhehui Wang, Mahdi Nikdast, Jiang Xu 0001, Peng Yang 0003, Zhe Wang 0003, Rafael Kioji Vivas Maeda, Haoran Li 0002, Xuan Wang 0001, Sébastien Le Beux, Yvain Thonnart
IEEE Trans. Very Large Scale Integr. Syst.2
2016 An Adaptive Process-Variation-Aware Technique for Power-Gating-Induced Power/Ground Noise Mitigation in MPSoC
abstract
Power gating (PG) is one of the most effective techniques to reduce the leakage power in multiprocessor system-on-chips (MPSoCs). However, the power-mode transition during the PG period of an individual processing unit (PU) will introduce serious power/ground (P/G) noise to the neighboring PUs. As technology scales, the P/G noise problem becomes a severe reliability threat to MPSoCs. At the same time, the increasing manufacturing process variations (PVs) also bring uncertainties to the P/G noise problem and make it difficult to predict and mitigate. To tackle this problem, in this paper, we analyze the PG-induced P/G noise in the presence of PVs and propose a hardware–software collaborated runtime technique to adaptively protect PUs from P/G noise. Sensor network-on-chip is used to gather noise information and coordinate different system components. An online PV-aware algorithm is developed to effectively decide the noise impact range and arrange protections for affected PUs based on the collected noise information. We evaluate the proposed technique through cycle-level Monte Carlo simulations of NoC-based MPSoCs in different scales. The experimental results on various realistic applications show that our technique could achieve comparable reliability to the most reliable static technique while improve on average 3.78%–29.5% the system energy efficiency and reduce 15.7%–70.4% the performance penalty on different MPSoC scales.
Zhe Wang 0003, Xuan Wang 0001, Jiang Xu 0001, Haoran Li 0002, Rafael Kioji Vivas Maeda, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong
IEEE Trans. Very Large Scale Integr. Syst.6
2016 A Holistic Modeling and Analysis of Optical-Electrical Interfaces for Inter/Intra-chip Interconnects
abstract
With the fast development of inter/intra-chip optical interconnects, the gap between the data rates of electrical interconnects and optical interconnects is continuously increasing. Electrical–optical (E-O) interfaces and optical–electrical (O-E) interfaces are a pair of components that convert data between parallel electrical interconnects and serial optical interconnects. This paper holistically models and analyzes E-O and O-E interfaces in terms of energy consumption, area, and latency. Traditional interfaces, where data are converted between parallel and serial ports by serializers and deserializers (SerDes), are studied. A new type of E-O and O-E interface, which serializes and deserializes data by optical weaving technologies, are proposed alongside. Traditional interfaces will become a bottleneck for the further development of optical interconnects in the near future because of the high energy consumption and large area of SerDes necessitating new technologies. Our analysis shows that optical weaving interfaces have a better overall performance than traditional interfaces. For example, if there are 64 parallel electrical interconnects and four optical wavelengths, optical weaving interfaces can achieve a 81.6% improvement in energy consumption and a 40.8% improvement in area, compared with traditional interfaces.
Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Luan H. K. Duong, Xuan Wang 0001, Zhe Wang 0003, Haoran Li 0002, Rafael Kioji Vivas Maeda
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Improve Chip Pin Performance Using Optical Interconnects
abstract
With the fast development of processor chips, power-efficient, high-bandwidth, and low-latency interchip interconnects become more and more important. Studies show that the bandwidth of traditional parallel interconnects with low I/O clock frequencies will become bottlenecks in the near future. To solve this problem, two types of high-bandwidth interchip interconnects are developed. Low-swing differential electrical interconnects have widely been used in high-speed I/O designs. On the other hand, optical interconnects promise high bandwidth, low latency, and could improve the chip pin performance for manycore processors. They are becoming potential alternatives for electrical interconnects. This paper systematically models these two types of interconnects in terms of crosstalk noises, attenuation, and receiver sensitivities. Based on the proposed models, we developed optical and electrical interfaces and links (OEIL) and an analysis tool for OEIL. The OEIL can be used to analyze the energy consumption, bandwidth density, and latency of interconnects. Analytical models are verified by the results of published experiments. It shows that the optical interconnects have much higher bandwidth densities than the electrical interconnects. With this feature, the optical interconnects can significantly reduce I/O pin count compared with the electrical interconnects. For example, they can save at least 92% signal pins when connecting chips more than 25 cm (10 in) apart. The energy consumption of optical interconnects is comparable with that of electrical interconnects, and the latency of polymer waveguide-based optical interconnects is 18% less than that of electrical interconnect.
Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Xuan Wang 0001, Zhe Wang 0003, Luan H. K. Duong, Rafael Kioji Vivas Maeda, Haoran Li 0002
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Alleviate chip I/O pin constraints for multicore processors through optical interconnects
abstract
Chip I/O pins are an increasingly limited resource and significantly affect the performance, power and cost of multicore processors. Optical interconnects promise low power and high bandwidth, and are potential alternatives to electrical interconnects. This work systematically developed a set of analytical models for electrical and optical interconnects to study their structures, receiver sensitivities, crosstalk noises, and attenuations. We verified the models by published implementation results. The analytical models quantitatively identified the advantages of optical interconnects in terms of bandwidth, energy consumption, and transmission distance. We showed that optical interconnects can significantly reduce chip pin counts. For example, compared to electrical interconnects, optical interconnects can save at least 92% signal pins when connecting chips more than 25 cm (10 inches) apart.
Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Xuan Wang 0001, Zhe Wang 0003, Luan H. K. Duong, Haoran Li 0002, Rafael Kioji Vivas Maeda, Xiaowen Wu, Yaoyao Ye, Qinfen Hao
ASP-DAC1
2015 Coherent crosstalk noise analyses in ring-based optical interconnects
Luan H. K. Duong, Mahdi Nikdast, Jiang Xu 0001, Zhehui Wang, Yvain Thonnart, Sébastien Le Beux, Peng Yang 0003, Xiaowen Wu
DATE4
2015 Adaptively tolerate power-gating-induced power/ground noise under process variations
Zhe Wang 0003, Xuan Wang 0001, Jiang Xu 0001, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong, Haoran Li 0002, Rafael Kioji Vivas Maeda
DATE5
2015 An Analytical Study of Power Delivery Systems for Many-Core Processors Using On-Chip and Off-Chip Voltage Regulators
abstract
Design of power delivery system has great influence on the power management in many-core processor systems. Moving voltage regulators from off-chip to on-chip gains more and more interest in the power delivery system design, because it is able to provide fine-grained dynamic voltage scaling. Previous works are proposed to implement power efficient on-chip voltage regulators. It is important to analyze the characteristics of the entire power delivery system to explore the tradeoff between the promising properties and costs of employing on-chip voltage regulators, especially the on-chip buck converters. In this paper, we present a novel analysis and design optimization platform of power delivery system called power supply on-chip (PowerSoC). It employs an analytical model to provide an accurate and fast evaluation of important characteristics, e.g., power efficiency, output stability, and dynamic voltage scaling, for the entire power delivery system consisting of on-chip/off-chip buck converters and power delivery network. Based on our model, geometric programming is utilized to find the optimal design for different power delivery systems and explore the tradeoff of using on-chip converters. Compared with SPICE simulations, our model achieves a simulation time reduction of six to seven orders of magnitude within 5% model error for the characteristic evaluation of different power delivery systems. By using PowerSoC, various architectures of power delivery systems are optimized for power efficiency under constraints of output stability, area, etc. Simulation results show that the hybrid architecture, consisting of both on-chip and off-chip converters, achieves 1.0% power efficiency improvement and 66.4% area reduction of converters, compared to the conventional design. We conclude the hybrid architecture has potential for efficient dynamic voltage scaling, small area, and the adaptability of the change of power delivery network parasitic, but careful account for the overhead of on-chip converters is needed.
Xuan Wang 0001, Jiang Xu 0001, Zhe Wang 0003, Kevin J. Chen, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2015 Fat-Tree-Based Optical Interconnection Networks Under Crosstalk Noise Constraint
abstract
Optical networks-on-chip (ONoCs) have shown the potential to be substituted for electronic networks-on-chip (NoCs) to bring substantially higher bandwidth and more efficient power consumption in both on- and off-chip communication. However, basic optical devices, which are the key components in constructing ONoCs, experience inevitable crosstalk noise and power loss; the crosstalk noise from the basic devices accumulates in large-scale ONoCs and considerably hurts the signal-to-noise ratio (SNR) as well as restricts the network scalability. For the first time, this paper presents a formal system-level analytical approach to analyze the worst-case crosstalk noise and SNR in arbitrary fat-tree-based ONoCs. The analyses are performed hierarchically at the basic optical device level, then at the optical router level, and finally at the network level. A general 4$\,\times\,$4 optical router model is considered to enable the proposed method to be adaptable to fat-tree-based ONoCs using an arbitrary 4$\,\times\,$4 optical router. Utilizing the proposed general router model, the worst-case SNR link candidates in the network are determined. Moreover, we apply the proposed analyses to a case study of fat-tree-based ONoCs using an optical turnaround router (OTAR). Quantitative simulation results indicate low values of SNR and scalability constraints in large scale fat-tree-based ONoCs, which is due to the high power of crosstalk noise and power loss. For instance, in fat-tree-based ONoCs using the OTAR, when the injection laser power equals 0 dBm, the crosstalk noise power is higher than the signal power when the number of processor cores exceeds 128; when it is equal to 256, the signal power, crosstalk noise power, and SNR are${-}{17.3}$,${-}{11.9}$, and${-}{\rm 5.5}~{\rm dB}$, respectively.
Mahdi Nikdast, Jiang Xu 0001, Luan H. K. Duong, Xiaowen Wu, Zhehui Wang, Xuan Wang 0001, Zhe Wang 0003
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Crosstalk Noise in WDM-Based Optical Networks-on-Chip: A Formal Study and Comparison
abstract
Optical networks-on-chip (ONoCs) using wavelength-division multiplexing (WDM) technology have progressively attracted more and more attention for their use in tackling the high-power consumption and low bandwidth issues in growing metallic interconnection networks in multiprocessor systems-on-chip. However, the basic optical devices employed to construct WDM-based ONoCs are imperfect and suffer from inevitable power loss and crosstalk noise. Furthermore, when employing WDM, optical signals of various wavelengths can interfere with each other through different optical switching elements within the network, creating crosstalk noise. As a result, the crosstalk noise in large-scale WDM-based ONoCs accumulates and causes severe performance degradation, restricts the network scalability, and considerably attenuates the signal-to-noise ratio (SNR). In this paper, we systematically study and compare the worst case as well as the average crosstalk noise and SNR in three well-known optical interconnect architectures, mesh-based, folded-torus-based, and fat-tree-based ONoCs using WDM. The analytical models for the worst case and the average crosstalk noise and SNR in the different architectures are presented. Furthermore, the proposed analytical models are integrated into a newly developed crosstalk noise and loss analysis platform (CLAP) to analyze the crosstalk noise and SNR in WDM-based ONoCs of any network size using an arbitrary optical router. Utilizing CLAP, we compare the worst case as well as the average crosstalk noise and SNR in different WDM-based ONoC architectures. Furthermore, we indicate how the SNR changes in respect to variations in the number of optical wavelengths in use, the free-spectral range, and the microresonators$\boldsymbol {Q}$factor. The analyses’ results demonstrate that the crosstalk noise is of critical concern to WDM-based ONoCs: in the worst case, the crosstalk noise power exceeds the signal power in all three WDM-based ONoC architectures, even when the number of processor cores is small, e.g., 64.
Mahdi Nikdast, Jiang Xu 0001, Luan H. K. Duong, Xiaowen Wu, Xuan Wang 0001, Zhehui Wang, Zhe Wang 0003, Peng Yang 0003, Yaoyao Ye, Qinfen Hao
IEEE Trans. Very Large Scale Integr. Syst.6
2015 Actively Alleviate Power Gating-Induced Power/Ground Noise Using Parasitic Capacitance of On-Chip Memories in MPSoC
abstract
By integrating multiple processing units (PUs) and memories on a single chip, multiprocessor system-on-chip (MPSoC) can provide higher performance per energy and lower cost per function to applications with growing complexity. On the other hand, shrinking feature sizes and reducing power supply voltages also make MPSoCs more susceptible to various reliability threats, such as power/ground (P/G) noises. Power gating is an effective technique to minimize leakage power. However, it also introduces significant P/G noises in MPSoCs. With significant area, power and performance overheads, traditional methods rely on reinforced circuits or fixed protection strategies to reduce P/G noises caused by power gating. In this paper, we propose a systematic approach to actively alleviating P/G noises using the parasitic capacitance of on-chip memories through sensor network on-chip (SENoC). We use the parasitic capacitance of on-chip memories as dynamic decoupling capacitance to suppress P/G noises and develop a detailed HSPICE model for related study. SENoC is developed to not only monitor and report P/G noises, but also coordinate PUs and memories to alleviate such transient threats at run time. Extensive evaluations show that compared with traditional method, our approach saves 12.6%–62.8% energy consumption and achieves 14.3%–69.8% performance improvement for different applications and MPSoCs with different scales. We implement the circuit details of our approach and show its low area and energy consumption overheads.
Xuan Wang 0001, Jiang Xu 0001, Wei Zhang 0012, Xiaowen Wu, Yaoyao Ye, Zhehui Wang, Mahdi Nikdast, Zhe Wang 0003
IEEE Trans. Very Large Scale Integr. Syst.6
2015 An Inter/Intra-Chip Optical Network for Manycore Processors
abstract
Manycore processor system is becoming an attractive platform for applications seeking both high performance and high energy efficiency. However, huge communication demands among cores, large power density, and low process yield will be three significant limitations for the scalability of future manycore processors. Breaking a large chip into multiple smaller ones can alleviate the problems of power density and yield, but would worsen the problem of communication efficiency due to the limited off-chip bandwidth. In response, we propose an inter/intra-chip optical network, which will not only fulfill the intra-chip communication requirements but also address the inter-chip communication, by exploiting the advantages of optical links with high bandwidth and energy efficiency. The network is composed of an inter-chip subnetwork and multiple intra-chip subnetworks, and the subnetworks closely coordinate with each other to balance the traffic. The proposed network effectively explores the distinctive properties of optical signals and photonic devices, and dynamically partitions each data channel into multiple sections. Each section can be utilized independently to boost performance as well as reduce energy consumption. Simulation results show that our network can achieve higher throughput with lower power consumption than alternative designs under most of synthetic traffics and real applications.
Xiaowen Wu, Jiang Xu 0001, Yaoyao Ye, Xuan Wang 0001, Mahdi Nikdast, Zhehui Wang, Zhe Wang 0003
IEEE Trans. Very Large Scale Integr. Syst.6
2014 Characterizing power delivery systems with on/off-chip voltage regulators for many-core processors
abstract
Design of power delivery system has great influence on the power management in many-core processor systems. Moving voltage regulators from off-chip to on-chip gains more and more interest in the power delivery system design, because it is able to provide fast voltage scaling and multiple power domains. Previous works are proposed to implement power efficient on-chip regulators. It is also important to analyze the characteristics of the entire power delivery system to explore the tradeoff between the promising properties and costs of employing on-chip regulators. In this work, we develop an analytical model to evaluate important characteristics of the power delivery system, including on-chip/off-chip voltage regulators and the passive on-chip/on-board parasitic. Compared with SPICE simulations, our model achieves a fast system-level evaluation with comparable accuracy. Based on the model, geometric programming is utilized to find the optimal power efficiency of different architectures of power delivery systems under constraints of output voltage stability and area. Experiments show that compared with the conventional architecture using off-chip regulators, the hybrid one using both on-chip and off-chip voltage regulators achieves 1.0% power efficiency improvement and 68% area reduction of voltage regulators on average. We conclude that the hybrid architecture has potential for high power efficiency and small area at heavy workload, but careful account for the overhead of on-chip regulators is needed.
Xuan Wang 0001, Jiang Xu 0001, Zhe Wang 0003, Kevin J. Chen, Xiaowen Wu, Zhehui Wang
DATE6
2014 CLAP: a crosstalk and loss analysis platform for optical interconnects
abstract
Basic photonic devices in inter- and intra-chip optical networks suffer from inevitable power loss and crosstalk noise. Incoherent crosstalk introduces quick power fluctuations, while coherent crosstalk varies the optical power of the optical signal in optical interconnection networks (OINs). As a result, the accumulative crosstalk in large scale OINs considerably hurts the signal-to-noise ratio (SNR) and imposes high power penalties. In this work, we aim at studying the worst-case incoherent and coherent crosstalk in OINs at the system level. The proposed analytical models are integrated into a newly developed crosstalk and loss analysis platform, called CLAP, to facilitate the SNR analyses in arbitrary OINs.
Mahdi Nikdast, Luan H. K. Duong, Jiang Xu 0001, Sébastien Le Beux, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Yaoyao Ye
NOCS6
2014 On-chip sensor networks for soft-error tolerant real-time multiprocessor systems-on-chip
abstract
As transistor density continues to increase with the advent of nanotechnology, reliability issues raised by the more frequent appearance of soft errors are becoming critical for future embedded multiprocessor systems design. State-of-the-art techniques for soft error protections targeting multiprocessor systems result either high chip cost and area overhead or high performance degradation and energy consumption, and do not fulfill the increasing requirements for high performance and dependability. In this article we present a systematic approach, that is, the Sensor Networks-on-Chip (SENoC), to collaboratively and efficiently manage on-chip applications and overcome reliability threats to Multiprocessor Systems-on-Chip (MPSoC). A hardware-software collaborative approach is proposed to solve soft error problems: a hardware-based on-chip sensor network is built for soft error detection, and a software-based recovery mechanism is applied for soft error correction. A two-step scheduling scheme is presented for reliable application and chip management, combining an off-line static optimization stage for application performance maximization and an online lightweight dynamic adjustment stage to handle runtime variations and exceptions. This strategy introduces only trivial overhead on hardware design and much lower overhead on software control and execution, and hence performance degradation and energy consumption is greatly reduced. We build a cycle-accurate simulator using SystemC, and verify the effectiveness of our technique by comparing performance with related techniques on several real-world applications.
Weichen Liu 0001, Xuan Wang 0001, Jiang Xu 0001, Wei Zhang 0012, Yaoyao Ye, Xiaowen Wu, Mahdi Nikdast, Zhehui Wang
ACM J. Emerg. Technol. Comput. Syst.8
2014 SUOR: Sectioned Undirectional Optical Ring for Chip Multiprocessor
abstract
Chip multiprocessor (CMP) is becoming an attractive platform for applications seeking both high performance and high energy efficiency. In large-scale CMPs, the communication efficiency among cores is crucial for the overall system performance and energy consumption. In this article, we propose a ring-based optical network-on-chip, called SUOR, to fulfill the communication requirement of CMPs. SUOR effectively explores the distinctive properties of optical signals and photonic devices, and dynamically partitions each data channel into multiple sections. Each section can be utilized independently to boost performance as well as reduce energy consumption. We develop a set of distributed control protocols and algorithms for SUOR, but physically allocate the corresponding cluster agents close to each other to benefit from the strengths of optical interconnects at long distances as well as electrical interconnects at short distances. Simulation results show that SUOR outperforms the alternative optical networks under a wide range of traffic patterns. For example, compared with MWSR design, SUOR achieves 2.58× throughput as well as saves 64% energy consumption on average in a 256-core CMP. Compared with MWMR design, SUOR achieves 1.52× throughput and reduces 73% energy consumption on average.
Xiaowen Wu, Jiang Xu 0001, Yaoyao Ye, Zhehui Wang, Mahdi Nikdast, Xuan Wang 0001
ACM J. Emerg. Technol. Comput. Syst.4
2014 Floorplan Optimization of Fat-Tree-Based Networks-on-Chip for Chip Multiprocessors
abstract
Chip multiprocessor (CMP) is becoming increasingly popular in the processor industry. Efficient network-on-chip (NoC) that has similar performance to the processor cores is important in CMP design. Fat-tree-based on-chip network has many advantages over traditional mesh or torus-based networks in terms of throughput, power efficiency, and latency. It has a bright future in the development of CMP. However, the floorplan design of the fat-tree-based NoC is very challenging because of the complexity of topology. There are a large number of crossings and long interconnects, which cause severe performance degradation in the network. In electronic NoCs, the parasitic capacitance and inductance will be significant. In optical ones, large crosstalk noise and power loss will be introduced. The novel contribution of this paper is to propose a method to optimize the fat-tree floorplan, which can effectively reduce the number of crossings and minimize the interconnect length. Two types of floorplans are proposed, which could be applied to fat-tree-based networks of arbitrary size. Compared with the traditional one, our floorplans could reduce more than 87% of the crossings. Since the traversal distance for signals is related to the aspect ratio of the processor cores, we also present a method to calculate the optimum aspect ratio of the processor cores to minimize the traversal distance.
Zhehui Wang, Jiang Xu 0001, Xiaowen Wu, Yaoyao Ye, Wei Zhang 0012, Mahdi Nikdast, Xuan Wang 0001, Zhe Wang 0003
IEEE Trans. Computers1
2014 Systematic Analysis of Crosstalk Noise in Folded-Torus-Based Optical Networks-on-Chip
abstract
Photonic devices are widely used in optical networks-on-chip (ONoCs) and suffer from crosstalk noise. The accumulative crosstalk noise in large scale ONoCs diminishes the signal-to-noise ratio (SNR), causes severe performance degradation, and constrains the network scalability. For the first time, this paper systematically analyzes and models the worst-case crosstalk noise and SNR in folded-torus-based ONoCs. Formal analytical models for the worst-case crosstalk noise and SNR are presented. The crosstalk noise analysis is hierarchically performed at the basic photonic device level, then at the optical router level, and finally at the network level. We consider a general 5$\,\times\,$5 optical router model to enable crosstalk noise and SNR analyses in folded-torus-based ONoCs using an arbitrary 5$\,\times\,$5 optical router. Using the general optical router model, the worst-case SNR link candidates, which restrict the network scalability, are found. Also, we present a novel crosstalk noise and loss analysis platform, called CLAP, which can analyze the crosstalk noise and SNR of arbitrary ONoCs. Case studies of optimized crossbar and Crux optical routers using recent photonic device parameters are presented. Moreover, we compare the worst-case crosstalk noise and SNR in folded-torus-based and mesh-based ONoCs using optimized crossbar and Crux optical routers. The quantitative simulation results show the critical behavior of crosstalk noise in large scale ONoCs. For example, in folded-torus-based ONoCs using the Crux optical router, the noise power exceeds the signal power for network sizes larger than 12$\,\times\,$12; when the network size is 20$\,\times\,$20 and the injection signal power equals 0 dBm, the signal power and noise power are${-}{\rm 9.4}~{\rm dBm}$and${-}{\rm 6.1}~{\rm dBm}$, respectively.
Mahdi Nikdast, Jiang Xu 0001, Xiaowen Wu, Wei Zhang 0012, Yaoyao Ye, Xuan Wang 0001, Zhehui Wang, Zhe Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2014 System-Level Modeling and Analysis of Thermal Effects in WDM-Based Optical Networks-on-Chip
abstract
Multiprocessor systems-on-chip show a trend toward integration of tens and hundreds of processor cores on a single chip. With the development of silicon photonics for short-haul optical communication, wavelength division multiplexing (WDM)-based optical networks-on-chip (ONoCs) are emerging on-chip communication architectures that can potentially offer high bandwidth and power efficiency. Thermal sensitivity of photonic devices is one of the main concerns about the on-chip optical interconnects. We systematically modeled thermal effects in optical links in WDM-based ONoCs. Based on the proposed thermal models, we developed OTemp, an optical thermal effect modeling platform for optical links in both WDM-based ONoCs and single-wavelength ONoCs. OTemp can be used to simulate the power consumption as well as optical power loss for optical links under temperature variations. We use case studies to quantitatively analyze the worst-case power consumption for one wavelength in an eight-wavelength WDM-based optical link under different configurations of low-temperature-dependence techniques. Results show that the worst-case power consumption increases dramatically with on-chip temperature variations. Thermal-based adjustment and optimal device settings can help reduce power consumption under temperature variations. Assume that off-chip vertical-cavity surface-emitting lasers are used as the laser source with WDM channel spacing of 1 nm, if we use thermal-based adjustment with guard rings for channel remapping, the worst-case total power consumption is 6.7 pJ/bit under the maximum temperature variation of 60 °C; larger channel spacing would result in a larger worst-case power consumption in this case. If we use thermal-based adjustment without channel remapping, the worst-case total power consumption is around 9.8 pJ/bit under the maximum temperature variation of 60 °C; in this case, the worst-case power consumption would benefit from a larger channel spacing.
Yaoyao Ye, Zhehui Wang, Peng Yang 0003, Jiang Xu 0001, Xiaowen Wu, Xuan Wang 0001, Mahdi Nikdast, Zhe Wang 0003, Luan H. K. Duong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 Active power-gating-induced power/ground noise alleviation using parasitic capacitance of on-chip memories
abstract
By integrating multiple processing units and memories on a single chip, multiprocessor system-on-chip (MPSoC) can provide higher performance per energy and lower cost per function to applications with growing complexity. In order to maintain the power budget, power gating technique is widely used to reduce the leakage power. However, it will introduce significant power/ground (P/G) noises, and threat the reliability of MPSoCs. With significant area, power and performance overheads, traditional methods rely on reinforced circuits or fixed protection strategies to reduce P/G noises caused by power gating. In this paper, we propose a systematic approach to actively alleviating P/G noises using the parasitic capacitance of on-chip memories through sensor network on-chip (SENoC). We utilize the parasitic capacitance of on-chip memories as dynamic decoupling capacitance to suppress P/G noises and develop a detailed Hspice model for related study. SENoC is developed to not only monitor and report P/G noises but also coordinate processing units and memories to alleviate such transient threats at run time. Extensive evaluations show that compared with traditional methods, our approach saves 11.7% to 62.2% energy consumption and achieves 13.3% to 69.3% performance improvement for different applications and MPSoCs with different scales. We implement the circuit details of our approach and show its low area and energy consumption overheads.
Xuan Wang 0001, Jiang Xu 0001, Wei Zhang 0012, Xiaowen Wu, Yaoyao Ye, Zhehui Wang, Mahdi Nikdast, Zhe Wang 0003
DATE6
2013 System-level analysis of mesh-based hybrid optical-electronic network-on-chip
abstract
Network-on-chip (NoC) can improve the performance, power efficiency, and scalability of multiprocessor system-on-chip (MPSoC). Optical NoCs, which are based on CMOS-compatible optical waveguides and microresonators, have significant bandwidth and power advantages over metallic interconnects. We propose a low-cost mesh-based hybrid optical-electronic NoC, HOME, with non-blocking 5×5, 4×4 and 3×3 optical switching fabrics. We systematically analyzed the key characteristics of HOME for a 64-core MPSoC in 45nm under different traffic conditions. Besides, we quantitatively analyzed the thermal effects in the 64-core HOME under temperature variations.
Yaoyao Ye, Xiaowen Wu, Jiang Xu 0001, Mahdi Nikdast, Zhehui Wang, Xuan Wang 0001, Zhe Wang 0003
ISCAS5
2013 3-D Mesh-Based Optical Network-on-Chip for Multiprocessor System-on-Chip
abstract
Optical networks-on-chip (ONoCs) are emerging communication architectures that can potentially offer ultrahigh communication bandwidth and low latency to multiprocessor systems-on-chip (MPSoCs). In addition to ONoC architectures, 3-D integrated technologies offer an opportunity to continue performance improvements with higher integration densities. In this paper, we present a 3-D mesh-based ONoC for MPSoCs, and new low-cost nonblocking 4$\,\times\,$4, 5$\,\times\,$5, 6$\,\times\,$6, and 7$\,\times\,$7 optical routers for dimension-order routing in the 3-D mesh-based ONoC. Besides, we propose an optimized floorplan for the 3-D mesh-based ONoC. The floorplan follows the regular 3-D mesh topology but implements all optical routers in a single optical layer. The floorplan is optimized to minimize the number of extra waveguide crossings caused when merging the 3-D ONoC to one optical layer. Based on a set of real applications and uniform traffic pattern, we develop a SystemC-based cycle-accurate NoC simulator and compare the 3-D mesh-based ONoC with the matched 2-D mesh-based ONoC and 2-D electronic NoC for performance and energy efficiency. Additionally, we quantitatively analyze thermal effects on the 3-D 8$\,\times\,$8$\,\times\,$2 mesh-based ONoC.
Yaoyao Ye, Jiang Xu 0001, Baihan Huang, Xiaowen Wu, Wei Zhang 0012, Xuan Wang 0001, Mahdi Nikdast, Zhehui Wang, Weichen Liu 0001, Zhe Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2013 Formal Worst-Case Analysis of Crosstalk Noise in Mesh-Based Optical Networks-on-Chip
abstract
Crosstalk noise is an intrinsic characteristic as well as a potential issue of photonic devices. In large scale optical networks-on-chips (ONoCs), crosstalk noise could cause severe performance degradation and prevent ONoC from communicating properly. The novel contribution of this paper is the systematical modeling and analysis of the crosstalk noise and the signal-to-noise ratio (SNR) of optical routers and mesh-based ONoCs using a formal method. Formal analytical models for the worst-case crosstalk noise and minimum SNR in mesh-based ONoCs are presented. The crosstalk analysis is performed at device, router, and network levels. A general 5$\,\times\,$5 optical router model is proposed for router level analysis. The minimum SNR optical link candidates, which constrain the scalability of mesh-based ONoCs, are identified. It is also shown that symmetric mesh-based ONoCs have the best SNR performance. The presented formal analyses can be easily applied to other optical routers and mesh-based ONoCs. Finally, we present case studies of mesh-based ONoCs using the optimized crossbar and Crux optical routers to evaluate the proposed formal method. We find that crosstalk noise can significantly limit the scalability of mesh-based ONoCs. For example, when the mesh-based ONoC size, using optimized crossbar, is larger than 8$\,\times\,$8, the optical signal power is smaller than the crosstalk noise power; when the network size is 16$\,\times\,$16 and the input power is 0 dBm, in the worst-case, the signal power is${-}{\rm 24.9}~{\rm dBm}$and the crosstalk noise power is${-}{\rm 11}~{\rm dBm}$.
Yiyuan Xie, Mahdi Nikdast, Jiang Xu 0001, Xiaowen Wu, Wei Zhang 0012, Yaoyao Ye, Xuan Wang 0001, Zhehui Wang, Weichen Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.8
2013 System-Level Modeling and Analysis of Thermal Effects in Optical Networks-on-Chip
abstract
The performance of multiprocessor systems, such as chip multiprocessors (CMPs), is determined not only by individual processor performance, but also by how efficiently the processors collaborate with one another. It is the communication architecture that determines the collaboration efficiency on the hardware side. Optical networks-on-chip (ONoCs) are emerging communication architectures that can potentially offer ultra-high communication bandwidth and low latency to multiprocessor systems. Thermal sensitivity is an intrinsic characteristic of photonic devices used by ONoCs as well as a potential issue. This paper systematically modeled and quantitatively analyzed the thermal effects in ONoCs. We used an 8$\times$8 mesh-based ONoC as a case study and evaluated the impacts of thermal effects in the average power efficiency for real MPSoC applications. We revealed three important factors regarding ONoC power efficiency under temperature variations, and proposed several techniques to reduce the temperature sensitivity of ONoCs. These techniques include the optimal initial setting of microresonator resonant wavelength, increasing the 3-dB bandwidth of optical switching elements by parallel coupling multiple microresonators, and the use of passive-routing optical router Crux to minimize the number of switching stages in mesh-based ONoCs. We gave a mathematical analysis of periodically parallel coupling of multiple microresonators and show that the 3-dB bandwidth of optical switching elements can be widened nearly linearly with the ring number. Evaluation results for different real MPSoC applications show that, on the basis of thermal tuning, the optimal device setting improves the average power efficiency by 54% to 1.2 pJ/bit when chip temperature reaches 85$^{\circ}$C. The findings in this paper can help support the further development of this emerging technology.
Yaoyao Ye, Jiang Xu 0001, Xiaowen Wu, Wei Zhang 0012, Xuan Wang 0001, Mahdi Nikdast, Zhehui Wang, Weichen Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.7