Lei Yang 0018

dblp:50/2484-18 · DBLP profile ↗
← Back
59ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0002-0646-440XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 50 · 9 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Computational Performance Bounds Prediction in Quantum Computing With Unstable Noise
abstract
Quantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. In this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach.
Jinyang Li 0001, Samudra Dasgupta, Yuhong Song, Lei Yang 0018, Travis S. Humble, Weiwen Jiang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 DiGiT: A Diffusion-based Modular Geophysical Toolkit for On-device Multi-modal Data Generation
abstract
Full-wave inversion (FWI), as a fundamental scientific approach to deducing unknown or unobservable subsurface properties, holds significant value in geophysics applications. Traditional FWI methods rely on physics-driven approaches that demand substantial computational resources. Recently, with the breakthroughs in machine learning (ML) and the prevalence of AI for science, data-driven approaches have been applied to FWI, showing promising results. However, as these applications often necessitate deployment in diverse regions with remote and extreme environments, localization of ML models on edge devices becomes imperative. A promising approach involves leveraging Generative AI models and governing wave equations to generate paired training data, including geophysical measurements (i.e., seismic waveform) as data and corresponding velocity maps as labels for model fine-tuning. However, the limited resources on edge devices pose significant challenges to achieving high software efficiency and low latency. In this article, we present a toolkit, namely DiGiT, a di ffusion-based modular g eophys i cal t oolkit platform. One key component is a library of decomposed modules from the widely used geophysical designs. Benefiting from the flexibility of combining modules, we composite a toolkit for the generation of on-device diffusion-based paired geophysical training data. The toolkit includes a 1-in-2-out network structure and diffusion model distillation, both of which can significantly reduce the computational time. Experiments on the OpenFWI dataset show that the DiGiT toolkit can generate paired seismic waveform and velocity map in seconds, which is over 100× speedup compared with the sequential execution of the diffusion model and the wave equation-based forward modeling.
Junhuan Yang, Yi Sheng 0001, Youzuo Lin, Weiwen Jiang, Lei Yang 0018
ACM Trans. Embed. Comput. Syst.6
2025 A Novel Diffusion Model for Pairwise Geoscience Data Generation with Unbalanced Training Dataset
abstract
Recently, the advent of generative AI technologies has made transformational impacts on our daily lives, yet its application in scientific applications remains in its early stages. Data scarcity is a major, well-known barrier in data-driven scientific computing, so physics-guided generative AI holds significant promise. In scientific computing, most tasks study the conversion of multiple data modalities to describe physical phenomena, for example, spatial and waveform in seismic imaging, time and frequency in signal processing, and temporal and spectral in climate modeling; as such, multi-modal pairwise data generation is highly required instead of single-modal data generation, which is usually used in natural images (e.g., faces, scenery). Moreover, in real-world applications, the unbalance of available data in terms of modalities commonly exists; for example, the spatial data (i.e., velocity maps) in seismic imaging can be easily simulated, but real-world seismic waveform is largely lacking. While the most recent efforts enable the powerful diffusion model to generate multi-modal data, how to leverage the unbalanced available data is still unclear. In this work, we use seismic imaging in subsurface geophysics as a vehicle to present "UB-Diff", a novel diffusion model for multi-modal paired scientific data generation. One major innovation is a one-in-two-out encoder-decoder network structure, which can ensure pairwise data is obtained from a co-latent representation. Then, the co-latent representation will be used by the diffusion process for pairwise data generation. Experimental results on the OpenFWI dataset show that UB-Diff significantly outperforms existing techniques in terms of Fréchet Inception Distance (FID) score and pairwise evaluation, indicating the generation of reliable and useful multi-modal pairwise data.
Junhuan Yang, Yi Sheng 0001, Youzuo Lin, Lei Yang 0018
AAAI5
2024 Tutorial on Novel Toolkits toward AI for Science on Resource-Constrained Computing Systems
abstract
Full Waveform Inversion (FWI) is a technique used to visualize and analyze wave propagation through a medium in order to infer its physical properties. This method relies on computational models and algorithms to simulate and interpret the behavior of waves—such as sound, electromagnetic, or seismic waves—as they travel through different materials. By analyzing how these waves are reflected, refracted, or absorbed by the medium, FWI can provide detailed information about the medium’s internal structure, composition, and physical properties, such as density, elasticity, or internal defects. The traditional process typically involves: 1) Wave Simulation: Using physics-based models to simulate how waves propagate through a medium. This may involve solving complex differential equations that describe wave behavior in different contexts. 2) Data Acquisition: Collecting data on wave interactions with the medium using sensors or other measurement devices. This could include data on wave speed, direction, amplitude, and phase changes. 3) Image Reconstruction: Applying computational techniques, such as inverse problems or tomographic reconstruction, to create images or maps of the medium based on the acquired wave data. 4) Analysis: Interpreting the reconstructed images to deduce the physical properties of the medium. This can involve identifying features like boundaries, interfaces, or anomalies within the medium.
Yi Sheng 0001, Junhuan Yang, Hanchen Wang 0003, Yinan Feng, Yinpeng Chen, Youzuo Lin, Weiwen Jiang, Lei Yang 0018
CODES+ISSS9
2024 EdGeo: A Physics-guided Generative AI Toolkit for Geophysical Monitoring on Edge Devices
abstract
Full-waveform inversion (FWI) plays a vital role in geoscience to explore the subsurface. It utilizes the seismic wave to image the subsurface velocity map. As the machine learning (ML) technique evolves, the data-driven approaches using ML for FWI tasks have emerged, offering enhanced accuracy and reduced computational cost compared to traditional physics-based methods. However, a common challenge in geoscience --- the unprivileged data --- severely limits ML effectiveness. The issue becomes even worse during model pruning, a step essential in geoscience due to environmental complexities. To tackle this, we introduce the EdGeo toolkit, which employs a diffusion-based model guided by physics principles to generate high-fidelity velocity maps. The toolkit uses the acoustic wave equation to generate corresponding seismic waveform data, facilitating the fine-tuning of pruned ML models. Our results demonstrate significant improvements in SSIM scores and reduction in both MAE and MSE across various pruning ratios. Notably, the ML model fine-tuned using data generated by EdGeo yields superior quality of velocity maps, especially in representing unprivileged features, outperforming other existing methods.
Junhuan Yang, Hanchen Wang 0003, Yi Sheng 0001, Youzuo Lin, Lei Yang 0018
DAC5
2024 Toward Fair Ultrasound Computing Tomography: Challenges, Solutions and Outlook
abstract
Medical image reconstruction plays a pivotal role in early cancer detection, which can significantly enhance both the quality and longevity of a patient’s life through timely treatment. However, the extent to which current image reconstruction methods accurately represent all populations, and whether they underperform for certain groups, remains largely unexplored. In this work, we will examine the deep learning (DL)–based approach to image reconstruction and its associated fairness concerns. Initially, our experiments confirmed the unfairness’s presence. Subsequently, by addressing the issue from two perspectives, we gained valuable insights, which deepened our understanding of the problem. To assess a model’s fairness, it’s crucial to evaluate it from various perspectives, as relying on a single metric can often yield misleading results.
Yi Sheng 0001, Junhuan Yang, Youzuo Lin, Weiwen Jiang, Lei Yang 0018
ACM Great Lakes Symposium on VLSI5
2024 Enhanced AI for Science using Diffusion-based Generative AI - A Case Study on Ultrasound Computing Tomography
abstract
Ultrasound computed tomography (USCT) is an emerging imaging modality that holds great promise for breast imaging. Full-waveform inversion (FWI)-based image reconstruction methods leverage accurate wave physics to generate high spatial resolution quantitative images of the breast tissue’s acoustic properties, such as speed of sound, from USCT measurement data. However, the significant computational demand for FWI reconstruction poses a considerable challenge to its widespread adoption in clinical settings. Data-driven machine learning approaches offer a faster and more efficient means of translating waveform data into images. Yet, the effectiveness of machine learning methods is constrained by the diversity and quality of the training data. Given the heterogeneous distribution of breast tissue characteristics, such as fat content and size, the performance of machine learning varies across different sizes. This variability is problematic, particularly in medical diagnostics, where precision is crucial. In response to the limited data in certain categories, we propose utilizing generative AI to augment data samples, thereby enhancing FWI’s performance on limited-sample data and addressing issues of AI fairness.
Junhuan Yang, Yi Sheng 0001, Hanchen Wang 0003, Youzuo Lin, Lei Yang 0018
ACM Great Lakes Symposium on VLSI6
2024 APS-USCT: Ultrasound Computed Tomography on Sparse Data via AI-Physic Synergy
Yi Sheng 0001, Hanchen Wang 0003, Yipei Liu, Junhuan Yang, Weiwen Jiang, Youzuo Lin, Lei Yang 0018
MICCAI (7)7
2024 Data-Algorithm-Architecture Co-Optimization for Fair Neural Networks on Skin Lesion Dataset
Junhuan Yang, James Alaina, Xiaowei Xu 0004, Yiyu Shi 0001, Jingtong Hu, Weiwen Jiang, Lei Yang 0018
MICCAI (10)9
2024 Self-guided Knowledge-Injected Graph Neural Network for Alzheimer's Diseases
Zhepeng Wang 0001, Runxue Bao, Yawen Wu, Lei Yang 0018, Liang Zhan, Feng Zheng 0001, Weiwen Jiang, Yanfu Zhang
MICCAI (2)5
2024 Advancements in Accelerating Deep Neural Network Inference on AIoT Devices: A Survey
abstract
The amalgamation of artificial intelligence with Internet of Things (AIoT) devices have seen a rapid surge in growth, largely due to the effective implementation of deep neural network (DNN) models across various domains. However, the deployment of DNNs on such devices comes with its own set of challenges, primarily related to computational capacity, storage, and energy efficiency. This survey offers an exhaustive review of techniques designed to accelerate DNN inference on AIoT devices, addressing these challenges head-on. We delve into critical model compression techniques designed to adapt to the limitations of devices and hardware optimization strategies that aim to boost efficiency. Furthermore, we examine parallelization methods that leverage parallel computing for swift inference, as well as novel optimization strategies that fine-tune the execution process. This survey also casts a future-forward glance at emerging trends, including advancements in mobile hardware, the co-design of software and hardware, privacy and security considerations, and DNN inference on AIoT devices with constrained resources. All in all, this survey aspires to serve as a holistic guide to advancements in the acceleration of DNN inference on AIoT devices, aiming to provide sustainable computing for upcoming IoT applications driven by artificial intelligence.
Long Cheng 0003, Qingzhi Liu, Lei Yang 0018, Cheng Liu 0008, Ying Wang 0001
IEEE Trans. Sustain. Comput.4
2023 Toward Fair and Efficient Hyperdimensional Computing
abstract
We are witnessing the evolution that Machine Learning (ML) is applied to varied applications, such as intelligent security systems, medical diagnoses, etc. With this trend, it has high demand to run ML on end devices with limited resources. What's more, the fairness in these ML algorithms is mounting important, since these applications are not designed for specific users (e.g., people with fair skin in skin disease diagnosis) but need to be applied to all possible users (i.e., people with different skin tones). Brain-inspired hyperdimensional computing (HDC) has demonstrated its ability to run ML tasks on edge devices with a small memory footprint; yet, it is unknown whether HDC can satisfy the fairness requirements from applications (e.g., medical diagnosis for people with different skin tones). In this paper, for the first time, we reveal that the vanilla HDC has severe bias due to its sensitivity to color information. Toward a fair and efficient HDC, we propose a holistic framework, namely FE-HDC, which integrates the image processing and input compression techniques in HDC's encoder. Compared with the vanilla HDC, results show that the proposed FE-HDC can reduce the unfairness score by 90%, achieving fairer architectures with competitively high accuracy.
Yi Sheng 0001, Junhuan Yang, Weiwen Jiang, Lei Yang 0018
ASP-DAC4
2023 Muffin: A Framework Toward Multi-Dimension AI Fairness by Uniting Off-the-Shelf Models
abstract
Model fairness (a.k.a., bias) has become one of the most critical problems in a wide range of AI applications. An unfair model in autonomous driving may cause a traffic accident if corner cases (e.g., extreme weather) cannot be fairly regarded; or it will incur healthcare disparities if the AI model misdiagnoses a certain group of people (e.g., brown and black skin). In recent years, there are emerging research works on addressing unfairness, and they mainly focus on a single unfair attribute, like skin tone; however, real-world data commonly have multiple attributes, among which unfairness can exist in more than one attribute, called "multi-dimensional fairness". In this paper, we first reveal a strong correlation between the different unfair attributes, i.e., optimizing fairness on one attribute will lead to the collapse of others. Then, we propose a novel Multi-Dimension Fairness framework, namely Muffin, which includes an automatic tool to unite off-the-shelf models to improve the fairness on multiple attributes simultaneously. Case studies on dermatology datasets with two unfair attributes show that the existing approach can achieve 21.05% fairness improvement on the first attribute while it makes the second attribute unfair by 1.85%. On the other hand, the proposed Muffin can unite multiple models to achieve simultaneously 26.32% and 20.37% fairness improvement on both attributes; meanwhile, it obtains 5.58% accuracy gain.
Yi Sheng 0001, Junhuan Yang, Lei Yang 0018, Yiyu Shi 0001, Jingtong Hu, Weiwen Jiang
DAC3
2023 Late Breaking Results: Fast Fair Medical Applications? Hybrid Vision Models Achieve the Fairness on the Edge
abstract
As edge devices become readily available and indispensable, there is an urgent need for effective and efficient intelligent applications to be deployed widespread. However, fairness has always been an issue, especially in edge medical applications. Compared to convolutional neuron networks (CNNs), Vision Transformer (ViT) has a better ability to extract global information, which will contribute to alleviating the unfairness problem. Typically, ViTs consume large amounts of computational and memory resources, which hinders their usage on edge. In this work, we propose a novel hardware-efficient Vision Model search framework for the fair dermatology classification, namely HeViFa. Experimental results show that HeViFa could search for a hybrid ViT model that reaches 173.1 FPS on a Samsung S21 mobile phone with 85.71% accuracy on the light skin dataset and 80.85% accuracy on the dark skin dataset. Note that HeViFa can reach both the highest accuracy and fairness under similar latency constrain on multiple edge devices (Samsung S21 mobile phone, iPhone 13 Pro and Raspberry PI).
Changdi Yang, Yi Sheng 0001, Peiyan Dong, Zhenglun Kong, Yanyu Li, Pinrui Yu, Lei Yang 0018, Xue Lin 0001
DAC7
2023 On-Device Unsupervised Image Segmentation
abstract
Along with the breakthrough of convolutional neural networks, in particular encoder-decoder and U-Net, learning-based segmentation has emerged in many research works. Most of them are based on supervised learning, requiring plenty of annotated data; however, to support segmentation, a label for each pixel is required, which is obviously expensive. As a result, the issue of lacking annotated segmentation data commonly exists. Continuous learning is a promising way to deal with this issue; however, it still has high demands on human labor for annotation. What’s more, privacy is highly required in segmentation data for real-world applications, which further calls for on-device learning. In this paper, we aim to resolve the above issue in an alternative way: Instead of supervised segmentation, we propose to develop efficient unsupervised segmentation which can be executed on edge devices without annotated data. Based on our observation that segmentation can obtain high performance when pixels are mapped to a high-dimension space using their position and color information, we for the first time bring brain-inspired hyperdimensional computing (HDC) to the segmentation task. We build the HDC-based unsupervised segmentation framework, namely "SegHDC". In SegHDC, we devise a novel encoding approach, which follows the Manhattan distance. A clustering algorithm is further developed on top of the encoded high-dimension vectors to obtain segmentation results. Experimental results show that SegHDC can significantly surpass neural network-based unsupervised segmentation. On a standard segmentation dataset, DSB2018, SegHDC can achieve a 28.0% improvement in Intersection over Union (IoU) score; meanwhile, it achieves over 300× speedup on Raspberry PI. What’s more, for a larger size image in the BBBC005 dataset, the existing approach cannot be accommodated to Raspberry PI due to out of memory; on the other hand, SegHDC can obtain segmentation results within 3 minutes while achieving a 0.9587 IoU score.
Junhuan Yang, Yi Sheng 0001, Weiwen Jiang, Lei Yang 0018
DAC5
2023 Fast and Fair Medical AI on the Edge Through Neural Architecture Search for Hybrid Vision Models
abstract
As edge devices become readily available and indispensable, there is an urgent need for effective and efficient intelligent applications to be deployed widespread. However, fairness has always been an issue, especially in edge medical applications. Although many approaches have been proposed to mitigate the unfairness problem, their edge performance is not desirable. By examining the fairness performance of different network architectures, we observed that compared to pure convolutional neuron network (CNN) architecture, hybrid models with CNN and Vision Transformer (ViT) have exhibited better performance in terms of fairness and accuracy. After further analyzing the feature maps of intermediate layers of CNNs, ViTs, and hybrid models, we found that ViT has a strong ability to extract global information, which contributes to alleviating the unfairness problem. However, ViTs consume large amounts of computational and memory resources, which hinders their application on edge devices. To address the challenges abovementioned, we propose the first hardware-oriented co-design NAS framework to explore hybrid ViT-CNN architecture for the fair dermatology classification, namely HeViFa, which can produce light-weight models for edge devices with low unfairness scores and high classification accuracy. Experimental results show that compared with FaHaNa-Small, HeViFa-Small could search for a hybrid ViT model that reaches 10.57% and 4.03% higher accuracy as well as 0.179 and 0.0403 higher PQD score on Mix and Fitzpatrick17k dataset, repectively, and speed up by 1.21 × on Samsung S21 mobile phone, 1.18 × on iPhone 13 Pro and 1.37 × on Raspberry Pi.
Changdi Yang, Yi Sheng 0001, Peiyan Dong, Zhenglun Kong, Yanyu Li, Pinrui Yu, Lei Yang 0018, Xue Lin 0001, Yanzhi Wang 0001
ICCAD7
2023 On Hyperdimensional Computing-based Federated Learning: A Case Study
abstract
Federated learning is a decentralized machine learning strategy that trains the model by using data stored across multiple decentralized edge devices or servers. Studies on federated learning currently focus primarily on neural network-based learning methods, which usually require powerful hardware and are relatively not energy-efficient. Recently, hyperdimensional computing (HDC) emerges as a potential alternative solution to neural networks, particularly on resource-constrained platforms such as edge intelligence systems. HDC mimics the “human brain” at the functionality level that learns with the attributes of brain circuits, including high-dimensionality and fully distributed holographic representation. Although there are existing works related to HDC-based federated learning, a comprehensive study on how HDC-based federated learning performs in different settings is still absent. To bridge this gap, we present a comprehensive case study on federated learning using HDC under two model aggregation strategies: hypervector aggregation and associative memory aggregation. We also perform extensive experiments with various settings, including data distribution, number of clients, and local training epochs. We also analyze their communication costs under these settings. Our results show that using the strategy of associative memory aggregation can achieve up to 95% communication cost reduction compared to hypervector aggregation. In addition, HDC-based federated learning system shows high robustness in training with Non-IID data. This study aims to shed light and provide guidance in opening up new directions and challenges for future HDC-based federated learning system design and optimization.
Sizhe Zhang, Dongning Ma, Song Bian 0001, Lei Yang 0018, Xun Jiao 0002
IJCNN4
2023 Hardware-aware neural architecture search for stochastic computing-based neural networks on tiny devices
Yuhong Song, Edwin H.-M. Sha, Qingfeng Zhuge, Rui Xu 0013, Xiaowei Xu 0004, Bingzhe Li, Lei Yang 0018
J. Syst. Archit.7
2022 BSC: Block-based Stochastic Computing to Enable Accurate and Efficient TinyML
abstract
Along with the progress of AI democratization, machine learning (ML) has been successfully applied to edge applications, such as smart phones and automated driving. Nowadays, more applications require ML on tiny devices with extremely limited resources, like implantable cardioverter de-fibrillator (ICD), which is known as TinyML. Unlike ML on the edge, TinyML with a limited energy supply has higher demands on low-power execution. Stochastic computing (SC) using bitstreams for data representation is promising for TinyML since it can perform the fundamental ML operations using simple logical gates, instead of the complicated binary adder and multiplier. However, SC commonly suffers from low accuracy for ML tasks due to low data precision and inaccuracy of arithmetic units. Increasing the length of the bitstream in the existing works can mitigate the precision issue but incur higher latency. In this work, we propose a novel SC architecture, namely Block-based Stochastic Computing (BSC). BSC divides inputs into blocks, such that the latency can be reduced by exploiting high data parallelism. Moreover, optimized arithmetic units and output revision (OUR) scheme are proposed to improve accuracy. On top of it, a global optimization approach is devised to determine the number of blocks, which can make a better latency-power trade-off. Experimental results show that BSC can outperform the existing designs in achieving over 10% higher accuracy on ML tasks and over$6\times$power reduction.
Yuhong Song, Edwin H.-M. Sha, Qingfeng Zhuge, Rui Xu 0013, Yongzhuo Zhang, Bingzhe Li, Lei Yang 0018
ASP-DAC7
2022 The larger the fairer?: small neural networks can achieve fairness for edge devices
abstract
Along with the progress of AI democratization, neural networks are being deployed more frequently in edge devices for a wide range of applications. Fairness concerns gradually emerge in many applications, such as face recognition and mobile medical. One fundamental question arises: what will be the fairest neural architecture for edge devices? By examining the existing neural networks, we observe that larger networks typically are fairer. But, edge devices call for smaller neural architectures to meet hardware specifications. To address this challenge, this work proposes a novel Fairness- and Hardware-aware Neural architecture search framework, namely FaHaNa. Coupled with a model freezing approach, FaHaNa can efficiently search for neural networks with balanced fairness and accuracy, while guaranteed to meet hardware specifications. Results show that FaHaNa can identify a series of neural networks with higher fairness and accuracy on a dermatology dataset. Target edge devices, FaHaNa finds a neural architecture with slightly higher accuracy, 5.28X smaller size, 15.14% higher fairness score, compared with MobileNetV2; meanwhile, on Raspberry PI and Odroid XU-4, it achieves 5.75X and 5.79X speedup.
Yi Sheng 0001, Junhuan Yang, Yawen Wu, Kevin Mao, Yiyu Shi 0001, Jingtong Hu, Weiwen Jiang, Lei Yang 0018
DAC8
2022 On the Design of Quantum Graph Convolutional Neural Network in the NISQ-Era and Beyond
abstract
The rapid growth in the size of Graph Convolutional Neural Networks (GCNs) encounters both computational- and memory-wall on classical computing platforms (e.g., CPU, GPU, FPGA, etc.). Quantum computing, on the other hand, provides extremely high parallelism for computation. Although quantum neural networks have been recently studied, the research on quantum graph neural networks is still in its infancy. The key challenge here is how to integrate both the graph topology information and the learning ability of GCNs into quantum circuits. In this work, we leverage the Givens rotations and its quantum implementation to encode graph information; in addition, we employ the widely used variational quantum circuit to bring the learnable parameters. On top of these, we present a full-quantum design of Graph Convolutional Neural Networks, namely "QuGCN", for semi-supervised learning on graph-structured data. Experiment results show our design is competitive with classical GCNs in terms of node classification accuracy on Cora sub-dataset. More importantly, we show the potential advantages that can be achieved by the proposed quantum GCN design when the number of features grows.
Zhirui Hu, Jinyang Li 0001, Zhenyu Pan, Shanglin Zhou, Lei Yang 0018, Caiwen Ding, Omer Khan, Tong Geng, Weiwen Jiang
ICCD5
2022 RT-DNAS: Real-Time Constrained Differentiable Neural Architecture Search for 3D Cardiac Cine MRI Segmentation
Qing Lu 0001, Xiaowei Xu 0004, Shunjie Dong, Cong Hao, Lei Yang 0018, Cheng Zhuo, Yiyu Shi 0001
MICCAI (5)5
2021 HMC-TRAN: A Tensor-core Inspired Hierarchical Model Compression for Transformer-based DNNs on GPU
abstract
Although Transformer-based deep learning models have been widely used in many natural language processing (NLP) tasks as well as computer vision, they suffer from gigantic model size and long latency. Network pruning can reduce the computational cost and model size. However, existing works mainly focus on irregular(sparse) pruning, which often causes irregular computations and extra indices per remained weight. In this work, we propose a Tensor-core inspired hierarchical model compression method to push the performance limit on modern GPUs. We present two modes of the two-step process. In the first mode, we use the Tensor-core aware block-based weight pruning method to exploit model sparsity in a coarse-grained manner and then use low-rank [33] decomposition to further reduce the weight storage in a fine-grained manner.In the second mode, we first use irregular pruning to achieve a highly sparse model and then apply the Tensor-core aware weight constraint on the sparse model to decompose the sparse matrix to several smaller but Tensor-core friendly sub-matrices. Experiments on Transformer, BERTBASE models show the proposed method outperforms the state-of-the-art.
Shaoyi Huang, Shiyang Chen 0004, Hongwu Peng, Daniel Manu, Zhenglun Kong, Geng Yuan, Lei Yang 0018, Shusen Wang, Hang Liu 0001, Caiwen Ding
ACM Great Lakes Symposium on VLSI7
2021 Co-Exploration of Graph Neural Network and Network-on-Chip Design Using AutoML
abstract
Recently, Graph Neural Networks (GNNs) have exhibited high efficiency in several graph-based machine learning tasks. Compared with the neural networks for computer vision or speech tasks (e.g., Convolutional Neural Networks), GNNs have much higher requirements on communication due to the complicated graph structures; however, when applying GNNs for real-world applications, say in recommender systems (e.g. Uber Eats), it commonly has the real-time requirements. To deal with the tradeoff between the complicated architecture and the high-demand timing performance, both GNN architecture and hardware accelerator need to be optimized. Network-on-Chip (NoC), derived for efficiently managing the high-volume of communications, naturally becomes one of the top candidates to accelerate GNNs. However, there is a missing link between the optimize of GNN architecture and the NoC design.
Daniel Manu, Shaoyi Huang, Caiwen Ding, Lei Yang 0018
ACM Great Lakes Symposium on VLSI4
2021 Can Noise on Qubits Be Learned in Quantum Neural Network? A Case Study on QuantumFlow (Invited Paper)
abstract
In the noisy intermediate-scale quantum (NISQ) era, one of the key questions is how to deal with the high noise level existing in physical quantum bits (qubits). Quantum error correction is promising but requires an extensive number (e.g., over 1,000) of physical qubits to create one ”perfect” qubit, exceeding the capacity of the existing quantum computers. This paper aims to tackle the noise issue from another angle: instead of creating perfect qubits for general quantum algorithms, we investigate the potential to mitigate the noise issue for dedicate algorithms. Specifically, this paper targets quantum neural network (QNN), and proposes to learn the errors in the training phase, so that the identified QNN model can be resilient to noise. As a result, the implementation of QNN needs no or a small number of additional physical qubits, which is more realistic for the near-term quantum computers. To achieve this goal, an application-specific compiler is essential: on the one hand, the error cannot be learned if the mapping from logical qubits to physical qubits exists randomness; on the other hand, the compiler needs to be efficient so that the lengthy training procedure can be completed in a reasonable time. In this paper, we utilize the recent QNN framework, QuantumFlow, as a case study. Experimental results show that the proposed approach can optimize QNN models for different errors in qubits, achieving up to 28% accuracy improvement compared with the model obtained by the error-agnostic training.
Zhiding Liang, Zhepeng Wang 0001, Junhuan Yang, Lei Yang 0018, Yiyu Shi 0001, Weiwen Jiang
ICCAD4
2021 FL-DISCO: Federated Generative Adversarial Network for Graph-based Molecule Drug Discovery: Special Session Paper
abstract
The outbreak of the global COVID-19 pandemic emphasizes the importance of collaborative drug discovery for high effectiveness; however, due to the stringent data regulation, data privacy becomes an imminent issue needing to be addressed to enable collaborative drug discovery. In addition to the data privacy issue, the efficiency of drug discovery is another key objective since infectious diseases spread exponentially and effectively conducting drug discovery could save lives. Advanced Artificial Intelligence (AI) techniques are promising to solve these problems: (1) Federated Learning (FL) is born to keep data privacy while learning data from distributed clients; (2) graph neural network (GNN) can extract structural properties of molecules whose underlying architecture is the connected atoms; and (3) generative adversarial network (GAN) can generate novel molecules while retaining the properties learned from the training data. In this work, we make the first attempt to build a holistic collaborative and privacy-preserving FL framework, namely FL-DISCO, which integrates GAN and GNN to generate molecular graphs. Experimental results demonstrate the effectiveness of FL-DISCO on: (1) IID data for ESOL and QM9, where FL-DISCO can generate highly novel compounds with high drug-likeliness, uniqueness and LogP scores compared to the baseline; (2) non-IID data for ESOL and QM9, where FL-DISCO generates 100% novel compounds with high validity and LogP scores compared to the baseline. We also demonstrate how different fractions of clients, generator and discriminator architectures affect our evaluation scores.
Daniel Manu, Yi Sheng 0001, Junhuan Yang, Jieren Deng, Tong Geng, Ang Li 0006, Caiwen Ding, Weiwen Jiang, Lei Yang 0018
ICCAD9
2021 Federated Contrastive Learning for Dermatological Disease Diagnosis via On-device Learning (Invited Paper)
abstract
Deep learning models have been deployed in an increasing number of edge and mobile devices to provide healthcare. These models rely on training with a tremendous amount of labeled data to achieve high accuracy. However, for medical applications such as dermatological disease diagnosis, the private data collected by mobile dermatology assistants exist on distributed mobile devices of patients, and each device only has a limited amount of data. Directly learning from limited data greatly deteriorates the performance of learned models. Federated learning (FL) can train models by using data distributed on devices while keeping the data local for privacy. Existing works on FL assume all the data have ground-truth labels. However, medical data often comes without any accompanying labels since labeling requires expertise and results in prohibitively high labor costs. The recently developed self-supervised learning approach, contrastive learning (CL), can leverage the unlabeled data to pre-train a model for learning data representations, after which the learned model can be fine-tuned on limited labeled data to perform dermatological disease diagnosis. However, simply combining CL with FL as federated contrastive learning (FCL) will result in ineffective learning since CL requires diverse data for accurate learning but each device in FL only has limited data diversity. In this work, we propose an on-device FCL framework for dermatological disease diagnosis with limited labels. Features are shared among devices in the FCL pre-training process to provide diverse and accurate contrastive information without sharing raw data for privacy. After that, the pre-trained model is fine-tuned with local labeled data independently on each device or collaboratively with supervised federated learning on all devices. Experiments on dermatological disease datasets show that the proposed framework effectively improves the recall and precision of dermatological disease diagnosis compared with state-of-the-art methods.
Yawen Wu, Dewen Zeng, Zhepeng Wang 0001, Yi Sheng 0001, Lei Yang 0018, Alaina J. James, Yiyu Shi 0001, Jingtong Hu
ICCAD5
2021 HuGE: An Entropy-driven Approach to Efficient and Scalable Graph Embeddings
abstract
Graph embedding is becoming widely adopted as an efficient way to learn graph representations required to solve graph analytics problems. However, most existing graph embedding methods, owing to computation-efficiency challenges for large-scale graphs, generally employ a one-size-fits-all strategy to extract information, resulting in a large amount of redundant or inaccurate representations. In this work, we propose HuGE, an efficient and scalable graph embedding method enabled by an entropy-driven mechanism. Specifically, HuGE leverages hybrid-property heuristic random walk to capture node features, which considers both node degree and the number of common neighbors in each walking step. More importantly, to guarantee information effectiveness of sampling, HuGE adopts two heuristic methods to decide the random walk length and the number of walks per node, respectively. Extensive experiments on real-world graphs demonstrate that HuGE achieves both efficiency and performance advantages over recent popular graph embedding approaches. For link prediction and multi-label classification, our approach not only offers >10% average gains, but also exhibits 22×-126× speedup compared with existing sampling-based methods.
Peng Fang 0002, Fang Wang 0001, Zhan Shi 0001, Hong Jiang 0001, Dan Feng 0001, Lei Yang 0018
ICDE6
2021 Device-Circuit-Architecture Co-Exploration for Computing-in-Memory Neural Accelerators
abstract
Co-exploration of neural architectures and hardware design is promising due to its capability to simultaneously optimize network accuracy and hardware efficiency. However, state-of-the-art neural architecture search algorithms for the co-exploration are dedicated for the conventional von-Neumann computing architecture, whose performance is heavily limited by the well-known memory wall. In this article, we are the first to bring the computing-in-memory architecture, which can easily transcend the memory wall, to interplay with the neural architecture search, aiming to find the most efficient neural architectures with high network accuracy and maximized hardware efficiency. Such a novel combination makes opportunities to boost performance, but also brings a bunch of challenges: The optimization space spans across multiple design layers from device type and circuit topology to neural architecture; and the presence of device variation may drastically degrade the neural network performance. To address these challenges, we propose a cross-layer exploration framework, namely NACIM, which jointly explores device, circuit and architecture design space and takes device variation into consideration to find the most robust neural architectures, coupled with the most efficient hardware design. Experimental results demonstrate that NACIM can find the robust neural network with 0.45 percent accuracy loss in the presence of device variation, compared with a 76.44 percent loss from the state-of-the-art NAS without consideration of variation; in addition, NACIM achieves an energy efficiency up to 16.3 TOPs/W, 3.17x higher than the state-of-the-art NAS.
Weiwen Jiang, Qiuwen Lou, Zheyu Yan, Lei Yang 0018, Jingtong Hu, Xiaobo Sharon Hu, Yiyu Shi 0001
IEEE Trans. Computers4
2021 Reduced Worst-Case Communication Latency Using Single-Cycle Multihop Traversal Network-on-Chip
abstract
The communication latency in traditional network-on-chip (NoC) with hop-by-hop traversal is inherently restricted by the distance between source-destination communicating pairs. SMART, as one of the dynamically reconfigurable NoC architectures, enables the new feature of single-cycle long-distance communication by building a direct bypass path between distant cores dynamically at runtime. With the increasing of the number of integrated cores in multi/many-core systems, SMART has been deemed a promising communication backbone in such systems. However, SMART is generally optimized for average-case performance for best-effort traffics, not offering real-time guaranteed services for real-time traffics, and thus SMART often shows extremely poor real-time performance (e.g., schedulability). To make SMART latency-predictable for real-time traffics, by combining with the single-cycle bypass forwarding technique, in this article, we first propose a priority-preemptive scheduling to allow contending packets to be arbitrated according to predefined priorities. Based on the priority-based scheduling, for the real-time packet flows with given flow mapping and predefined priorities, we then propose a real-time communication analysis model, by considering shared virtual channels (or priority levels) and arbitrary-deadline real-time packet flows, to predict theworst-case communication latencyand validate the schedulability. Through theoretical and experimental comparison, theworst-case communication latencyof the analyzed packet flows is reduced significantly compared with that of the traditional priority-preemptive NoCs with hop-by-hop traversal and the original distance-based SMART, thus improving the schedulability.
Peng Chen 0027, Weichen Liu 0001, Hui Chen 0016, Shiqing Li, Mengquan Li, Lei Yang 0018, Nan Guan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 Contention-Aware Routing for Thermal-Reliable Optical Networks-on-Chip
abstract
Optical network-on-chip (ONoC) architecture offers ultrahigh bandwidth, low latency, and low power dissipation for new-generation manycore systems. However, the benefits in communication performance and energy efficiency will be diminished by communication contention. The intrinsic thermal susceptibility is another challenge for ONoC designs. Under on-chip temperature variations, core functional devices suffer from significant thermal-induced optical power loss, which seriously threatens ONoCs' reliability. In this article, we develop novel routing techniques to resolve both issues for ONoCs. By analyzing the thermal effect in ONoCs, we first present a routing criterion at the network level. Combined with device-level thermal tuning, it can implement thermal-reliable ONoCs. Two routing approaches, including a mixed-integer linear programming (MILP) model and a heuristic algorithm (called CAR), are further proposed to minimize communication conflicts based on guaranteed thermal reliability, and meanwhile, maximize the communication energy efficiency in the presence of on-chip thermal variations. By applying the criterion, our approaches achieve excellent performance with largely reduced complexity of design space exploration. The evaluation results based on both synthetic traffic patterns and realistic benchmarks validate the effectiveness of our approaches with an average of 126.95% improvement in communication performance and 16.12% reduction in energy overhead compared to state-of-the-art techniques. CAR only introduces 7.20% performance difference compared to the MILP model and is more scalable to large-size ONoCs.
Mengquan Li, Weichen Liu 0001, Luan H. K. Duong, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Contention Minimized Bypassing in SMART NoC
abstract
SMART, a recently proposed dynamically reconfigurable NoC, enables single-cycle long-distance communication by building single-bypass paths. However, such a single-cycle single-bypass path will be broken when contention occurs. Thus, lower-priority packets will be buffered at intermediate routers with blocking latency from higher-priority packets, and extra router-stage latency to rebuild remaining path, reducing the bypassing benefits that SMART offers. In this paper, we for the first time propose an effective routing strategy to achieve nearly contention-free bypassing in SMART NoC. Specifically, we identify two different routes for communication pairs: direct route, with which data can reach the destination in a single bypass; and indirect route, with which data can reach the destination in two bypasses via an intermediate router. If a direct route is not found, we would alternatively resort to an indirect route in advance to eliminate the blocking latency, at the cost of only one router-stage latency. Compared with the current routing, our new approach can effectively isolate conflicting communication pairs, greatly balance the traffic loads and fully utilize bypass paths. Experiments show that our approach makes 22.6% performance improvement on average in terms of communication latency.
Peng Chen 0027, Weichen Liu 0001, Mengquan Li, Lei Yang 0018, Nan Guan
ASP-DAC4
2020 Co-Exploring Neural Architecture and Network-on-Chip Design for Real-Time Artificial Intelligence
abstract
Hardware-aware Neural Architecture Search (NAS), which automatically finds an architecture that works best on a given hardware design, has prevailed in response to the ever-growing demand for real-time Artificial Intelligence (AI). However, in many situations, the underlying hardware is not pre-determined. We argue that simply assuming an arbitrary yet fixed hardware design will lead to inferior solutions, and it is best to co-explore neural architecture space and hardware design space for the best pair of neural architecture and hardware design. To demonstrate this, we employ Network-on-Chip (NoC) as the infrastructure and propose a novel framework, namely NANDS, to co-explore NAS space and NoC Design Search (NDS) space with the objective to maximize accuracy and throughput. Since two metrics are tightly coupled, we develop a multi-phase manager to guide NANDS to gradually converge to solutions with the best accuracy-throughput tradeoff. On top of it, we propose techniques to detect and alleviate timing performance bottleneck, which allows better and more efficient exploration of NDS space. Experimental results on common datasets, CIFAR10, CIFAR-100 and STL-10, show that compared with state-of-the-art hardware-aware NAS, NANDS can achieve 42.99% higher throughput along with 1.58% accuracy improvement. There are cases where hardware-aware NAS cannot find any feasible solutions while NANDS can.
Lei Yang 0018, Weiwen Jiang, Weichen Liu 0001, Edwin H.-M. Sha, Yiyu Shi 0001, Jingtong Hu
ASP-DAC1
2020 Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple Tasks
abstract
Neural Architecture Search (NAS) has demonstrated its power on various AI accelerating platforms such as Field Programmable Gate Arrays (FPGAs) and Graphic Processing Units (GPUs). However, it remains an open problem how to integrate NAS with Application-Specific Integrated Circuits (ASICs), despite them being the most powerful AI accelerating platforms. The major bottleneck comes from the large design freedom associated with ASIC designs. Moreover, with the consideration that multiple DNNs will run in parallel for different workloads with diverse layer operations and sizes, integrating heterogeneous ASIC sub-accelerators for distinct DNNs in one design can significantly boost performance, and at the same time further complicate the design space. To address these challenges, in this paper we build ASIC template set based on existing successful designs, described by their unique dataflows, so that the design space is significantly reduced. Based on the templates, we further propose a framework, namely ASICNAS, which can simultaneously identify multiple DNN architectures and the associated heterogeneous ASIC accelerator design, such that the design specifications (specs) can be satisfied, while the accuracy can be maximized. Experimental results show that compared with successive NAS and ASIC design optimizations which lead to design spec violations, ASICNAS can guarantee the results to meet the design specs with 17.77%, 2.49×, and 2.32× reductions on latency, energy, and area and less than 1.6% accuracy loss. To the best of the authors’ knowledge, this is the first work on neural architecture and ASIC accelerator design co-exploration.
Lei Yang 0018, Zheyu Yan, Meng Li 0004, Hyoukjun Kwon, Liangzhen Lai, Tushar Krishna, Vikas Chandra, Weiwen Jiang, Yiyu Shi 0001
DAC1
2020 Standing on the Shoulders of Giants: Hardware and Neural Architecture Co-Search With Hot Start
abstract
Hardware and neural architecture co-search that automatically generates artificial intelligence (AI) solutions from a given dataset are promising to promote AI democratization; however, the amount of time that is required by current co-search frameworks is in the order of hundreds of GPU hours for one target hardware. This inhibits the use of such frameworks on commodity hardware. The root cause of the low efficiency in existing co-search frameworks is the fact that they start from a “cold” state (i.e., search from scratch). In this article, we propose a novel framework, namely, HotNAS, that starts from a “hot” state based on a set of existing pretrained models (also known as model zoo) to avoid lengthy training time. As such, the search time can be reduced from 200 GPU hours to less than 3 GPU hours. In HotNAS, in addition to hardware design space and neural architecture search space, we further integrate a compression space to conduct model compressing during the co-search, which creates new opportunities to reduce latency, but also brings challenges. One of the key challenges is that all of the above search spaces are coupled with each other, e.g., compression may not work without hardware design support. To tackle this issue, HotNAS builds a chain of tools to design hardware to support compression, based on which a global optimizer is developed to automatically co-search all the involved search spaces. Experiments on ImageNet dataset and Xilinx FPGA show that, within the timing constraint of 5 ms, neural architectures generated by HotNAS can achieve up to 5.79% Top-1 and 3.97% Top-5 accuracy gain, compared with the existing ones.
Weiwen Jiang, Lei Yang 0018, Sakyasingha Dasgupta, Jingtong Hu, Yiyu Shi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Hardware/Software Co-Exploration of Neural Architectures
abstract
We propose a novel hardware and software co-exploration framework for efficient neural architecture search (NAS). Different from existing hardware-aware NAS which assumes a fixed hardware design and explores theNAS spaceonly, our framework simultaneously explores both the architecture search space and thehardware design spaceto identify the best neural architecture and hardware pairs that maximize both test accuracy and hardware efficiency. Such a practice greatly opens up the design freedom and pushes forward the Pareto frontier between hardware efficiency and test accuracy for better design tradeoffs. The framework iteratively performs a two-level (fast and slow) exploration. Without lengthy training, the fast exploration can effectively fine-tune hyperparameters and prune inferior architectures in terms of hardware specifications, which significantly accelerates the NAS process. Then, the slow exploration trains candidates on a validation set and updates a controller using the reinforcement learning to maximize the expected accuracy together with the hardware efficiency. In this article, we demonstrate that the co-exploration framework can effectively expand the search space to incorporate models with high accuracy, and we theoretically show that the proposed two-level optimization can efficiently prune inferior solutions to better explore the search space. The experimental results on ImageNet show that the co-exploration NAS can find solutions with the same accuracy, 35.24% higher throughput, 54.05% higher energy efficiency, compared with the hardware-aware NAS.
Weiwen Jiang, Lei Yang 0018, Edwin H.-M. Sha, Qingfeng Zhuge, Shouzhen Gu, Sakyasingha Dasgupta, Yiyu Shi 0001, Jingtong Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Routing in optical network-on-chip: minimizing contention with guaranteed thermal reliability
abstract
Communication contention and thermal susceptibility are two potential issues in optical network-on-chip (ONoC) architecture, which are both critical for ONoC designs. However, minimizing conflict and guaranteeing thermal reliability are incompatible in most cases. In this paper, we present a routing criterion in the network level. Combined with device-level thermal tuning, it can implement thermal-reliable ONoC. We further propose two routing approaches (including a mixed-integer linear programming (MILP) model and a heuristic algorithm (CAR)) to minimize communication conflict based on the guaranteed thermal reliability, and meanwhile, mitigate the energy overheads of thermal regulation in the presence of chip thermal variations. By applying the criterion, our approaches achieve excellent performance with largely reduced complexity of design space exploration. Evaluation results on synthetic communication traces and realistic benchmarks show that the MILP-based approach achieves an average of 112.73% improvement in communication performance and 4.18% reduction in energy overhead compared to state-of-the-art techniques. Our heuristic algorithm only introduces 4.40% performance difference compared to the optimal results and is more scalable to large-size ONoCs.
Mengquan Li, Weichen Liu 0001, Lei Yang 0018, Peng Chen 0027, Duo Liu 0002, Nan Guan
ASP-DAC3
2019 Accuracy vs. Efficiency: Achieving Both through FPGA-Implementation Aware Neural Architecture Search
abstract
A fundamental question lies in almost every application of deep neural networks: what is the optimal neural architecture given a specific data set? Recently, several Neural Architecture Search (NAS) frameworks have been developed that use reinforcement learning and evolutionary algorithm to search for the solution. However, most of them take a long time to find the optimal architecture due to the huge search space and the lengthy training process needed to evaluate each candidate. In addition, most of them aim at accuracy only and do not take into consideration the hardware that will be used to implement the architecture. This will potentially lead to excessive latencies beyond specifications, rendering the resulting architectures useless. To address both issues, in this paper we use Field Programmable Gate Arrays (FPGAs) as a vehicle to present a novel hardware-aware NAS framework, namely FNAS, which will provide an optimal neural architecture with latency guaranteed to meet the specification. In addition, with a performance abstraction model to analyze the latency of neural architectures without training, our framework can quickly prune architectures that do not satisfy the specification, leading to higher efficiency. Experimental results on common data set such as ImageNet show that in the cases where the state-of-the-art generates architectures with latencies 7.81× longer than the specification, those from FNAS can meet the specs with less than 1% accuracy loss. Moreover, FNAS also achieves up to 11.13× speedup for the search process. To the best of the authors' knowledge, this is the very first hardware aware NAS.
Weiwen Jiang, Xinyi Zhang 0001, Edwin H.-M. Sha, Lei Yang 0018, Qingfeng Zhuge, Yiyu Shi 0001, Jingtong Hu
DAC4
2019 XFER: A Novel Design to Achieve Super-Linear Performance on Multiple FPGAs for Real-Time AI
abstract
Real-time inference with low latency requirement has become increasingly important for numerous applications in both cloud computing and edge computing. The FPGA-based Deep Neural Network (DNN) accelerators have demonstrated the superior performance and energy efficiency over CPUs and GPUs; in addition, for real-time AI with low batch size, FPGA is expected to achieve further performance improvement over the general purpose computing platform. However, the performance gain of the single-FPGA design is hindered by the limited on-chip resource. In this paper, we leverage a cluster of FPGAs to fully exploit the parallelism in DNNs with the objective of obtaining super-linear performance. To achieve this goal, a novel design, "XFER", is proposed to deploy DNNs to FPGA cluster by splitting the DNN layer to multiple FPGAs and moving traffics from memory bus to inter-FPGA links. The resultant system can achieve both workload balance and traffic balance. As a case study, we implement Convolutional Neural Networks (CNNs) on ZCU102 FPGA boards. Evaluation results demonstrate that XFER on two FPGAs can achieve 3.48x speedup compared with state-of-the-art FPGA designs, achieving super-linear speedup.
Weiwen Jiang, Xinyi Zhang 0001, Edwin H.-M. Sha, Qingfeng Zhuge, Lei Yang 0018, Yiyu Shi 0001, Jingtong Hu
FPGA5
2019 Optimal Application Mapping and Scheduling for Network-on-Chips with Computation in STT-RAM Based Router
abstract
Spin-Torque Transfer Magnetic RAM (STT-RAM), one of the emerging nonvolatile memory (NVM) technologies explored as the replacement for SRAM memory architectures, is particularly promising due to the fast access speed, high integration density, and zero standby power consumption. Recently, hybrid deigns with SRAM and STT-RAM buffers for routers in Network-on-Chip (NoC) systems have been widely implemented to maximize the mutually complementary characteristics of different memory technologies, and leverage the efficiency of intra-router latency and system power consumption. With the realization of Processing-in-Memory enabled by STT-RAM, in this paper, we novelly offload the execution from processors to the STT-RAM based on-chip routers to improve the application performance. On top of the hybrid buffer design in routers, we further present system-level approaches, including an ILP model and polynomial-time heuristic algorithms, to fine-tune the application mapping and scheduling on NoCs, with the objectives of improving system performance-energy efficiency. Network overhead caused by flit conflict in conventional communication circumstances can be ideally avoided by computing the contended flits in intermediate routers; meanwhile, the pressure of heavy workload on processors can be relieved by transferring partial operations to routers, such that network latency and system power consumption can be significantly reduced. Experimental results demonstrate that application schedule length and system energy consumption can be reduced by 35.62, 32.87 percent on average, respectively, in extensive evaluation experiments on PARSEC benchmark applications. In particular, the achievements of application performance and energy efficiency, averagely 36.44 and 33.19 percent, for the CNN application AlexNet have verified the practicability and effectiveness of our presented approaches.
Lei Yang 0018, Weichen Liu 0001, Nan Guan, Nikil Dutt
IEEE Trans. Computers1
2019 On the Design of Time-Constrained and Buffer-Optimal Self-Timed Pipelines
abstract
Pipelining is a powerful technique to achieve high performance in computing systems. However, as computing platforms become large-scale and integrate with heterogeneous processing elements (PEs) (CPUs, GPUs, field-programmable gate arrays, etc.), it is difficult to employ a global clock to achieve synchronous pipelines. Therefore, self-timed (or asynchronous) pipelines are usually adopted. Nevertheless, due to their complex running behavior, the performance modeling and systematic optimizations for self-timed pipeline (STP) systems are more complicated than those for synchronous ones. This paper employs marked graph theory to model STPs and presents algorithms to detect performance bottlenecks. Based on the proposed model, we observe that the system performance can be improved by inserting buffers. Due to the limited memory resources on the PEs, it is critical to minimize the number of buffers for STPs while satisfying the required timing constraints. In this paper, we propose integer linear programming formulations to obtain the optimal solutions and devise efficient algorithms to obtain the near-optimal solutions. Experimental results show that the proposed algorithms can achieve 53.10% improvement in the maximum performance and 54.04% reduction in the number of buffers, compared with the technique for the slack matching problem.
Weiwen Jiang, Edwin H.-M. Sha, Qingfeng Zhuge, Lei Yang 0018, Xianzhang Chen, Jingtong Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Energy-Efficient Application Mapping and Scheduling for Lifetime Guaranteed MPSoCs
abstract
Energy optimization is one of the most critical objectives for the synthesis of multiprocessor system-on-chip (MPSoC). Besides, to ensure a long processor lifetime and to maintain a safe chip temperature are also important for multiprocessor manufactures under deep submicrometer process technologies. This paper presents a mixed integer linear programming (MILP) model to determine the mapping and scheduling of real-time applications onto embedded MPSoC platforms, such that the total energy consumption is minimized with the lifetime reliability constraint and the temperature threshold constraint satisfied. We develop a lightweight temperature model that can be integrated in the MILP model to predict the chip temperature accurately and efficiently. By exploiting the dynamic voltage and frequency scaling capability of modern processors, processor voltage/frequency assignment is also considered in our MILP model. Extensive performance evaluations on synthetic and real-world applications demonstrate the effectiveness of the proposed approach. Our MILP model achieves an average reduction of 19.09% and 28.53% total energy in comparison with two state-of-the-art techniques on the basis of guaranteeing the safe chip temperature and system lifetime reliability.
Weichen Liu 0001, Juan Yi, Mengquan Li, Peng Chen 0027, Lei Yang 0018
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2019 Achieving Super-Linear Speedup across Multi-FPGA for Real-Time DNN Inference
abstract
Real-time Deep Neural Network (DNN) inference with low-latency requirement has become increasingly important for numerous applications in both cloud computing (e.g., Apple’s Siri) and edge computing (e.g., Google/Waymo’s driverless car). FPGA-based DNN accelerators have demonstrated both superior flexibility and performance; in addition, for real-time inference with low batch size, FPGA is expected to achieve further performance improvement. However, the performance gain from the single-FPGA design is obstructed by the limited on-chip resource. In this paper, we employ multiple FPGAs to cooperatively run DNNs with the objective of achieving super-linear speed-up against single-FPGA design. In implementing such systems, we found two barriers that hinder us from achieving the design goal: (1) the lack of a clear partition scheme for each DNN layer to fully exploit parallelism, and (2) the insufficient bandwidth between the off-chip memory and the accelerator due to the growing size of DNNs. To tackle these issues, we propose a general framework, “Super-LIP”, which can support different kinds of DNNs. In this paper, we take Convolutional Neural Network (CNN) as a vehicle to illustrate Super-LIP. We first formulate an accurate system-level model to support the exploration of best partition schemes. Then, we develop a novel design methodology to effectively alleviate the heavy loads on memory bandwidth by moving traffic from memory bus to inter-FPGA links. We implement Super-LIP based on ZCU102 FPGA boards. Results demonstrate that Super-LIP with 2 FPGAs can achieve 3.48× speedup, compared to the state-of-the-art single-FPGA design. What is more, as the number of FPGAs scales up, the system latency can be further reduced while maintaining high energy efficiency.
Weiwen Jiang, Edwin H.-M. Sha, Xinyi Zhang 0001, Lei Yang 0018, Qingfeng Zhuge, Yiyu Shi 0001, Jingtong Hu
ACM Trans. Embed. Comput. Syst.4
2018 On the Design of Reliable Heterogeneous Systems via Checkpoint Placement and Core Assignment
abstract
This paper studies two basic problems in the design of high-performance and high-reliability heterogeneous systems: (1) what type of core to execute each task, and (2) where to place checkpoints in the execution of tasks. The implementation of checkpointing techniques on the novel persistent memory (e.g., 3D Xpoint memory) based heterogeneous systems faces a bundle of new problems. First, the assignments of tasks may greatly influence the execution time of the whole application. Therefore, with the same time constraint, the reliability of the resultant system can be significantly affected. Second, creating checkpoints will incur heavy writes on persistent memories and reduce the lifetime of devices. In this paper, we optimally construct reliable systems by assigning tasks to the most suitable cores and placing minimum number of checkpoints in the application, such that the resultant system can satisfy the time constraint in the presence of faults. We devise an efficient dynamic programming algorithm to obtain the optimal assignment and checkpoint placement. Experimental results demonstrate that, compared with existing approaches, our technique can achieve 44% reductions on the number of checkpoints on average.
Edwin H.-M. Sha, Hailiang Dong, Weiwen Jiang, Qingfeng Zhuge, Xianzhang Chen, Lei Yang 0018
ACM Great Lakes Symposium on VLSI6
2018 User Experience-Enhanced and Energy-Efficient Task Scheduling on Heterogeneous Multi-Core Mobile Systems
abstract
Heterogeneous Multi-Core Mobile Systems has been widely used to improve performance. However, it faces with the challenge of tradeoff between energy saving and user experience. ARM big. LITTLE architecture, a heterogeneous computing architecture, is a power-optimization technology. In most big. LITTLE devices, however, it still cannot achieve excellent user experience and higher energy saving. In this paper, we propose an improved task scheduling (UCES-GTS) by introducing the concept of user-centric task on big. LITTLE mobile device. In order to enhance user experience, the response time of user-centric tasks is shortened with reducing slack time of them properly. We then present a detailed algorithm to compute appropriate frequency and allocate the CPU resources to each task. The experimental evaluation results show that our improved global task scheduling model can achieve 17 % and 8 % energy saving average compared with the clustered switching scheduling and the original global task scheduling respectively. And the response time of user-centric tasks can decrease 27 % average, which means excellent user experience.
Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye
ICPADS5
2018 Fine-Grained Task-Level Parallel and Low Power H.264 Decoding in Multi-Core Systems
abstract
In the past few years, the extinction of Moore's Law makes people reconsider the solutions for dealing with the low computing resource utilization of applications on multicore processor systems. However, making good use of computing resources in multi-core processors systems is not easy due to the differences between single-core and multi-core architecture. Nowadays short video apps like Instagram and Tik Tok have successfully caught people's eyes by fascinating short videos, typically just 10 to 30 seconds long, uploaded by the users of apps. And almost all of these videos are recorded by their mobile devices, which are typically HD (High Definition) or FHD (Full High Definition) videos, which prefer to be encoded/decoded by H.264/AVC rather then HEVC (High Efficiency Video Coding) on mobile devices in view of the energy consumption and decoding speed. How to dive the huge potential of the computing resource on multi-core mobile devices to speed up decoding these videos while consuming low energy, is a big challenge. In our previous work [1], a relatively simple parallel framework was proposed to implement a parallel H.264/ AV C decoder. This work further proposes a more detailed systematic task-level parallel framework, together with an energy saving strategy based on this framework, to research a new H.264/AVC decoder on multi-core processor systems. The proposed parallel method is composed of a set of rules to guide parallel software programming (PSPR) and a software parallelization framework (SPF). The PSPR is applied in pre-processing steps to address the potential issues limiting the inherent parallelism, and the SPF is applied to parallelize the original serial programs. After the parallelization is successfully deployed, DVFS technique would be applied to decrease the power dissipation based on the SPF. Results show that proposed solutions make a significant improvement in decoding speed of 32% at 720p, 27% at 1080p and 29% at 2160p, and in energy savings of 25% at 720p, 25% at 1080p and 23% at 2160p on a four-core workstation running Linux, compared to the original serial H.264/ AV C decoder. The results demonstrate our methods are effective and scalable, served as a reference for future parallel software development.
Wenyang Liu, Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye
ICPADS5
2018 Thermal-Aware Task Mapping on Dynamically Reconfigurable Network-on-Chip Based Multiprocessor System-on-Chip
abstract
Dark silicon is the phenomenon that a fraction of many-core chip has to be turned off or run in a low-power state in order to maintain the safe chip temperature. System-level thermal management techniques normally map application on non-adjacent cores, while communication efficiency among these cores will be oppositely affected over conventional network-on-chip (NoC). Recently, SMART NoC architecture is proposed, enabling single-cycle multi-hop bypass channels to be built between distant cores at runtime, to reduce communication latency. However, communication efficiency of SMART NoC will be diminished by communication contention, which will in turn decrease system performance. In this paper, we first propose an Integer-Linear Programming (ILP) model to properly address communication problem, which generates the optimal solutions with the consideration of inter-processor communication. We further present a novel heuristic algorithm for task mapping in dark silicon many-core systems, called TopoMap, on top of SMART architecture, which can effectively solve communication contention problem in polynomial time. With fine-grained consideration of chip thermal reliability and inter-processor communication, presented approaches are able to control the reconfigurability of NoC communication topology in task mapping and scheduling. Thermal-safe system is guaranteed by physically decentralized active cores, and communication overhead is reduced by the minimized communication contention and maximized bypass routing. Performance evaluation on PARSEC shows the applicability and effectiveness of the proposed techniques, which achieve on average 42.5 and 32.4 percent improvement in communication and application performance, and 32.3 percent reduction in system energy consumption, compared with state-of-the-art techniques. TopoMap only introduces 1.8 percent performance difference compared to ILP model and is more scalable to large-size NoCs.
Weichen Liu 0001, Lei Yang 0018, Weiwen Jiang, Liang Feng 0001, Nan Guan, Wei Zhang 0012, Nikil Dutt
IEEE Trans. Computers2
2018 Heterogeneous FPGA-Based Cost-Optimal Design for Timing-Constrained CNNs
abstract
Field programmable gate array (FPGA) has been one of the most popular platforms to implement convolutional neural networks (CNNs) due to its high performance and cost efficiency; however, limited by the on-chip resources, the existing single-FPGA architectures cannot fully exploit the parallelism in CNNs. In this paper, we explore heterogeneous FPGA-based designs to effectively leverage both task and data parallelism, such that the resultant system can achieve the minimum cost while satisfying timing constraints. In order to maximize the task parallelism, we investigate two critical problems: 1) buffer placement, where to place buffers to partition CNNs into pipeline stages and 2) task assignment, what type of FPGA to implement different CNN layers. We first formulate the system-level optimization problem with a mixed integer linear programming model. Then, we propose an efficient dynamic programming algorithm to obtain the optimal solutions. On top of that, we devise an efficient algorithm that exploits data parallelism within CNN layers to further improve cost efficiency. Evaluations on well-known CNNs demonstrate that the proposed techniques can obtain an average of 30.82% reduction in system cost under the same timing constraint, and an average of 1.5 times speedup in performance under the same cost budget, compared with the state-of-the-art techniques.
Weiwen Jiang, Edwin H.-M. Sha, Qingfeng Zhuge, Lei Yang 0018, Xianzhang Chen, Jingtong Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 Chip Temperature Optimization for Dark Silicon Many-Core Systems
abstract
In the dark silicon era, a fundamental problem is given a real-time computation demand, how to determine if an on-chip multiprocessor system is able to accept this demand and to maintain its reliability by keeping every core within a safe temperature range. In this paper, a practical thermal model is described for quick chip temperature prediction. Integrated with the thermal model, we present a mixed integer linear programming (MILP) model to find the optimal task-to-core assignment with the minimum chip peak temperature. For the worst case where even the minimum chip peak temperature exceeds the safe temperature, a heuristic algorithm, called temperature-constrained task selection (TCTS), is proposed to optimize the system performance within chip safe temperature. The optimality of the TCTS algorithm is formally proven. Extensive performance evaluations show that our thermal model achieves an average prediction accuracy of 0.0741 °C within 0.2392 ms. The MILP model reduces chip peak temperature of ~10 °C comparing with traditional techniques. The system performance is increased by 19.8% under safe temperature limitation. Due to the satisfying scalability of our MILP formulation, the chip peak temperature is further decreased by 5.06 °C via the TCTS algorithm. The feasibility of this systematical technique is testified in a real case study as well.
Mengquan Li, Weichen Liu 0001, Lei Yang 0018, Peng Chen 0027, Chao Chen 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Dark silicon-aware hardware-software collaborated design for heterogeneous many-core systems
abstract
ARM's big. LITTLE architecture coupled with Heterogeneous Multi-Processing (HMP) has enabled energy-efficient solutions in the dark silicon era. System-level techniques activate nonadjacent cores to eliminate chip thermal hotspot. However, it unexpectedly increases communication delay due to longer distance in network architectures, and in turn degrades application performance and system energy efficiency. In this paper, we present a novel hierarchical hardware-software collaborated approach to address the performance/temperature conflict in dark silicon many-core systems. Optimizations on interprocessor communication, application performance, chip temperature and energy consumption are well isolated and addressed in different phases. Evaluation results show that on average 22.57% reduction of communication latency, 23.04% improvement on energy efficiency and 6.11°C reduction of chip peak temperature are achieved compared with state-of-the-art techniques.
Lei Yang 0018, Weichen Liu 0001, Nan Guan, Mengquan Li, Peng Chen 0027, Edwin H.-M. Sha
ASP-DAC1
2017 Task Mapping on SMART NoC: Contention Matters, Not the Distance
abstract
On-chip communication is the bottleneck of system performance for NoC-based MPSoCs. SMART, a recently proposed NoC architecture, enables single-cycle multi-hop communications. In SMART NoCs, unconflicted messages can go through an express bypass and the communication efficiency is significantly improved, while conflicted messages have to be buffered for guaranteed delivery with extra delays. Therefore, that performance of SMART NoC may be seriously degraded when communication contention increases. In this paper, we present task mapping techniques to address this problem for SMART NoCs, with the consideration of communication contention, rather than inter-processor distance, by minimizing conflicts and thus maximizing bypass utilization. We first model the entire problem by ILP formulations to find the theoretically optimal solution, and further propose polynomial-time algorithms for contention-aware task mapping and message priority assignment. Communicating tasks can be mapped to distant processors in SMART NoCs as long as conflict-free communication paths can be established and bypass can be enabled. Evaluation results on real benchmarks show an average of 44.1% and 32.8% improvement in communication efficiency and application performance compared to state-of-the-art techniques. The proposed heuristic algorithms only introduce 1.9% performance difference compared to the ILP model and are more scalable to large-size NoCs.
Lei Yang 0018, Weichen Liu 0001, Peng Chen 0027, Nan Guan, Mengquan Li
DAC1
2017 Hardware-software collaboration for dark silicon heterogeneous many-core systems
Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Chao Chen 0004, Mengquan Li, Peng Chen 0027, Edwin H.-M. Sha
Future Gener. Comput. Syst.1
2017 FoToNoC: A Folded Torus-Like Network-on-Chip Based Many-Core Systems-on-Chip in the Dark Silicon Era
abstract
Dark silicon refers to the phenomenon that a fraction of a many-core chip has to become “dark” or “dim” in order to guarantee the system to be kept in a safe temperature range and allowable power budget. Techniques have been developed to selectively activate non-adjacent cores on many-core chip to avoid temperature hotspot, while resulting unexpected increase of communication overhead due to the longer average distance between active cores, and in turn affecting application performance and energy efficiency, when Network-on-Chip (NoC) is used as a scalable communication subsystem. To address the brand-new challenges brought by dark silicon, in this paper, we present FoToNoC, a Folded Torus-like NoC, coupled with a hierarchical management strategy for heterogeneous many-core systems. On top of it, objectives of maximizing application performance, energy efficiency and chip reliability are isolated and well achieved by hardware-software co-design in several different phases, including application mapping and scheduling, cluster management and DVFS control. Evaluations on PARSEC benchmark applications demonstrate the significance of the entire strategy. Compared with state-of-the-art approaches, the proposed FoToNoC organization can achieve on average 35.4 and 35.2 percent on communication efficiency and application performance improvement, respectively, when maintaining the safe chip temperature. The hierarchical cluster-based management strategy can further reduce an average 34.6 percent of the total energy consumption with a notable reduction on the chip peak temperature. The significant achievements on system energy efficiency and the reduction on chip temperature of H.264 decoder and DSP-stone benchmarks additionally verify the effectiveness of the proposed methods.
Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Mengquan Li, Peng Chen 0027, Edwin H.-M. Sha
IEEE Trans. Parallel Distributed Syst.1
2017 Optimal Functional-Unit Assignment for Heterogeneous Systems Under Timing Constraint
abstract
In high-level synthesis for real-time systems, it typically employs heterogeneous functional-unit types to achieve high-performance and low-cost designs. In the design phase, it is critical to determine which functional-unit type to be mapped for each operation in a given application such that the total cost is minimized while the deadline can be met. For a path or tree structured application, existing approaches can obtain the minimum-cost assignment, called “optimal assignment”, under which the resultant system satisfies a given timing constraint. However, it is still an open question whether there exist efficient algorithms to obtain the optimal assignment for the directed acyclic graph (DAG), or more generally, the data-flow graph with cycles (cyclic DFG). For DAGs, by analyzing the property of the problem, this paper designs an efficient algorithm to obtain the optimal assignments. For cyclic DFGs, we approach this problem with the combination of retiming technique to thoroughly explore the design space. We formulate a Mixed Integer Linear Programming (MILP) model to give the optimal solution. But because of the high degree of its time complexity, we devise a practical algorithm to obtain near-optimal solutions within a minute. Experimental results show the effectiveness of our algorithms. Specifically, compared with existing techniques, we can achieve 25.70 and 30.23 percent reductions in total cost on DAGs and cyclic DFGs, respectively.
Weiwen Jiang, Edwin H.-M. Sha, Xianzhang Chen, Lei Yang 0018, Lei Zhou 0020, Qingfeng Zhuge
IEEE Trans. Parallel Distributed Syst.4
2016 FoToNoC: A hierarchical management strategy based on folded lorus-like Network-on-Chip for dark silicon many-core systems
abstract
In this dark silicon era, techniques have been developed to selectively activate nonadjacent cores in physical locations to maintain the safe temperature and allowable power budget on a many-core chip. This will result in unexpected increase in the communication overhead due to longer average distance between active cores in a typical mesh-based Network-on-Chip (NoC), and in turn reduce the system performance and energy efficiency. In this paper, we present FoToNoC, a Folded Torus-like NoC, and a hierarchical management strategy on top of it, to address this tradeoff problem for heterogeneous many-core systems. Optimizations of chip temperature, inter-core communication, application performance, and system energy consumption are well isolated in FoToNoC, and addressed in different design phases and aspects. A cluster-based hierarchical strategy is proposed to manage the system adaptively in several different control levels. Compared with mesh-based systems on a set of synthetic and real benchmarks, FoToNoC can achieve on average 39.4% performance improvement when similar temperature conditions are maintained, and the proposed strategy can further reduce the total energy consumption by up to 42.0%.
Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Mengquan Li, Juan Yi, Edwin H.-M. Sha
ASP-DAC1
2016 Application Mapping and Scheduling for Network-on-Chip-Based Multiprocessor System-on-Chip With Fine-Grain Communication Optimization
abstract
Network-on-chip (NoC) is promising for the communication paradigm of the next-generation multiprocessor system-on-chip (MPSoC). As communication has become an integral part of on-chip computing, and even the performance bottleneck, researchers are paying much attention to its implementation and optimization. Traditional techniques that model communication inaccurately will lead to unexpected runtime performance, which is on average 90.8% worse than the predicted results based on observation, and are not suitable for the deep optimization of communication-intensive scenarios. In this paper, techniques are presented for the NoC-based MPSoCs that integrate optimization on interprocessor communications with the objective of minimizing the schedule length. A fine-grained integer-linear programming (ILP) model is proposed to properly address the communication latency with a network contention, which generates runtime scheduling with trivial performance difference from the predictions. We further propose a heuristic algorithm, unified priority-based scheduling (UPS), to effectively solve the contention problem in polynomial time by assigning priorities to messages. Evaluation results show that the solutions obtained by the ILP model outperform the state-of-the-art techniques by 31.1%, and UPS improves application performance by 34.7% and 44.4% compared with acquainted first-in-first-out (FIFO)-based and random-based methods. In addition, UPS achieves averagely 8.3% approximated results with the optimal solutions generated by ILP. A case study on H.264 high-definition television (HDTV) decoder and the digital signal processor (DSP) filter benchmarks achieves significant improvement on the performance and the results prediction accuracy, as well as the prominent reduction in the number of network contention and energy consumption.
Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Mengquan Li, Juan Yi, Edwin H.-M. Sha
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Prevent Deadlock and Remove Blocking for Self-Timed Systems
Edwin H.-M. Sha, Weiwen Jiang, Qingfeng Zhuge, Xianzhang Chen, Lei Yang 0018
ICA3PP (1)5
2014 On self-timed ring for consistent mapping and maximum throughput
abstract
Multiprocessor System-on-Chip employing self-timed technique becomes increasingly attractive due to its ability for exploiting high parallelism of applications. There have been many research efforts on studying self-timed techniques on hardware layer. However, these research results are unable to be applied to system synthesis; in particular, how to correctly and optimally map an application represented by a Data Flow Graph to a self-timed ring architecture remains unknown. Self-timed ring (STR) is a popular and easy to implemented architecture. This paper establishes a series of theorems about the setting of initial configuration to achieve correct mappings and the formulas of calculating corresponding throughputs of STR. Based on the understanding, we can obtain a correct initial configuration of STR. And an algorithm presented in the paper can also find the best initial configuration that achieves the maximum throughput of STR. Examples show maximum throughput algorithm achieves 51.11% improvement of throughput compared with non-optimized ones.
Weiwen Jiang, Qingfeng Zhuge, Juan Yi, Lei Yang 0018, Edwin H.-M. Sha
RTCSA4
2014 Contention-aware task and communication co-scheduling for network-on-chip based Multiprocessor System-on-Chip
abstract
To satisfy the ever increasing performance requirement of applications, Multiprocessor System-on-Chip (MPSoC) plays an irreplaceable role in embedded system these days. It is significant to effectively optimize communication for achieving maximum parallelism on MPSoC, especially on Network-on-Chip (NoC) based architectures. The problem of how to make an arbitration of communication congestion is remained unsolved. In this paper, we propose a reasonable Unified Priority-Based Scheduling (UPS) algorithm for task and communication co-scheduling with communication contention, which is based on a novel Task Communication Graph (TCG) model of an application. The proposed method is more accurate and effective to describe the overall process of applications. The experimental results show that the performance is improved by 31.1% on average of scheduling generated by our algorithm. It verifies that the proposed method in this paper can improve the performance of contention-aware task and communication scheduling on NoC-based MPSoC architecture.
Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Juan Yi, Duo Liu 0002, Qingfeng Zhuge
RTCSA1