Di Liu 0002

dblp:15/1777-2 · DBLP profile ↗
← Back
55ranked-venue papers
7as first author
38since 2021 · last 2026
0000-0002-4365-2768ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46 · 3 first-author · 34 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 4 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Bit-Shareable Inference on Microcontrollers
abstract
Many embedded systems are now being deployed in energy-constrained environments, with some systems utilizing energy-harvesting technologies. Consequently, the energy available to these systems is dynamic. For example, energy harvesting from the sun can provide excess energy during the daytime, but energy levels run low at night. In such energy-harvesting environments, low-power microcontroller (MCU) platforms are used to run machine learning inference, but their software is not adaptive to the energy fluctuations. BitSIM is the first to provide a clear methodology to train and deploy switchable-precision networks (SP-nets) that tackle the challenges of an MCU platform.BitSIM employs a novel quantizer, PolyQAT, which not only enables weight-sharing but also bit-shareable weights. In bit-shareable weights, the narrower-precision weight can be directly extracted from the wider weight. With PolyQAT, SP-nets can be trained with low precision (i.e., with weights of four bits or less), which enables the deployment of large networks with respect to the memory size of MCUs. For the deployment of the SP-nets, BitSIM considers one minimalistic MCU hardware extension that enables efficient execution of sub-byte quantized neural networks.
Charalampos Bezaitis, Yaman Umuroglu, Di Liu 0002, Magnus Själander
DATE3
2025 LODAP: On-device incremental learning via lightweight operations and data pruning
Biqing Duan, Di Liu 0002, Wei Zhou 0011, Zhenli He, Shengfa Miao
J. Syst. Archit.3
2025 Efficient Deep Learning Infrastructures for Embedded Computing Systems: A Comprehensive Survey and Future Envision
abstract
Deep neural networks (DNNs) have recently achieved impressive success across a wide range of real-world vision and language processing tasks, spanning from image classification to many other downstream vision tasks, such as object detection, tracking, and segmentation. However, previous well-established DNNs, despite being able to maintain superior accuracy, have also been evolving to be deeper and wider and thus inevitably necessitate prohibitive computational resources for both training and inference. This trend further enlarges the computational gap between computation-intensive DNNs and resource-constrained embedded computing systems, making it challenging to deploy powerful DNNs in real-world embedded computing systems towards ubiquitous embedded intelligence. To alleviate this computational gap and enable ubiquitous embedded intelligence, we focus in this survey on discussing recent efficient deep learning infrastructures for embedded computing systems, spanning from training to inference , from manual to automated , from convolutional neural networks to transformers , from transformers to vision transformers , from vision models to large language models , from software to hardware , and from algorithms to applications . Specifically, we discuss recent efficient deep learning infrastructures for embedded computing systems from the lens of (1) efficient manual network design for embedded computing systems, (2) efficient automated network design for embedded computing systems, (3) efficient network compression for embedded computing systems, (4) efficient on-device learning for embedded computing systems, (5) efficient large language models for embedded computing systems, (6) efficient deep learning software and hardware for embedded computing systems, and (7) efficient intelligent applications for embedded computing systems. We also envision promising future directions and trends, which have the potential to deliver more ubiquitous embedded intelligence. We believe this survey has its merits and can shed light on future research, which can largely help researchers to quickly and smoothly get started in this emerging field.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Hui Chen 0016, Guochu Xiong, Weichen Liu 0001
ACM Trans. Embed. Comput. Syst.2
2024 Pearls Hide Behind Linearity: Simplifying Deep Convolutional Networks for Embedded Hardware Systems via Linearity Grafting
abstract
The increasing complexity of convolutional neural networks (CNNs) has fueled a huge demand for compression. Nonetheless, network pruning, as the most effective knob, fails to deliver Pareto-optimal networks. To tackle this issue, we introduce a novel pruning-free compression framework dubbed Domino, pioneering to revisit the trade-off dilemma between accuracy and efficiency from a fresh perspective of linearity and non-linearity. Specifically, Domino leverages two predictors, including one vanilla latency predictor and one meta-accuracy predictor, to identify the less important non-linear building blocks, which are then grafted with the linear counterparts. And next, the grafted network is trained on target task to obtain decent accuracy, after which the grafted linear building block that contains multiple consecutive linear layers is reparameterized into one single linear layer to boost the efficiency on target hardware without degrading the accuracy on target task. Extensive experiments on two popular Nvidia Jetson embedded platforms (i.e., Xavier and Nano) and two representative networks (i.e., MobileNetV2 and ResNet50) clearly demonstrate the superiority of Domino. For example, Domino-Aggressive achieves +10.6%/+8.8% higher top-l/top-5 accuracy on ImageNet than ${\mathrm {MobileNetV}} 2 \times 0.2$, while bringing $\times 1.9/\times 1.3$ speedup on Xavier/Nano.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Hui Chen 0016, Shiqing Li, Guochu Xiong, Weichen Liu 0001
ASPDAC2
2024 Double-Win NAS: Towards Deep-to-Shallow Transformable Neural Architecture Search for Intelligent Embedded Systems
abstract
Thanks to the evolving network depth, convolutional neural networks (CNNs) have achieved impressive performance across various intelligent embedded scenarios towards embedded intelligence. Nonetheless, this trend also leads to degraded hardware efficiency as the network evolves deeper and deeper. In contrast, shallow networks exhibit superior hardware efficiency, which, unfortunately, suffer from inferior accuracy. To tackle this dilemma, we establish the first deep-to-shallow transformable neural architecture search (NAS) paradigm, namely Double-Win NAS (DW-NAS), which is dedicated to automatically exploring deep-to-shallow transformable networks to marry the best of both worlds. Extensive experiments on two NVIDIA Jetson intelligent embedded systems clearly show the superiority of DW-NAS over previous state-of-the-art methods.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Weichen Liu 0001
DAC2
2024 FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual Inspection
abstract
Federated learning (FL) is a collaborative learning scheme to train deep learning models, where collaborating parties can consolidate their models without sharing local data with other parties, hence preserving data privacy. Nevertheless, when implementing FL in Industrial visual inspection (IVI), the constraints posed by limited data availability and the intricate nature of the inspection tasks significantly impact the performance of the resulting model. This paper introduces FedTR, a novel FL framework incorporating transfer learning designed for Autonomous IVI, focusing on the challenging task of identifying label defects through end-to-end text recognition. Transfer learning is a method that leverages the knowledge of a pre-trained model to adapt to a different dataset. FedTR initially trains the model using a publicly available dataset, after which performs the essential federated learning process with model fine-tuning on the distributed and limited private data. Extensive experiment results demonstrate the effectiveness and feasibility of FedTR on private ink cartridge datasets for label defect identification. FedTR achieves an end-to-end text recognition word-level accuracy of 95.5% and 94.2% on homogeneous and heterogeneous data respectively. Additionally, it attains performance levels that are on par with those achieved through centralized training.
Vikash Sathiamoorthy, Shuo Huai, Hao Kong 0001, Di Liu 0002, Wendy Yong Yi Loy, Christian Makaya, Daren Ho, Ravi Subramaniam, Qian Lin 0001, Weichen Liu 0001
ACM Great Lakes Symposium on VLSI4
2024 Integrating Branching and Pruning for Efficient Hyperdimensional Computing
abstract
As an emerging brain-inspired computing method, hyperdimensional computing (HDC) has attracted increasing attention. HDC encodes data into a high-dimensional hypervector and makes predictions based on the similarity comparison of the hypervector. However, the computational overhead of an HDC model increases, and the accuracy decreases as the number of predicted classes increases. Branching has been proposed to address this issue in HDC models, but it also introduces a new challenge: increased memory overhead for storing new branching hypervectors. In this paper, we aim to address this issue by integrating branching with the pruning method in HDC models. To this end, we propose a simple yet effective two-level branching structure that can reduce the memory overhead required by branching hypervectors. In addition, we propose a novel variable-position pruning method capable of removing redundant dimensions from hypervectors at various positions. This can best identify the important dimensions of a hypervector and further mitigate the memory issue caused by the branching structure. Experimental results demonstrate that our two-level branching structure can slightly increase the accuracy across various datasets. With only a 0.5 % accuracy loss, our pruning method can remove more than half of the dimensions, thereby improving the HDC model's inference latency and energy consumption.
Zhiqian Guan, Di Liu 0002, Shengfa Miao, Fei Dai 0002
ICCD3
2024 Domino-Pro-Max: Toward Efficient Network Simplification and Reparameterization for Embedded Hardware Systems
abstract
The prohibitive complexity of convolutional neural networks (CNNs) has triggered an increasing demand for network simplification. To this end, one natural solution is to remove the redundant channels or layers to explore simplified network structures. However, the resulting simplified network structures often suffer from suboptimal accuracy-efficiency tradeoffs. To overcome such limitations, we, in this work, introduce a simple yet effective network simplification approach, namely Domino, which aims to comprehensively revisit the tradeoff dilemma between accuracy and efficiency from a new perspective of linearity and nonlinearity through linearity grafting. Furthermore, we also draw insights from Domino and introduce two enhanced variants, namely Domino-Pro and Domino-Pro-Max, to improve the attainable accuracy on target task without degrading the runtime efficiency on target hardware. Extensive experiments are conducted on two popular Nvidia Jetson embedded hardware systems (i.e., Xavier and Nano) and two representative deep convolutional networks (i.e., MobileNetV2 and ResNet50), which clearly demonstrate the superiority of Domino and its two enhanced variants over previous state-of-the-art methods.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Guochu Xiong, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 MUGNoC: A Software-Configured Multicast-Unicast-Gather NoC for Accelerating CNN Dataflows
abstract
Current communication infrastructures for convolutional neural networks (CNNs) only focus on specific transmission patterns, not applicable to benefit the whole system if the dataflow changes or different dataflows run in one system. To reduce data movement, various CNN dataflows are presented. For these dataflows, parameters and results are delivered using different traffic patterns, i.e., multicast, unicast, and gather, preventing dataflow-specific communication backbones from benefiting the entire system if the dataflow changes or different dataflows run in the same system. Thus, in this paper, we propose MUG-NoC to support typical traffic patterns and accelerate them, therefore boosting multiple dataflows. Specifically, (i) we for the first time support multicast in 2D-mesh software configurable NoC by revising router configuration and proposing the efficient multicast routing; (ii) we decrease unicast latency by transmitting data through the different routes in parallel; (iii) we reduce output gather overheads by pipelining basic dataflow units. Experiments show that at least our proposed design can reduce 39.2% total data transmission time compared with the state-of-the-art CNN communication backbone.
Hui Chen 0016, Di Liu 0002, Shiqing Li, Shuo Huai, Weichen Liu 0001
ASP-DAC2
2023 Crossbar-Aligned & Integer-Only Neural Network Compression for Efficient in-Memory Acceleration
abstract
Crossbar-based In-Memory Computing (IMC) accelerators preload the entire Deep Neural Network (DNN) into crossbars before inference. However, devices with limited crossbars cannot infer increasingly complex models. IMC-pruning can reduce the usage of crossbars, but current methods need expensive extra hardware for data alignment. Meanwhile, quantization can represent weights of DNNs by integers, but they employ non-integer scaling factors to ensure accuracy, requiring costly multipliers. In this paper, we first propose crossbar-aligned pruning to reduce the usage of crossbars without hardware overhead. Then, we introduce a quantization scheme to avoid multipliers in IMC devices. Finally, we design a learning method to complete above two schemes and cultivate an optimal compact DNN with high accuracy and large sparsity during training. Experiments demonstrate that our framework, compared to state-of-the-art methods, achieves larger sparsity and lower power consumption with higher accuracy. We even improve the accuracy by 0.43% for VGG-16 with an 88.25% sparsity rate on the Cifar-10 dataset. Compared to the original model, we reduce computing power and area by 19.8x and 18.8x, respectively.
Shuo Huai, Di Liu 0002, Hui Chen 0016, Weichen Liu 0001, Ravi Subramaniam
ASP-DAC2
2023 Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional Pruning
abstract
In this paper, we propose TECO, a multi-dimensional pruning framework to collaboratively prune the three dimensions (depth, width, and resolution) of convolutional neural networks (CNNs) for better execution efficiency on embedded hardware. In TECO, we first introduce a two-stage importance evaluation framework, which efficiently and comprehensively evaluates each pruning unit according to both the local importance inside each dimension and the global importance across different dimensions. Based on the evaluation framework, we present a heuristic pruning algorithm to progressively prune the three dimensions of CNNs towards the optimal trade-off between accuracy and efficiency. Experiments on multiple benchmarks validate the advantages of TECO over existing state-of-the-art (SOTA) approaches. The code and pre-trained models are available anonymously at https://github.com/ntuliuteam/Teco.
Hao Kong 0001, Di Liu 0002, Shuo Huai, Ravi Subramaniam, Christian Makaya, Qian Lin 0001, Weichen Liu 0001
DAC2
2023 EMNAPE: Efficient Multi-Dimensional Neural Architecture Pruning for EdgeAI
abstract
In this paper, we propose a multi-dimensional pruning framework, EMNAPE, to jointly prune the three dimensions (depth, width, and resolution) of convolutional neural networks (CNNs) for better execution efficiency on embedded hardware. In EMNAPE, we introduce a two-stage evaluation strategy to evaluate the importance of each pruning unit and identify the computational redundancy in the three dimensions. Based on the evaluation strategy, we further present a heuristic pruning algorithm to progressively prune redundant units from the three dimensions for better accuracy and efficiency. Experiments demonstrate the superiority of EMNAPE over existing methods.
Hao Kong 0001, Shuo Huai, Di Liu 0002, Ravi Subramaniam, Christian Makaya, Qian Lin 0001, Weichen Liu 0001
DATE4
2023 Multi-Layer Seasonal Perception Network for Time Series Forecasting
abstract
Seasonal time series contain rich long-term dependencies. How to make good use of the seasonal information to predict the future is still a challenging problem. In this paper, we propose a neural network model called Multilayer Seasonal Perception Network (MSPNet) to predict seasonal time series. Firstly, we propose the idea of seasonal alignment, which converts univariate time series into multivariate time series, in order to capture seasonal features more effectively. Secondly, we extract the seasonal features and historical dependencies, using the Multi-layer Seasonal Perception Attention. Finally, we combine the obtained nonlinear features with linear features to conduct the final prediction. Experimental verification shows that the proposed MSPNet model is significantly superior to the baseline methods on multiple public datasets. The source code and datasets are available at https://github.com/MasterofEating/MSPNet
Ruoshu Wang, Shengfa Miao, Di Liu 0002, Xin Jin 0005, Weisheng Zhang
ICASSP3
2023 Energy-efficient computation offloading strategy with task priority in cloud assisted multi-access edge computing
Zhenli He, Di Liu 0002, Wei Zhou 0011, Keqin Li 0001
Future Gener. Comput. Syst.3
2023 Latency-constrained DNN architecture learning for edge systems using zerorized batch normalization
Shuo Huai, Di Liu 0002, Hao Kong 0001, Weichen Liu 0001, Ravi Subramaniam, Christian Makaya, Qian Lin 0001
Future Gener. Comput. Syst.2
2023 Improving robustness of convolutional neural networks using element-wise activation scaling
Zhi-Yuan Zhang, Zhenli He, Wei Zhou 0011, Di Liu 0002
Future Gener. Comput. Syst.5
2023 OCAP: On-device Class-Aware Pruning for personalized edge DNN models
Ye-Da Ma, Zhi-chao Zhao, Di Liu 0002, Zhenli He, Wei Zhou 0011
J. Syst. Archit.3
2023 SurgeNAS: A Comprehensive Surgery on Hardware-Aware Differentiable Neural Architecture Search
abstract
Differentiable neural architecture search (NAS) is an emerging paradigm to automate the design of top-performing convolutional neural networks (CNNs). Nonetheless, existing differentiable NAS methods suffer from several crucial weaknesses, such as inaccurate gradient estimation, high memory consumption, search fairness,etc. In this work, we introduce a novel hardware-aware differentiable NAS framework, namely SurgeNAS, in which we leverage the one-level optimization to avoid inaccuracy in gradient estimation. To this end, we propose an effective identity mapping regularization to alleviate the over-selecting issue. Besides, to mitigate the memory bottleneck, we propose an ordered differentiable sampling approach, which significantly reduces the search memory consumption to the single-path level, thereby allowing to directly search on target tasks instead of small proxy tasks. Meanwhile, it guarantees the strict search fairness. Moreover, we introduce a graph neural networks (GNNs) based predictor to approximate the on-device latency, which is further integrated into SurgeNAS to enable the latency-aware architecture search. Finally, we analyze the resource underutilization issue, in which we propose to scale up the searched SurgeNets withinComfort Zoneto balance the computation and memory access, which brings considerable accuracy improvement without deteriorating the execution efficiency. Extensive experiments are conducted on ImageNet with diverse hardware platforms, which clearly show the effectiveness of SurgeNAS in terms of accuracy, latency, and search efficiency.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Hui Chen 0016, Weichen Liu 0001
IEEE Trans. Computers2
2023 EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI
abstract
Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resource-constrained embedded devices. To address this issue, we propose EdgeCompress, a comprehensive compression framework to reduce the computational overhead of CNNs. In EdgeCompress, we first introduce dynamic image cropping (DIC), where we design a lightweight foreground predictor to accurately crop the most informative foreground object of input images for inference, which avoids redundant computation on background regions. Subsequently, we present compound shrinking (CS) to collaboratively compress the three dimensions (depth, width, and resolution) of CNNs according to their contribution to accuracy and model computation. DIC and CS together constitute a multidimensional CNN compression framework, which is able to comprehensively reduce the computational redundancy in both input images and neural network architectures, thereby improving the inference efficiency of CNNs. Further, we present a dynamic inference framework to efficiently process input images with different recognition difficulties, where we cascade multiple models with different complexities from our compression framework and dynamically adopt different models for different input images, which further compresses the computational redundancy and improves the inference efficiency of CNNs, facilitating the deployment of advanced CNNs onto embedded hardware. Experiments on ImageNet-1K demonstrate that EdgeCompress reduces the computation of ResNet-50 by 48.8% while improving the top-1 accuracy by 0.8%. Meanwhile, we improve the accuracy by 4.1% with similar computation compared to HRank. The state-of-the-art compression framework. The source code and models are available athttps://github.com/ntuliuteam/edge-compress.
Hao Kong 0001, Di Liu 0002, Shuo Huai, Ravi Subramaniam, Christian Makaya, Qian Lin 0001, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Efficient FPGA-Based Sparse Matrix-Vector Multiplication With Data Reuse-Aware Compression
abstract
Sparse matrix–vector multiplication (SpMV) on FPGAs has gained much attention. The performance of SpMV is mainly determined by the number of multiplications between nonzero matrix elements and the corresponding vector values per cycle. On the one side, the off-chip memory bandwidth limits the number of nonzero matrix elements transferred from the off-chip DDR to the FPGA chip per cycle. On the other side, the irregular vector access pattern poses challenges to fetch the corresponding vector values. Besides, the read-after-write (RAW) dependency in the accumulation process shall be solved to enable a fully pipelined design. In this work, we propose an efficient FPGA-based SpMV accelerator with data reuse-aware compression. The key observation is that repeated accesses to a vector value can be omitted by reusing the fetched data. Based on the observation, we propose a reordering algorithm to manually exploit the data reuse of fetched vector values. Further, we propose a novel compressed format called data reuse-aware compressed (DRC) to take full advantage of the data reuse and a fast format conversion algorithm to shorten the preprocessing time. Meanwhile, we propose an HLS-friendly accumulator to solve the RAW dependency. Finally, we implement and evaluate our proposed design on the Xilinx Zynq-UltraScale ZCU106 platform with a set of sparse matrices from the SuiteSparse matrix collection. Our proposed design achieves an average$1.18\times $performance speedup without the DRC format and an average$1.57\times $performance speedup with the DRC format w.r.t. the state-of-the-art work, respectively.
Shiqing Li, Di Liu 0002, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 LightNAS: On Lightweight and Scalable Neural Architecture Search for Embedded Platforms
abstract
Neural architecture search (NAS) is an emerging paradigm to automate the design of competitive deep neural networks (DNNs). In practice, DNNs are subject to strict latency constraints and any violation may lead to catastrophic consequences (e.g., autonomous vehicles). However, to obtain the architecture that strictly satisfies the required latency constraint, previous hardware-aware differentiable NAS methods have to repeat a plethora of search runs to tune relevant hyperparameters by trial and error, and as a result, the total design cost increases proportionally (empirically by ten times). To tackle this, we, in this article, introduce a lightweight and scalable hardware-aware NAS framework named LightNAS, which consists of two separate stages. In the first stage, we strive to search for the architecture that strictly satisfies the required latency constraint at the macro level in a differentiable manner, and more importantly, through a one-time search (i.e., you only search once). The architectures searched in the first stage are denoted as LightNets. After that, in the second stage, we introduce an efficient evolutionary scheme to further explore the micro-level channel configuration of each LightNet at low cost. To achieve this, we propose an effective yet computationally cheap proxy, namely, batchwise training estimation (BTE), as a plug-in complement to enable the channel-level exploration of LightNets on the fly such that the accuracy of LightNets can be improved without degrading the runtime latency on target hardware. Finally, extensive experiments are conducted on one popular embedded platform (i.e., Nvidia Jetson AGX Xavier) to demonstrate the efficacy of the proposed approach over previous state-of-the-art counterparts.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Hui Chen 0016, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 FAT: An In-Memory Accelerator With Fast Addition for Ternary Weight Neural Networks
abstract
Convolutional neural networks (CNNs) demonstrate excellent performance in various applications but have high computational complexity. Quantization is applied to reduce the latency and storage cost of CNNs. Among the quantization methods, binary and ternary weight networks (BWNs and TWNs) have a unique advantage over 8 and 4-bit quantization. They replace the multiplication operations in CNNs with additions, which are favored on in-memory-computing (IMC) devices. IMC acceleration for BWNs has been widely studied. However, though TWNs have higher accuracy and better sparsity than BWNs, IMC acceleration for TWNs has limited research. TWNs on the existing IMC devices are inefficient because the sparsity is not well utilized, and the addition operation is not efficient. In this article, we propose FAT as a novel IMC accelerator for TWNs. First, we propose a sparse addition control unit, which utilizes the sparsity of TWNs to skip the null operations on zero weights. Second, we propose a fast addition scheme based on the memory sense amplifier (SA) to avoid the time overhead of both carry propagation and writing back the carry to memory cells. Third, we further propose a combined-stationary data mapping to reduce the data movement of activations and weights and increase the parallelism across memory columns. Simulation results show that for addition operations at the SA level, FAT achieves$2.00\times $speedup,$1.22\times $power efficiency, and$1.22\times $area efficiency compared with a state-of-the-art IMC accelerator ParaPIM. FAT achieves$10.02\times $speedup and$12.19\times $energy efficiency compared with ParaPIM on networks with 80% average sparsity.
Shien Zhu, Luan H. K. Duong, Hui Chen 0016, Di Liu 0002, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 HACScale: Hardware-Aware Compound Scaling for Resource-Efficient DNNs
abstract
Model scaling is an effective way to improve the accuracy of deep neural networks (DNNs) by increasing the model capacity. However, existing approaches seldom consider the underlying hardware, causing inefficient utilization of hardware resources and consequently high inference latency. In this paper, we propose HACScale, a hardware-aware model scaling strategy to fully exploit hardware resources for higher accuracy. In HACScale, different dimensions of DNNs are jointly scaled with consideration of their contributions to hardware utilization and accuracy. To improve the efficiency of width scaling, we introduce importance-aware width scaling in HACScale, which computes the importance of each layer to the accuracy and scales each layer accordingly to optimize the trade-off between accuracy and model parameters. Experiments show that HACScale improves the hardware utilization by 1.92× on ImageNet, as a result, it achieves 2.41% accuracy improvement with a negligible latency increase of 0.6%. On CIFAR-10, HACScale improves the accuracy by 2.23% with only 6.5% latency growth.
Hao Kong 0001, Di Liu 0002, Weichen Liu 0001, Ravi Subramaniam
ASP-DAC2
2022 Efficient On-Device Incremental Learning by Weight Freezing
abstract
On-device learning has become a new trend for edge intelligence systems. In this paper, we investigate the on-device in-cremental learning problem, which targets to learn new classes on top of a well-trained model on the device. Incremental learning is known to suffer from catastrophic forgetting, i.e., a model learns new classes at the cost of forgetting the old classes. Inspired by model pruning techniques, we propose a new on-device incremental learning method based on weight freezing. The weight freezing in our framework plays two roles: 1) preserving the knowledge of the old classes; 2) boosting the training procedure. By means of weight freezing, we build up an efficient incremental learning framework which combines knowledge distillation to fine-tune the new model. We conduct extensive experiments on CIFAR100 and compare our method with two existing methods. The experimental results show that our method can achieve higher accuracy after incrementally learning new classes.
Ze-Han Wang, Zhenli He, Yi-Xiong Huang, Zhi-Yuan Zhang, Di Liu 0002
ASP-DAC8
2022 Work-in-Progress: What to Expect of Early Training Statistics? An Investigation on Hardware-Aware Neural Architecture Search
abstract
Neural architecture search (NAS) is an emerging paradigm to automate the design of top-performing deep neural networks (DNNs). Specifically, the increasing success of NAS is attributed to the reliable performance estimation of different architectures. Despite significant progress to date, previous relevant methods suffer from prohibitive computational overheads. To avoid this, we propose an effective yet computationally efficient proxy, namely Trained Batchwise Estimation (TBE), to reliably estimate the performance of different architectures using the early batchwise training statistics. We then integrate TBE into the hardware-aware NAS scenario to search for hardware-efficient architecture solutions. Experimental results clearly show the superiority of TBE over previous relevant state-of-the-art approaches.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Hui Chen 0016, Weichen Liu 0001
CODES+ISSS2
2022 You only search once: on lightweight differentiable architecture search for resource-constrained embedded platforms
abstract
Benefiting from the search efficiency, differentiable neural architecture search (NAS) has evolved as the most dominant alternative to automatically design competitive deep neural networks (DNNs). We note that DNNs must be executed under strictly hard performance constraints in real-world scenarios, for example, the runtime latency on autonomous vehicles. However, to obtain the architecture that meets the given performance constraint, previous hardware-aware differentiable NAS methods have to repeat a plethora of search runs to manually tune the hyper-parameters by trial and error, and thus the total design cost increases proportionally. To resolve this, we introduce a lightweight hardware-aware differentiable NAS framework dubbed LightNAS, striving to find the required architecture that satisfies various performance constraints through a one-time search (i.e., you only search once). Extensive experiments are conducted to show the superiority of LightNAS over previous state-of-the-art methods. Related codes will be released at https://github.com/stepbuystep/LightNAS.
Di Liu 0002, Hao Kong 0001, Shuo Huai, Hui Chen 0016, Weichen Liu 0001
DAC2
2022 Once For All Skip: Efficient Adaptive Deep Neural Networks
abstract
In this paper, we propose a new module, namely once for all skip (OFAS), for adaptive deep neural networks to efficiently control the block skip within a DNN model. The novelty of OFAS is that it only needs to compute once for all skippable blocks to determine their execution states. Moreover, since adaptive DNN models with OFAS cannot achieve the best accuracy and efficiency in end-to-end training, we propose a reinforcement learning-based training method to enhance the training procedure. The experimental results with different models and datasets demonstrate the effectiveness and efficiency in comparison to the state of the arts. The code is available at https://github.com/ieslab-ynu/OFAS.
Di Liu 0002, Yi-Xiong Huang, Zhi-Yuan Zhang
DATE2
2022 Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded Hardware
abstract
Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI. However, as an image usually contains much spatial redundancy, e.g., background pixels, directly shrinking the whole image will lose important features of the foreground object and lead to severe accuracy degradation. In this paper, we propose a dynamic image cropping framework to reduce the spatial redundancy by accurately cropping the foreground object from images. To achieve the instance-aware fine cropping, we introduce a lightweight foreground predictor to efficiently localize and crop the foreground of an image. The finely cropped images can be correctly recognized even at a small resolution. Meanwhile, computational redundancy also exists in CNN architectures. To pursue higher execution efficiency on resource-constrained embedded devices, we also propose a compound shrinking strategy to coordinately compress the three dimensions (depth, width, resolution) of CNNs. Eventually, we seamlessly combine the proposed dynamic image cropping and compound shrinking into a unified compression framework, Smart Scissor, which is expected to significantly reduce the computational overhead of CNNs while still maintaining high accuracy. Experiments on ImageNet-1K demonstrate that our method reduces the computational cost of ResNet50 by 41.5% while improving the top-1 accuracy by 0.3%. Moreover, compared to HRank, the state-of-the-art CNN compression framework, our method achieves 4.1% higher top-1 accuracy at the same computational cost. The codes and data are available at https://github.com/ntuliuteam/smart-scissor
Hao Kong 0001, Di Liu 0002, Shuo Huai, Weichen Liu 0001, Ravi Subramaniam, Christian Makaya, Qian Lin 0001
ICCAD2
2022 Collate: Collaborative Neural Network Learning for Latency-Critical Edge Systems
abstract
Federated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while protecting data privacy. However, when deploying FL in real-time edge systems, the heterogeneity of devices among systems has a severe impact on the performance of the inferred model. Existing optimizations on FL focus on improving the training efficiency but fail to speed up inference, especially when there is a latency constraint. In this work, we propose Collate, a novel training framework that collaboratively learns heterogeneous models to meet the latency constraints of multiple edge systems simultaneously. We design a dynamic zeroizing-recovering method to adjust each local model architecture for high accuracy under its latency constraint. A proto-corrected federated aggregation scheme is also introduced to aggregate all heterogeneous local models, satisfying the latency constraint of different systems with only one training process and maintaining high accuracy. Extensive experiments indicate that, compared to state-of-the-art methods and under a latency constraint, our extended models can improve the accuracy by 1.96% on average, and our shrunk models can also obtain a 3.09% accuracy improvement on average, with almost no extra training overhead. The related codes and data will be available at https://github.com/ntuliuteam/Collate.
Shuo Huai, Di Liu 0002, Hao Kong 0001, Weichen Liu 0001, Ravi Subramaniam, Christian Makaya, Qian Lin 0001
ICCD2
2022 Bringing AI to edge: From deep learning's perspective
Di Liu 0002, Hao Kong 0001, Weichen Liu 0001, Ravi Subramaniam
Neurocomputing1
2022 CARTAD: Compiler-Assisted Reinforcement Learning for Thermal-Aware Task Scheduling and DVFS on Multicores
abstract
As the power density of modern CPUs is gradually increasing, thermal management has become one of the primary concerns for multicore systems, where task scheduling and dynamic voltage/frequency scaling (DVFS) play a pivotal role in effectively managing the system temperature. In this article, we proposeCARTAD, a new reinforcement learning (RL)-based task scheduling and DVFS method for temperature minimization and latency guarantee on multicore systems. The novelty ofCARTADframework is that we exploit the machine learning technique to analyze the applications’ intermediate representations (IRs) generated by a compiler and identify an important feature which is critical for predicting the application’s performance. With the newly explored feature, we construct an RL-based scheduler with the more effective state representation and reward function such that the system temperature can be minimized while guaranteeing applications’ latency. We implement and evaluateCARTADon real platforms in comparison with the state-of-the-art approaches. Experimental results showCARTADcan reduce the maximum temperature by up to 16 °C and the average temperature by up to 10 °C.
Di Liu 0002, Shi-Gui Yang, Zhenli He, Mingxiong Zhao 0001, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Designing Efficient DNNs via Hardware-Aware Neural Architecture Search and Beyond
abstract
Hardware systems integrated with deep neural networks (DNNs) are deemed to pave the way for future artificial intelligence (AI). However, manually designing efficient DNNs involves nontrivial computation resources since significant trial-and-errors are required to finalize the network configuration. To this end, we, in this article, introduce a novel hardware-aware neural architecture search (NAS) framework, namely, GoldenNAS, to automate the design of efficient DNNs. To begin with, we present a novel technique, called dynamic channel scaling, to enable the channel-level search since the number of channels has non-negligible impacts on both accuracy and efficiency. Besides, we introduce an efficient progressive space shrinking method to raise the awareness of the search space toward target hardware and alleviate the search overheads as well. Moreover, we propose an effective hardware performance modeling method to approximate the runtime latency of DNNs upon target hardware, which is further integrated into GoldenNAS to avoid the tedious on-device measurements. Then, we employ the evolutionary algorithm (EA) to search for the optimal operator/channel configurations of DNNs, denoted as GoldenNets. Finally, to enable the depthwise adaptiveness of GoldenNets under dynamic environments, we propose the adaptive batch normalization (ABN) technique, followed by the self-knowledge distillation (SKD) approach to improve the accuracy of adaptive subnetworks. We conduct extensive experiments directly on ImageNet, which clearly demonstrate the advantages of GoldenNAS over existing state-of-the-art approaches.
Di Liu 0002, Shuo Huai, Hao Kong 0001, Hui Chen 0016, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Toward the Predictability of Dynamic Real-Time DNN Inference
abstract
Deep neural networks (DNNs) have been widely used in many cyber–physical systems (CPSs). However, it is still a challenging work to deploy DNNs in real-time systems. In particular, the execution time of DNN inference must be predictable, s.t. it could be known whether the runtime inference can complete within a required timing constraint. Moreover, the timing constraints may change dynamically with the runtime environment in many embedded applications, such as autonomous cars. A possible way to meet such dynamic real-time requirements is to execute different subnetworks of a DNN at runtime. However, improper construction of subnetworks may not only introduce unpredictable inference time, s.t. the real-timing constraints could be violated unexpectedly, but also has poor compatibility with the well-optimized machine learning framework (e.g., TensorFlow). In this article, we study the predictability when executing different subnetworks of a DNN. In particular, we present a featurewise runtime adaptation framework for DNN inference, which is implemented and validated on NVIDIA Jetson TX2 and Nano with TensorFlow. The experimental results show that our method can achieve predictable inference time in comparison with the state-of-the-art methods.
Weiguang Pang, Xu Jiang 0004, Mingsong Lv, Teng Gao, Di Liu 0002, Wang Yi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 Layerwise Security Protection for Deep Neural Networks in Industrial Cyber Physical Systems
abstract
Although deep neural networks (DNNs) have been increasingly applied in industrial cyber physical systems (ICPSs), they are vulnerable to security attacks due to the tight interaction between cyber elements and physical elements. In this article, we aim to protect the core IP of DNNs, i.e., the model weights, against security attacks. Different from conventional approaches, a layerwise protection framework is proposed to ensure the confidentiality of DNN model weights during the inference procedure such that the security quality is maximized, while satisfying the latency constraint of the DNN task. Based on the layerwise execution characteristics of DNN tasks, the encrypted layer-related weights are decrypted and fed to the next layer of DNN in plaintext. CPU-field programmable gate array (FPGA) coscheduling is considered to accelerate the execution of confidentiality protection, where CPU is utilized to conduct the decryption of weights and FPGA is used to perform the layer execution of DNN. Considering to provide optimal confidential protection for each layer, the problem is transformed into a quality of security maximization problem subject to layerwise execution constraint and deadline constraint of the DNN application. Due to the problem being NP-hard, a fast approximation algorithm is proposed to obtain the near-optimal solution under given real-time and security constraints. Extensive experiments and a real-life ICPS application evaluate the efficiency of the proposed techniques.
Wei Jiang 0016, Jinyu Zhan, Di Liu 0002, Jiafu Wan
IEEE Trans. Ind. Informatics4
2021 ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge Systems
abstract
Edge devices have been widely adopted to bring deep learning applications onto low power embedded systems, mitigating the privacy and latency issues of accessing cloud servers. The increasingly computational demand of complex neural network models leads to large latency on edge devices with limited resources. Many application scenarios are real-time and have a strict latency constraint, while conventional neural network compression methods are not latency-oriented. In this work, we propose a novel compact neural networks training method to reduce the model latency on latency-critical edge systems. A latency predictor is also introduced to guide and optimize this procedure. Coupled with the latency predictor, our method can guarantee the latency for a compact model by only one training process. The experiment results show that, compared to state-of-the-art model compression methods, our approach can well-fit the ‘hard’ latency constraint by significantly reducing the latency with a mild accuracy drop. To satisfy a 34ms latency constraint, we compact ResNet-50 with 0.82% of accuracy drop. And for GoogLeNet, we can even increase the accuracy by 0.3%
Shuo Huai, Lei Zhang 0072, Di Liu 0002, Weichen Liu 0001, Ravi Subramaniam
DAC3
2021 HSCoNAS: Hardware-Software Co-Design of Efficient DNNs via Neural Architecture Search
abstract
In this paper, we present a novel multi-objective hardware-aware neural architecture search (NAS) framework, namely HSCoNAS, to automate the design of deep neural networks (DNNs) with high accuracy but low latency upon target hardware. To accomplish this goal, we first propose an effective hardware performance modeling method to approximate the runtime latency of DNNs on target hardware, which will be integrated into HSCoNAS to avoid the tedious on-device measurements. Besides, we propose two novel techniques, i.e., dynamic channel scaling to maximize the accuracy under the specified latency and progressive space shrinking to refine the search space towards target hardware as well as alleviate the search overheads. These two techniques jointly work to allow HSCoNAS to perform fine-grained and efficient explorations. Finally, an evolutionary algorithm (EA) is incorporated to conduct the architecture search. Extensive experiments on ImageNet are conducted upon diverse target hardware, i.e., GPU, CPU, and edge device to demonstrate the superiority of HSCoNAS over recent state-of-the-art approaches.
Di Liu 0002, Shuo Huai, Weichen Liu 0001
DATE2
2021 Optimized Data Reuse via Reordering for Sparse Matrix-Vector Multiplication on FPGAs
abstract
Sparse matrix-vector multiplication (SpMV) is of paramount importance in both scientific and engineering applications. The main workload of SpMV is multiplications between randomly distributed nonzero elements in sparse matrices and their corresponding vector elements. Due to irregular data access patterns of vector elements and the limited memory bandwidth, the computational throughput of CPUs and GPUs is lower than the peak performance offered by FPGAs. FPGA's large on-chip memory allows the input vector to be buffered on-chip and hence the off-chip memory bandwidth is only utilized to transfer the nonzero elements' values, column indices, and row indices. Multiple nonzero elements are transmitted to FPGA and then their corresponding vector elements are accessed per cycle. However, typical on-chip block RAMs (BRAM) in FPGAs only have two access ports. The mismatch between off-chip memory bandwidth and on-chip memory ports stalls the whole engine, resulting in inefficient utilization of off-chip memory bandwidth. In this work, we reorder the nonzero elements to optimize data reuse for SpMV on FPGAs. The key observation is that since the vector elements can be reused for nonzero elements with the same column index, memory requests of these elements can be omitted by reusing the fetched data. Based on this observation, a novel compressed format is proposed to optimize data reuse by reordering the matrix's nonzero elements. Further, to support the compressed format, we design a scalable hardware accelerator and implement it on the Xilinx UltraScale ZCU106 platform. We evaluate the proposed design with a set of matrices from the University of Florida sparse matrix collection. The experimental results show that the proposed design achieves an average 1.22x performance speedup w.r.t. the state-of-the-art work.
Shiqing Li, Di Liu 0002, Weichen Liu 0001
ICCAD2
2021 Reaching consensus in decentralized coordination of distributed microservices
Shuiguang Deng, Di Liu 0002, Zeming Yan
Comput. Networks3
2020 EdgeNAS: Discovering Efficient Neural Architectures for Edge Systems
abstract
Edge systems integrated with deep neural networks (DNNs) are deemed to pave the way for future artificial intelligence (AI). However, designing accurate and efficient DNNs for resource-limited edge systems is challenging as well as requires a huge amount of engineering efforts from human experts since the design space is highly complex and diverse. Also, previous works mostly focus on designing DNNs with less floating-point operations (FLOPs), but indirect FLOPs count does not necessarily reflect the complexity of DNNs. To tackle these, we, in this paper, propose a novel neural architecture search (NAS) approach, namely EdgeNAS, to automatically discover efficient DNNs for less capable edge systems. To this end, we propose an end-to-end learning-based latency estimator, which is able to directly approximate the architecture latency on edge systems while incurring negligible computational overheads. Further, we effectively incorporate the latency estimator into EdgeNAS with a uniform sampling strategy, which guides the architecture search towards an edge-efficient direction. Moreover, a search space regularization approach is introduced to balance the trade-off between efficiency and accuracy. We evaluate EdgeNAS on the edge platform, Nvidia Jetson Xavier, with three popular datasets. Experimental results demonstrate the superiority of EdgeNAS over state-of-the-art approaches in terms of latency, accuracy, number of parameters, and the search cost.
Di Liu 0002, Hao Kong 0001, Weichen Liu 0001
ICCD2
2020 Joint Offloading and Resource Allocation for Time-Sensitive Multi-Access Edge Computing Network
abstract
In this paper, we investigate offloading scheme and resource allocation strategy for Orthogonal Frequency-Division Multiple Access (OFDMA) based multi-access edge computing (MEC) network to minimize the total system energy consumption. Partial data offloading is studied where mobile date can be computed at both local devices and the edge cloud with the consideration of time-sensitive tasks for users. Since the NP-hardness of the considered optimization problem, we propose an iterative algorithm to decide the proportion of data to offload and design the resource allocation strategy in a sequence. Simulation results show that the proposed algorithm achieves better performance than the reference schemes.
Jun-Jie Yu, Mingxiong Zhao 0001, Di Liu 0002, Shaowen Yao 0001, Wei Feng 0014
WCNC4
2020 Energy-Efficient Parallel Real-Time Scheduling on Clustered Multi-Core
abstract
Energy-efficiency is a critical requirement for computation-intensive real-time applications on multi-core embedded systems. Multi-core processors enable intra-task parallelism, and in this work, we study energy-efficient real-time scheduling of constrained deadline sporadic parallel tasks, where each task is represented as a directed acyclic graph (DAG). We consider a clustered multi-core platform where processors within the same cluster run at the same speed at any given time. A new concept named speed-profile is proposed to model per-task and per-cluster energy-consumption variations during run-time to minimize the expected long-term energy consumption. To our knowledge, no existing work considers energy-aware real-time scheduling of DAG tasks with constrained deadlines, nor on a clustered multi-core platform. The proposed energy-aware real-time scheduler is implemented upon an ODROID XU-3 board to evaluate and demonstrate its feasibility and practicality. To complement our system experiments in large-scale, we have also conducted simulations that demonstrate a CPU energy saving of up to 67 percent through our proposed approach compared to existing methods.
Ashikahmed Bhuiyan, Di Liu 0002, Aamir Khan, Abusayeed Saifullah, Nan Guan, Zhishan Guo
IEEE Trans. Parallel Distributed Syst.2
2019 Analyzing GEDF Scheduling for Parallel Real-Time Tasks with Arbitrary Deadlines
abstract
Real-time and embedded systems are shifting from single-core to multi-core processors, on which software must be parallelized to fully utilize the computation capacity of hardware. Recently much work has been done on real-time scheduling of parallel tasks modeled as directed acyclic graphs (DAG). However, most of these studies assume tasks to have implicit or constrained deadlines. Much less work considered the general case of arbitrary deadlines (i.e., the relative deadline is allowed to be larger than the period), which is more difficult to analyze due to intra-task interference among jobs. In this paper, we study the analysis of Global Earliest Deadline First (GEDF) scheduling for DAG parallel tasks with arbitrary deadlines. We develop new analysis techniques for GEDF scheduling of a single DAG task, which not only outperform the state-of-the-art in general evidenced by empirical evaluation, but also guarantee a better capacity augmentation bound 2.41 (the best known result is 2.5). The proposed analysis techniques are also extended to and evaluated with the case of multiple DAG tasks using the federated scheduling approach.
Xu Jiang 0004, Nan Guan, Di Liu 0002, Weichen Liu 0001
DATE3
2019 Energy-Efficient Real-Time Scheduling of DAGs on Clustered Multi-Core Platforms
abstract
With the growth of computation-intensive real-time applications on multi-core embedded systems, energy-efficient real-time scheduling becomes crucial. Multi-core processors enable intra-task parallelism, and there has been much progress on exploiting that, while there has been only a little progress on energy-efficient multi-core real-time scheduling as yet. In this work, we study energy-efficient real-time scheduling of constrained deadline sporadic parallel tasks, where each task is represented as a directed acyclic graph (DAG). We consider a clustered multi-core platform where processors within the same cluster run at the same speed at any given time. A new concept named speed-profile is proposed to model per-task and per-cluster energy-consumption variations during run-time to minimize the expected long-term energy consumption. To our knowledge, no existing work considers energy-aware real-time scheduling of DAG tasks with constrained deadlines, nor on a clustered multi-core platform. The proposed energy-aware realtime scheduler is implemented upon an ODROID XU-3 board to evaluate and demonstrate its feasibility and practicality. To complement our system experiments in large-scale, we have also conducted simulations that demonstrate a CPU energy saving of up to 57% through our proposed approach compared to existing methods.
Zhishan Guo, Ashikahmed Bhuiyan, Di Liu 0002, Aamir Khan, Abusayeed Saifullah, Nan Guan
RTAS3
2019 CASS: Criticality-Aware Standby-Sparing for real-time systems
Mingxiong Zhao 0001, Di Liu 0002, Xu Jiang 0004, Weichen Liu 0001, Cheng Xie 0001, Yun Yang 0003, Zhishan Guo
J. Syst. Archit.2
2019 A process partitioning technique for constructing decentralized web service compositions
abstract
Summary Web service compositions have been widely applied in different applications. A service composition is usually implemented in either a centralized or decentralized manner. Compared with the centralized service composition, the decentralized composition has no central control component, and components interact with each other directly, thereby achieving better performance. Process partitioning is a technique to divide a process into multiple parts and has been shown that it can be successfully applied to decentralizing process‐driven service compositions. This paper proposes a new process partitioning technique for constructing decentralized service compositions. The proposed technique, which is based on typed digraphs and a graph transformation technique, is used for exploring available process partitioning solutions. For applications, this paper discusses the topology and interaction features about the partitioning solutions and summarizes a ranking method for them. Three experiments are conducted to evaluate the proposed methods in this paper. Experimental results show that the proposed methods can be applied in constructing decentralized service compositions effectively. In addition, the results also show that the decentralized compositions can have lower average response times and higher throughputs than the corresponding centralized compositions in the experiments.
Di Liu 0002, Junsong Liu, Shaowen Yao 0001
Softw. Pract. Exp.2
2019 Real-Time Scheduling of DAG Tasks with Arbitrary Deadlines
abstract
Real-time and embedded systems are shifting from single-core to multi-core processors, on which the software must be parallelized to fully utilize the computation capacity of the hardware. Recently, much work has been done on real-time scheduling of parallel tasks modeled as directed acyclic graphs (DAG). However, most of these studies assume tasks to have implicit or constrained deadlines. Much less work considered the general case of arbitrary deadlines (i.e., the relative deadline is allowed to be larger than the period), which is more difficult to analyze due to intra-task interference among jobs. In this article, we study the analysis of Global Earliest Deadline First (GEDF) scheduling for DAG parallel tasks with arbitrary deadlines. We develop new analysis techniques for GEDF scheduling of a single DAG task and this new analysis techniques can guarantee a better capacity augmentation bound 2.41 (the best known result is 2.5) in the case of a single task. Furthermore, the proposed analysis techniques are also extended to the case of multiple DAG tasks under GEDF and federated scheduling. Finally, through empirical evaluation, we justify the out-performance of our schedulability tests compared to the state-of-the-art in general.
Kankan Wang, Xu Jiang 0004, Nan Guan, Di Liu 0002, Weichen Liu 0001, Qingxu Deng
ACM Trans. Design Autom. Electr. Syst.4
2018 Brain Tumor Segmention Based on Dilated Convolution Refine Networks
abstract
A brain tumor is a growth of abnormal cells in the tissues of the brain, which is difficult for treatment and severely affects patients' cognitive ability. Recent year magnetic resonance imaging (MRI) has been widely used imaging technique to assess brain tumors. However manual segmentation and artificial extracting features block MRI's practice when facing with the huge amount of data produced by MRI. An efficient and automatic image segmentation of brain tumor is still needed. In this paper, a novel automatic segmentation framework of brain tumors, which have 5 parts and resnet-50 use as a backbone, is proposed based on convolutional neural network. A dilated convolution refine (DCR) structure is introduced to extract the local features and global features. After investigating different parameters of our framework, it is proved that DCR is an efficient and robust method in Brain Tumor Segmentation. The experiments are evaluated by Multimodal Brain Tumor Image Segmentation (BRATS 2015) dataset. The results show that our framework in complete tumor segmentation achieved excellent results with a DEC score of 0.87 and a PPV score of 0.92. (GitHub: https://github.com/wei-lab/DCR)
Di Liu 0002, Xiaojuan Yu, Shaowen Yao 0001, Wei Zhou 0011
SERA1
2018 Utilization-Based Scheduling of Flexible Mixed-Criticality Real-Time Tasks
abstract
Mixed-criticality models are an emerging paradigm for the design of real-time systems because of their significantly improved resource efficiency. However, formal mixed-criticality models have traditionally been characterized by two impractical assumptions: once any high-criticality task overruns, all low-criticality tasks are suspended and all other high-criticality tasks are assumed to exhibit high-criticality behaviors at the same time. In this paper, we propose a more realistic mixed-criticality model, called the flexible mixed-criticality (FMC) model, in which these two issues are addressed in a combined manner. In this new model, only the overrun task itself is assumed to exhibit high-criticality behavior, while other high-criticality tasks remain in the same mode as before. The guaranteed service levels of low-criticality tasks are gracefully degraded with the overruns of high-criticality tasks. We derive a utilization-based technique to analyze the schedulability of this new mixed-criticality model under EDF-VD scheduling. During run time, the proposed test condition serves an important criterion for dynamic service level tuning, by means of which the maximum available execution budget for low-criticality tasks can be directly determined with minimal overhead while guaranteeing mixed-criticality schedulability. Experiments demonstrate the effectiveness of the FMC scheme compared with state-of-the-art techniques.
Gang Chen 0023, Nan Guan, Di Liu 0002, Qingqiang He, Kai Huang 0001, Todor P. Stefanov, Wang Yi 0001
IEEE Trans. Computers3
2018 Scheduling Analysis of Imprecise Mixed-Criticality Real-Time Tasks
abstract
In this paper, we study the scheduling problem of the imprecise mixed-criticality model (IMC) under earliest deadline first with virtual deadline (EDF-VD) scheduling upon uniprocessor systems. Two schedulability tests are presented. The first test is a concise utilization-based test which can be applied to the implicit deadline IMC task set. The suboptimality of the proposed utilization-based test is evaluated via a widely-used scheduling metric, speedup factors. The second test is a more effective test but with higher complexity which is based on the concept of demand bound function (DBF). The proposed DBF-based test is more generic and can apply to constrained deadline IMC task set. Moreover, in order to address the high time cost of the existing deadline tuning algorithm, we propose a novel algorithm which significantly improve the efficiency of the deadline tuning procedure. Experimental results show the effectiveness of our proposed schedulability tests, confirm the theoretical suboptimality results with respect to speedup factor, and demonstrate the efficiency of our proposed algorithm over the existing deadline tunning algorithm. In addition, issues related to the implementation of the IMC model under EDF-VD are discussed.
Di Liu 0002, Nan Guan, Jelena Spasic, Gang Chen 0023, Songran Liu, Todor P. Stefanov, Wang Yi 0001
IEEE Trans. Computers1
2016 Exploiting resource-constrained parallelism in hard real-time streaming applications
Jelena Spasic, Di Liu 0002, Todor P. Stefanov
DATE2
2016 Energy-Efficient Scheduling of Real-Time Tasks on Heterogeneous Multicores Using Task Splitting
abstract
In this paper, we investigate the problem of using the state-of-the-art C=D task-splitting approach to energy efficiently schedule real-time tasks on a single-ISA heterogeneous multicore system. We first extend the existing task-splitting approach for heterogeneous multicore systems. Based on our extension, we propose an algorithm, called ASHM, to allocate and split realtime tasks on a heterogeneous multicore system. The experimental results demonstrate the effectiveness of our proposed ASHM algorithm compared to existing allocation approaches in terms of energy savings.
Di Liu 0002, Jelena Spasic, Peng Wang 0036, Todor P. Stefanov
RTCSA1
2016 EDF-VD Scheduling of Mixed-Criticality Systems with Degraded Quality Guarantees
abstract
This paper studies real-time scheduling of mixed-criticality systems where low-criticality tasks are still guaranteed some service in the high-criticality mode, with reduced execution budgets. First, we present a utilization-based schedulability test for such systems under EDF-VD scheduling. Second, we quantify the suboptimality of EDF-VD (with our test condition) in terms of speedup factors. In general, the speedup factor is a function with respect to the ratio between the amount of resource required by different types of tasks in different criticality modes, and reaches 4/3 in the worst case. Furthermore, we show that the proposed utilization-based schedulability test and speedup factor results apply to the elastic mixed-criticality model as well. Experiments show effectiveness of our proposed method and confirm the theoretical suboptimality results.
Di Liu 0002, Jelena Spasic, Nan Guan, Gang Chen 0023, Songran Liu, Todor P. Stefanov, Wang Yi 0001
RTSS1
2016 On the Improved Hard Real-Time Scheduling of Cyclo-Static Dataflow
abstract
Recently, it has been shown that the hard real-time scheduling theory can be applied to streaming applications modeled as acyclic Cyclo-Static Dataflow (CSDF) graphs. However, this recent approach is not always efficient in terms of throughput and processor utilization. Therefore, in this article, we propose an improved hard real-time scheduling approach to schedule streaming applications modeled as acyclic CSDF graphs on a Multiprocessor System-on-Chip (MPSoC) platform. The proposed approach converts each actor in a CSDF graph to a set of real-time periodic tasks. The conversion enables application of many hard real-time scheduling algorithms that offer fast calculation of the required number of processors for scheduling the tasks. In addition, we propose a method to reduce the graph latency when the converted tasks are scheduled as real-time periodic tasks. We evaluate the performance and time complexity of our approach in comparison to several existing scheduling approaches. Experiments on a set of real-life streaming applications demonstrate that our approach (1) results in systems with higher throughput and better processor utilization in comparison to the existing hard real-time scheduling approach for CSDF graphs, while requiring comparable time for the system derivation; (2) delivers shorter application latency by applying the proposed method for graph latency reduction while providing better throughput and processor utilization when compared to the existing hard real-time scheduling approach; (3) gives the same throughput as the existing periodic scheduling approach for CSDF graphs, but requires much shorter time to derive the task schedule and tasks’ parameters (periods, start times, and so on); and (4) gives the throughput that is equal to or very close to the maximum achievable throughput of an application obtained via self-timed scheduling, but requires much shorter time to derive the schedule. The total time needed for the proposed conversion approach and the calculation of the minimum number of processors needed to schedule the tasks and the calculation of the size of communication buffers between tasks is in the range of seconds.
Jelena Spasic, Di Liu 0002, Emanuele Cannella, Todor P. Stefanov
ACM Trans. Embed. Comput. Syst.2
2014 Resource optimization for CSDF-modeled streaming applications with latency constraints
abstract
In this paper, we study the problem of minimizing the number of processors required for scheduling latency-constrained streaming applications modeled as CSDF graphs, where the actors of a CSDF are executed as strictly periodic tasks. We formalize the problem and prove that due to the strict periodicity of actors the problem is an integer convex programming problem, that can be solved efficiently by using an existing convex programming solver. We evaluate our solution approach on a set of 13 real-life streaming applications modeled as CSDF graphs and demonstrate that it can reduce the number of processors in more than 52% of the conducted experiments in comparison to an existing approach.
Di Liu 0002, Jelena Spasic, Jiali Teddy Zhai, Todor P. Stefanov, Gang Chen 0023
DATE1
2014 Abstract: Shared L2 Cache Management in Multicore Real-Time System
abstract
In multicore system, shared cache interference has been recognized as one of the major factors that degrade the average performance as well as predictability of system. How to manage the shared cache in order to optimize the system performance while guaranteeing the system predictability is still an open issue. State-of-the-art techniques on this topic use page coloring to partition the shared cache at OS level. In this paper, we present a shared cache management scheme for multicore system. This shared cache management scheme supports way-based cache partitioning at hardware level, building task-level time-triggered reconfigurable-cache multicore system. We evaluated the proposed scheme w.r.t. different numbers of cores and cache modules and prototyped the constructed MPSoCs on FPGA.
Gang Chen 0023, Biao Hu 0001, Kai Huang 0001, Alois C. Knoll, Di Liu 0002
FCCM5