VLDB 2026 Research / reviewers in the wild / expert
Haodong Lu 0001
dblp:252/1159-1
· DBLP profile ↗
20ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-8628-2664ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-author · 14 since 2021Computer networks · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PipeViT: Accelerating Vision Transformers via Intra-Layer PipeliningabstractVision Transformers (ViTs) have achieved high performance across various computer vision tasks by leveraging the attention mechanism. However, the attention module in ViTs severely hindered inference performance due to its low operational intensity. Existing approaches improve ViTs efficiency through pruning, sparsity, and linearization, but at the cost of fine-tuning overhead and accuracy degradation. In this paper, we propose PipeViT, a memory-efficient and low-latency accelerator for ViTs inference. The key insight of PipeViT is to exploit intra-layer acceleration opportunities. Specifically, we first fuse the attention operations into a single operator to reduce memory access overhead. Then, we divide the input of attention into multiple tiles to reduce the on-chip memory requirement. Finally, we pipeline the tiled attention computation to improve overall throughput. Based on the optimized dataflow, we design a heterogeneous dual-core architecture for efficient pipeline execution. Furthermore, to maximize hardware utilization, the architecture can be reconfigured into a single core with higher parallelism during the execution of the feed-forward network. Experimental results show that PipeViT achieves up to $19.3 \times 1.5 \times, 2.1 \times$, and $2.0 \times$ improvements in Frames Per Second (FPS) compared to state-of-the-art accelerators, including ViTA, Auto-ViT, MEViT, and HeatViT. Additionally, PipeViT achieves up to $8.0 \times$ and $2.6 \times$ higher energy efficiency compared to CPU and GPU implementations, respectively. Xilang Zhou, Yiheng Xu, Haodong Lu 0001, Jun Yu 0010, Kun Wang 0005 |
ASP-DAC | 3 |
| 2026 | RouterAcc: FPGA Acceleration for VLSI Detailed Router via Hierarchical Storage MappingabstractDetailed routing constitutes a critical phase in the very large-scale integration (VLSI) physical design, widely regarded as the most time-consuming and computationally intensive step in the back-end design process. Due to its iterative nature and strong data dependencies, conventional parallel acceleration techniques often suffer from limited scalability and effectiveness. To address these challenges, we propose RouterAcc, an FPGA-based software–hardware co-design acceleration framework tailored for VLSI detailed routing. RouterAcc incorporates an access analysis mechanism and a termination condition strategy to accelerate convergence. Furthermore, we employ a hierarchical storage mapping scheme and a flexible dimension-partitioning architecture to alleviate memory bottlenecks and enhance data locality. Additionally, RouterAcc leverages a hierarchical comparison pipeline with fully parallelized computing units and a data preprocessing strategy to maximize computational efficiency. Experimental results on the ISPD’18 benchmarks demonstrate that RouterAcc achieves consistent speedups of 2.1×–2.3× over TritonRoute with less than 1% quality degradation. With further co-optimization, RouterAcc attains speedups of 2.7×–11.8× while maintaining routing quality comparable to TritonRoute and surpassing Dr.CU 2.0 as well as the state-of-the-art (SOTA) FPGA-based approaches. Ruiyuan Guo, Zexu Zhang, Da Tang, Weiqi Shen, Haodong Lu 0001, Xiqiong Bai, Kun Wang 0005, Jianli Chen, Jun Yu 0010 |
DATE | 6 |
| 2026 | A Co-optimization Framework for Resolving Via Coloring Conflict in Multiple Patterning LithographyabstractAs integrated circuit technology nodes scale down, high via density challenges multiple patterning lithography (MPL). Existing methods for addressing via coloring conflicts mainly focus on detailed routing, yet they cannot resolve conflicts arising from vias that are fixed before routing, such as Power/Ground vias and obstruction vias. This paper presents a co-optimization framework to eliminate such inherent via coloring conflicts and boost routing efficiency. It proposes a conflict detection method identifying odd cycle and odd wheel violation patterns, balancing efficiency and precision well. It dynamically marks Forbidden Box and Forbidden Pair for mask-decomposition-aware placement and routing. A placement adjustment based on Directed Acyclic Graph (DAG) simultaneously handles overlaps between a cell’s Forbidden Box and Power/Ground vias, as well as illegal abutment of Forbidden Pair cells. During routing, Forbidden Box constraints guide pin access and via locations, while a final check resolves remaining conflicts through rip-up and reroute. Industrial benchmark experiments show the framework completely eliminates via coloring conflicts, and slightly reduces wirelength, via count and runtime. Haodong Lu 0001, Jianli Chen, Kun Wang 0005 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | End-to-end Compilation is All FPGAs Need: A Unified Overlay-based FPGA Compiler for Deep LearningabstractField-Programmable Gate Array (FPGA) has shown great application potential in deploying Neural Networks (NNs) due to the characteristics of programmability, low power consumption, etc. However, deploying NNs on FPGA is non-trivial because (1) Mainstream NNs pose significant FPGA architecture design challenges due to their large number of parameters, complex operations, and the need for data optimization, and (2) Supporting the deployment of different machine learning frameworks to FPGA requires significant manual effort, consuming a large amount of time. In this paper, we propose AutoCompiler, a unified compiler for mapping NNs to different FPGAs, along with overlay techniques to enable fast and efficient implementation. To the best of our knowledge, we are the first work to support both Deep Neural Networks (DNNs) and Transformer-based networks for overlay-based FPGA deployment. AutoCompiler comprises three integrated enablers: (1) Model Translator, built on top of a topology-based NNs representation, which can optimize the topology and data representation of the models from an algorithmic level based on different hardware configurations, e.g., DSP utilization, (2) Instruction Generator, which generates pipeline data streams according to various FPGA resource configurations by manipulating the instruction set at the upper level rapidly, and (3) End-to-end optimization, which moves as much of the computational processes as possible onto the FPGA chip and minimizes the interaction between CPU and FPGA. Extensive experiments on various Xilinx FPGAs show that AutoCompiler outperforms state-of-the-art overlay-based compiler by 1.2× - 1.35× and same-level GPUs by 1.15× - 1.59× for classic DNN models, and ViT inference, respectively. Haodong Lu 0001, Yinqiu Liu, Zexu Zhang, Kun Wang 0005 |
ASP-DAC | 2 |
| 2025 | A Precision-Steerable Electromigration Solver with Physics-Informed Adaptive Graph PartitioningabstractElectromigration-related reliability concerns in very large-scale integration (VLSI) circuits have garnered increasing attention as technology continues to scale. As integrated circuits shrink and their density rises, solving Korhonen's equation for the multi-segment interconnect line model becomes increasingly challenging. Recent advances in neural network-based approaches have demonstrated notable efficacy in addressing differential equations arising in physical modeling frameworks. Inspired by Physics-Informed Graph Neural Network (PIGNN) methodologies, we propose a novel Physics-Informed Message Passing (PIMNEM) architecture designed to solve coupled multi-domain Korhonen equations. At the same time, we introduce AdaptEM, which incorporates a graph partitioning mechanism with a hierarchical training strategy and employs the PIM-NEM architecture as a subgraph computation unit. AdaptEM enables multi-scale decomposition of interconnected circuits and facilitates hierarchical unsupervised learning via its hierarchical architecture. Unsupervised training is first applied to partitioned subgraphs using the PIMP mechanism, followed by global graph fine-tuning, where inter-subgraph boundary constraints are explicitly enforced through differentiable penalty terms. AdaptEM achieves a 20× speedup over FEM-based methods at the cost of about 0.5% accuracy loss. While AdaptEM may not match the absolute computational speed of state-of-the-art EM tools, its end-to-end unsupervised training framework, enhanced by a hierarchical subgraph training strategy, offers superior generalization capabilities and greater tuning flexibility. Zhaoyuan Liu, Haodong Lu 0001, Jianli Chen, Jun Yu 0010, Kun Wang 0005 |
ICCAD | 2 |
| 2025 | LUT-HD: Accelerating Hyperdimensional Computing Inference via Efficient Table LookupabstractHyperdimensional computing (HDC) has emerged as a promising cognitive computing paradigm, offering exceptional robustness and energy efficiency for intelligent applications. However, the computational demands of HDC, particularly during the encoding and associative search phases, pose significant challenges due to their time and resource intensity. In this paper, we propose LUT-HD, a software-hardware co-design framework that accelerates HDC inference by leveraging efficient table lookup techniques. First, we introduce a binary code quantization (BCQ) algorithm based on a lookup table (LUT) that transforms costly matrix-vector multiplications in HDC into simple table lookups using precomputed results. Next, we propose a custom FPGA-based accelerator tailored for LUT-based HDC to strike a balance between accuracy and efficiency. This accelerator incorporates a performance-optimized pipeline for encoding and associative search, enhancing computational speed and resource utilization. Experimental results demonstrate that LUT-HD achieves up to 14.6 × inference speedup and reduces 97.3% energy consumption compared to the GPU platform. In addition, compared to state-of-the-art (SOTA) HDC solutions, LUT-HD offers a 5.5× speedup with negligible accuracy loss and reduces 44.8% energy consumption. Haodong Lu 0001, Da Tang, Xiqiong Bai, Zexu Zhang, Kun Wang 0005 |
ICCAD | 1 |
| 2024 | PipeFuser: Building Flexible Pipeline Architecture for DNN Accelerators via Layer FusionabstractIn this paper, we propose a fused-pipeline architecture that leverages the layer fusion technique to harness the strengths of both non-pipeline and full-pipeline architectures while mitigating their disadvantages. In particular, we observe that the performance of the fused-pipeline accelerators is significantly influenced by the layer fusion strategies and intra-layer mapping schemes. To optimize and rapidly employ the fused-pipeline architecture, we present an end-to-end automation framework, named PipeFuser. At the core of PipeFuser is a genetic algorithm (GA)-based co-design engine, which is used to acquire near-optimal hardware configurations in the vast design space. Experimental results demonstrate that our fused-pipeline architecture achieves 2.3 × to 3.3 × higher performance over the non-pipeline design and 1.9 × to 2.5 × speedup compared to the full-pipeline architecture, with greater deployment flexibility. Xilang Zhou, Haodong Lu 0001, Kun Wang 0005 |
ASPDAC | 3 |
| 2024 | TrafficHD: Efficient Hyperdimensional Computing for Real-Time Network Traffic AnalyticsabstractWith the evolution of network infrastructure, the pattern of network traffic becomes unprecedentedly complex. Conventional machine learning algorithms struggle to cope with the high-dimensional data and real-time processing speeds required in such complex networks. Fortunately, Hyperdimensional Computing (HDC), which is power-efficient and supports parallel processing, provides a potential solution to this challenge. In this paper, we present TrafficHD, a novel classification framework that leverages HDC to analyze network traffic in real-time. By transforming network traffic features into high-dimensional binary vectors, TrafficHD enables the rapid execution of recognition tasks within the constraints of real-time systems. Extensive evaluations on a wide range of network tasks show that TrafficHD is 30.57× and 98.32× faster than state-of-the-art (SOTA) machine learning and HDC algorithms while providing 3× higher robustness to network noise. Haodong Lu 0001, Shiyan Bi, Xiaoming He 0004, Kun Wang 0005 |
DAC | 1 |
| 2024 | AutoHammer: Breaking the Compilation Wall Between Deep Neural Network and Overlay-based FPGA AcceleratorabstractField-Programmable Gate Array (FPGA) has shown great potential in accelerating Deep Neural Networks (DNNs) due to its characteristics of programmability and high power efficiency. In address the compilation challenges between DNNs and FPGA, we propose AutoHammer, an automated compiler for mapping DNNs to different FPGAs. Specifically, AutoHammer leverages overlay techniques to enable fast and effective implementation. Moreover, three enablers are integrated into AutoHammer. First, the Model Translator optimizes the topology and predicts a DNN's results based on different hardware configurations, built on top of a topology-based representation of DNNs. Second, the Instruction Generator generates pipeline data streams in various FPGA resource configurations by manipulating the instruction set at the upper level rapidly. Last, we realize the End-to-end Optimization, moving the whole computational processes onto the FPGA. Extensive experimental results show that AutoHammer improves great deployment efficiency when validated by 14 types of DNN models on 3 companies' (Xilinx, Fudan Micro, and Pango Micro) mainstream FPGA chips. Yinqiu Liu, Haodong Lu 0001, Zexu Zhang, Ruiqiu Chen, Kun Wang 0005 |
FPGA | 4 |
| 2024 | DNNMapper: An Elastic Framework for Mapping DNNs to Multi-die FPGAsabstractDeep Neural Networks (DNNs) have stimulated intensive FPGA-based acceleration solutions, and multi-die FPGAs offer abundant resources for implementing large-scale DNN workloads. However, current FPGA frameworks overlook the opportunities of optimization on multi-die FPGAs. In this paper, we propose an automated framework named DNNMapper, for mapping DNNs to multi-die FPGAs. With careful consideration of the unique architectural characteristics and resource constraints of multi-die FPGAs, DNNMapper involves model partitioning and resource allocation as two critical processes that map DNN layers onto respective FPGA dies and efficiently allocate hardware resources. DNNMapper employs a co-design engine based on a genetic algorithm, which co-optimizes model partitioning and resource allocation. Experimental results demonstrate that accelerators generated by DNNMapper offer superior performance and scalability, achieving up to 2× higher throughput and 1.3× to 1.9× higher DSP density. Moreover, our accelerator demonstrates a frequency improvement from 1.28× to 1.69×. Xilang Zhou, Haodong Lu 0001, Kun Wang 0005 |
ISCAS | 3 |
| 2023 | An Efficient Piecewise Linear Approximation of Non-linear Operations for Transformer InferenceabstractTransformer-based models have achieved remarkable performance across various tasks, while the computational complexity presents an obstacle for deploying on resource-constrained devices. To this end, this paper proposes an efficient approximation framework termed NPLA for approximating non-linear operations during Transformer inference on hardware accelerators. Specifically, NPLA enables the approximation of non-linear operations using non-uniform piecewise linear functions and directly converts coefficients into LUTs for hardware implementation. Experimental results demonstrate that NPLA can reduce the hardware cost by 13.43× in LUTs and 1.98× in DSP compared to the state-of-the-art method. Haodong Lu 0001, Qichang Mei, Kun Wang 0005 |
FCCM | 1 |
| 2023 | microGEMM: An Effective CNN-Based Inference Acceleration for Edge ComputingabstractConvolutional Neural Networks (CNNs), a widely recognized deep learning algorithm, have been utilized in various domains such as smart cities and healthcare. However, the remarkable performance of CNNs is accompanied by high resource overhead and deployment complexity. To address these challenges, CNN compilers have been developed to simplify convolutional operations for edge device deployment. One of the crucial components in CNN models is the General Matrix Multiply (GEMM) operation, which serves as the main computational kernel. In previous studies, efforts were made to improve the computation speed of GEMM by modifying the matrix calculation sequence, but they did not fully exploit the computing resources of edge devices. In this paper, we propose a novel GEMM-based acceleration algorithm, named microGEMM. The microGEMM algorithm divides convolutional data to reduce the memory access times during the GEMM calculation process. Moreover, the algorithm employs instruction-level optimization in the GEMM calculation unit, decreasing the cache miss rate. To better evaluate the superiority of microGEMM on resource-constrained devices, two edge-oriented metrics are proposed, namely CCPS & CCPoE. The microGEMM algorithm is implemented in C++ and compared with the standard GEMM algorithm (naiveGEMM) and the GEMM of the open-source Basic Linear Algebra Subprograms (BLAS) library (openblasGEMM). The experimental results demonstrate that microGEMM achieves a significant speedup, ranging from 5.67 × to 14.19 ×, compared to naiveGEMM. Haodong Lu 0001, Yinqiu Liu, Siguang Chen, Kun Wang 0005 |
ICC | 4 |
| 2023 | Auto-LUT: Auto Approximation of Non-Linear Operations for Neural Networks on FPGAabstractThe approximation of non-linear operation can simplify the logic design and save the system resources during the neural network inference on Field-Programmable Gate Array (FPGA). Prior work can approximate the non-linear operations with piecewise linear (PWL) function, but such approximation neglects considering the hardware overhead simultaneously. This paper proposes a novel approximation framework called Auto-LUT, which leverages a neural network to automatically approximate the non-linear operations. The framework formulates the approximation error and hardware overhead as a multi-objective optimization problem and employs an automated search mechanism to find the minimum number of segments and data bit width. To improve the approximation accuracy, we propose a bias clipping operation during the training of approximation networks, which enforces the model to approximate within the range of interest. Moreover, a hardware-friendly quantization scheme is further introduced to simulate the hardware behavior, thereby reducing the hardware overhead. Finally, a customized hardware architecture based on FPGA is utilized to deploy the quantized result. The experimental results show that Auto-LUT costs 56.32% less LUTs and 32.31% less flip-flops (FF) while reducing 4.32% approximation error compared to the state-of-the-art method. Haodong Lu 0001, Qichang Mei, Kun Wang 0005 |
ISCAS | 1 |
| 2021 | Distributed Machine Learning based Mitigating Straggler in Big Data EnvironmentabstractIn big data era, utilizing the parameter server paradigm has been regarded as an efficient and practical way to improve performance in processing deep learning (DL) applications. One of the main problems is that straggler greatly hinders DL training progress, but the previous methods cannot fully consider the resource utilization of the cluster when dealing with straggler. To mitigate straggler problem in parameter server, we propose a Deep Reinforcement Learning (DRL)-based framework called Distributed Actor-critic Reinforcement Learning (DARL) that can automatically adapt each worker's training load to the dynamic cluster without parameter settings. DARL employs state-of-the-art techniques to stabilize training and improve convergence, including distributed framework, multiple actors and prioritized experience replay. Meanwhile, we also apply our customized experience sampling method to fully exploit potentially good samples. Experiments using real DL workloads show that DARL outperforms the representative Bulk Synchronous Parallel (BSP) scheme by 57.8% and Stale Synchronous Parallel (SSP) by 50.3% in terms of per-iteration time in heterogeneous environment. Haodong Lu 0001, Kun Wang 0005 |
ICC | 1 |
| 2021 | Falcon: Addressing Stragglers in Heterogeneous Parameter Server Via Multiple ParallelismabstractThe parameter server architecture has shown promising performance advantages when handling deep learning (DL) applications. One crucial issue in this regard is the presence of stragglers, which significantly retards DL training progress. Previous solutions for solving stragglers may not fully exploit the computation resource of the cluster as evidenced by our experiments, especially in the heterogeneous environment. This motivates us to design a heterogeneity-aware parameter server paradigm that addresses stragglers and accelerates DL training from the perspective of computation parallelism. We introduce a novel methodology named straggler projection to give a comprehensive inspection of stragglers and reveal practical guidelines to solve this problem in two aspects: (1) controlling each worker's training speed via elastic training parallelism control and (2) transferring blocked tasks from stragglers to pioneers to fully utilize the computation resource. Following these guidelines, we propose the abstraction of parallelism as an infrastructure and design the Elastic-Parallelism Synchronous Parallel (EPSP) algorithm to handle distributed training and parameter synchronization, supporting both enforcedand slack-synchronization schemes. The whole idea has been implemented into a prototype called Falcon which effectively accelerates the DL training speed with the presence of stragglers. Evaluation under various benchmarks with baseline comparison demonstrates the superiority of our system. Specifically, Falcon reduces the training convergence time, by up to 61.83, 55.19, 38.92, and 23.68 percent shorter than FlexRR, Sync-opt, ConSGD, and DynSGD, respectively. Qihua Zhou, Song Guo 0001, Haodong Lu 0001, Li Li 0012, Minyi Guo, Yanfei Sun, Kun Wang 0005 |
IEEE Trans. Computers | 3 |
| 2021 | QoE-Based Task Offloading With Deep Reinforcement Learning in Edge-Enabled Internet of VehiclesabstractIn the transportation industry, task offloading services of edge-enabled Internet of Vehicles (IoV) are expected to provide vehicles with the better Quality of Experience (QoE). However, the various status of diverse edge servers and vehicles, as well as varying vehicular offloading modes, make a challenge of task offloading service. Therefore, to enhance the satisfaction of QoE, we first introduce a novel QoE model. Specifically, the emerging QoE model restricted by the energy consumption: 1) intelligent vehicles equipped with caching spaces and computing units may work as carriers; 2) various computational and caching capacities of edge servers can empower the offloading; and 3) unpredictable routings of the vehicles and edge servers can lead to diverse information transmission. We then propose an improved deep reinforcement learning (DRL) algorithm named PS-DDPG with the prioritized experience replay (PER) and the stochastic weight averaging (SWA) mechanisms based on deep deterministic policy gradients (DDPG) to seek an optimal offloading mode, saving energy consumption. Specifically, the PER scheme is proposed to enhance the availability of the experience replay buffer, thus accelerating the training. Moreover, reducing the noise in the training process and thus stabilizing the rewards, the SWA scheme is introduced to average weights. Extensive experiments certify the better performance, i.e., stability and convergence, of our PS-DDPG algorithm compared to existing work. Moreover, the experiments indicate that the QoE value can be improved by the proposed algorithm. Xiaoming He 0004, Haodong Lu 0001, Miao Du, Yingchi Mao, Kun Wang 0005 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Canary: Decentralized Distributed Deep Learning Via Gradient Sketch and Partition in Multi-Interface NetworksabstractThe multi-interface networks are efficient infrastructures to deploy distributed Deep Learning (DL) tasks as the model gradients generated by each worker can be exchanged to others via different links in parallel. Although this decentralized parameter synchronization mechanism can reduce the time of gradient exchange, building a high-performance distributed DL architecture still requires the balance of communication efficiency and computational utilization, i.e., addressing the issues of traffic burst, data consistency, and programming convenience. To achieve this goal, we intend to asynchronously exchange gradient pieces without the central control in multi-interface networks. We propose the Piece-level Gradient Exchange and Multi-interface Collective Communication to handle parameter synchronization and traffic transmission, respectively. Specifically, we design the gradient sketch approach based on 8-bit uniform quantization to compress gradient tensors and introduce the colayerabstraction to better handle gradient partition, exchange and pipelining. Also, we provide general programming interfaces to capture the synchronization semantics and build the Gradient Exchange Index (GEI) data structures to make our approach online applicable. We implement our algorithms into a prototype system called Canary by using PyTorch-1.4.0. Experiments conducted in Alibaba Cloud demonstrate that Canary reduces 56.28 percent traffic on average and completes the training by up to 1.61x, 2.28x, and 2.84x faster than BML, Ako on PyTorch, and PS on TensorFlow, respectively. Qihua Zhou, Kun Wang 0005, Haodong Lu 0001, Wenyao Xu, Yanfei Sun, Song Guo 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | QoE-driven Task Offloading with Deep Reinforcement Learning in Edge intelligent IoVabstractIn the transportation industry, task offloading services of edge intelligent Internet of Vehicles (IoV) are expected to provide vehicles with the better Quality of Experience (QoE). However, the various status of diverse edge servers and vehicles, as well as varying vehicular offloading modes, make a challenge of task offloading service. Therefore, to enhance the satisfaction of QoE, we first introduce a novel QoE model. Specifically, the emerging QoE model restricted by the energy consumption, (1) intelligent vehicles equipped with caching spaces and computing units may work as carriers; (2) various computational and caching capacities of edge servers can empower the offloading; (3) unpredictable routings of the vehicles and edge servers can lead to diverse information transmission. We then propose an improved deep reinforcement learning (DRL) algorithm named RA-DDPG with the prioritized experience replay (PER) and the stochastic weight averaging (SWA) mechanisms based on deep deterministic policy gradients (DDPG) to seek an optimal offloading mode, saving energy consumption. Extensive experiments certify the better performance, i.e., stability and convergence, of our RA-DDPG algorithm compared to existing work. Moreover, the experiments indicate that the QoE value can be improved by the proposed algorithm. Xiaoming He 0004, Haodong Lu 0001, Yingchi Mao, Kun Wang 0005 |
GLOBECOM | 2 |
| 2020 | Edge QoE: Computation Offloading With Deep Reinforcement Learning for Internet of ThingsabstractIn edge-enabled Internet of Things (IoT), computation offloading service is expected to offer users with better Quality of Experience (QoE) than traditional IoT. Unfortunately, the growing multiple tasks from users are occuring with the emergence of the IoT environment. Meanwhile, the current computation offloading with QoE is solved by deep reinforcement learning (DRL) with the issue of instability and slow convergence. Therefore, improving the QoE in edge-enabled IoT is still the ultimate challenge. In this article, to enhance the QoE, we propose a new QoE model to study the computation offloading. Specifically, the emerged QoE model can capture three influential elements: 1) service latency determined by local computing latency and transmission latency; 2) energy consumption according to local calculation and transmission consumption; and 3) task success rate based on the coding error probability. Moreover, we improve the deep deterministic policy gradients (DDPG) algorithm and propose a algorithm named the double-dueling-deterministic policy gradients (D3PG) based on the proposed model. Specifically, the actor network highly relies on the critic network, which makes the performance of the DDPG sensitive to the critic and thus leads to poor stability and slow convergence in the computation offloading process. To solve this, we redesign the critic network by using Double Q -learning and Dueling networks. Extensive experiments verify the better stability and faster convergence of our proposed algorithm than existing methods. In addition, experiments also indicate that our proposed algorithm can improve the QoE performance. Haodong Lu 0001, Xiaoming He 0004, Miao Du, Xiukai Ruan, Yanfei Sun, Kun Wang 0005 |
IEEE Internet Things J. | 1 |
| 2019 | Falcon: Towards Computation-Parallel Deep Learning in Heterogeneous Parameter ServerabstractParameter server paradigm has shown great performance superiority for handling deep learning (DL) applications. One crucial issue in this regard is the presence of stragglers, which significantly retards DL training progress. Previous solutions for solving straggler may not fully exploit the computation capacity of a cluster as evidenced by our experiments. This motivates us to make an attempt at building a new parameter server architecture that mitigates and addresses stragglers in heterogeneous DL from the perspective of computation parallelism. We introduce a novel methodology named straggler projection to give a comprehensive inspection of stragglers and reveal practical guidelines for resolving this problem: (1) reducing straggler emergence frequency via elastic parallelism control and (2) transferring blocked tasks to pioneer workers for fully exploiting cluster computation capacity. Following the guidelines, we propose the abstraction of parallelism as an infrastructure and elaborate the Elastic-Parallelism Synchronous Parallel (EPSP) that supports both enforced-and slack-synchronization schemes. The whole idea has been implemented in a prototype called Falcon which efficiently accelerates the DL training progress with the presence of stragglers. Evaluation under various benchmarks with baseline comparison evidences the superiority of our system. Specifically, Falcon yields shorter convergence time, by up to 61.83%, 55.19%, 38.92% and 23.68% reduction over FlexRR, Sync-opt, ConSGD and DynSGD, respectively. Qihua Zhou, Kun Wang 0005, Song Guo 0001, Haodong Lu 0001, Li Li 0012, Minyi Guo, Yanfei Sun |
ICDCS | 4 |