EDBT 2026 Demo / reviewers in the wild / expert
An Zou
dblp:202/9504
· DBLP profile ↗
32ranked-venue papers
7as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 7 first-author · 19 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TimeBill: Time-Budgeted Inference for Large Language ModelsabstractLarge Language Models (LLMs) are increasingly deployed in time-critical systems, such as robotics, autonomous driving, embodied intelligence, and industrial automation, where generating accurate responses within a given time budget is crucial for decision-making, control, or safety-critical tasks. However, the auto-regressive generation process of LLMs makes it challenging to model and estimate the end-to-end execution time. Furthermore, existing efficient inference methods based on a fixed key-value (KV) cache eviction ratio struggle to adapt to varying tasks with diverse time budgets, where an improper eviction ratio may lead to incomplete inference or a drop in response performance. In this paper, we propose TimeBill, a novel time-budgeted inference framework for LLMs that balances the inference efficiency and response performance. To be more specific, we propose a fine-grained response length predictor (RLP) and an execution time estimator (ETE) to accurately predict the end-to-end execution time of LLMs. Following this, we develop a time-budgeted efficient inference approach that adaptively adjusts the KV cache eviction ratio based on execution time prediction and the given time budget. Finally, through extensive experiments, we demonstrate the advantages of TimeBill in improving task completion rate and maintaining response performance under various overrun strategies. An Zou, Yehan Ma |
AAAI | 2 |
| 2026 | River-LLM: Large Language Model Seamless Exit Based on KV ShareabstractLarge Language Models (LLMs) have demonstrated exceptional performance across diverse domains but are increasingly constrained by high inference latency.Early Exit has emerged as a promising solution to accelerate inference by dynamically bypassing redundant layers.However, in decoder-only architectures, the efficiency of Early Exit is severely bottlenecked by the KV Cache Absence problem, where skipped layers fail to provide the necessary historical states for subsequent tokens.Existing solutions, such as recomputation or masking, either introduce significant latency overhead or incur severe precision loss, failing to bridge the gap between theoretical layer reduction and practical wall-clock speedup.In this paper, we propose River-LLM, a training-free framework that enables seamless token-level Early Exit.River-LLM introduces a lightweight KV-Shared Exit River that allows the backbone's missing KV cache to be naturally generated and preserved during the exit process, eliminating the need for costly recovery operations.Furthermore, we utilize state transition similarity within decoder blocks to predict cumulative KV errors and guide precise exit decisions.Extensive experiments on mathematical reasoning and code generation tasks demonstrate that River-LLM achieves 1.71× to 2.16× practical speedup while maintaining high generation quality. Yingtao Shen, An Zou |
ACL (1) | 2 |
| 2026 | Compression Space Search: RL-Based Combinational Compression for Neural NetworksabstractThe rising demand for lightweight, high-performance models on mobile and embedded platforms has accelerated the development of model compression techniques. Among these, combinational compression methods—which integrate multiple techniques such as pruning and quantization—offer complementary advantages over using individual methods alone. However, existing research typically focuses on specific combinations designed for a particular model architecture or task. These approaches often overlook the need for a general approach capable of identifying the optimal combination strategy, including the selection, sequence, and degree of applying compression methods. In this paper, we formalize the challenge of combining compression methods—specifically their selection, ordering, and compression degree—as a customized Markov Decision Process defined in a configurable compression space. To solve this, we introduce Compression Space Search (CSS), a practical RL-based framework for automatically and efficiently discovering optimal compression strategies. Experiments across CNN and transformer based vision models demonstrate that the proposed CSS achieves a 30 to 101 times reduction in bit operations while maintaining an accuracy drop of no more than 2%. Yingtao Shen, Yinchen Ni, Jiace Zhu, An Zou |
DATE | 5 |
| 2026 | LEAP: Lightweight Neural Network Inference Through Proactive Early-Exiting PredictionabstractIn recent years, the incorporation of early exit layers into deep neural networks has allowed inference to terminate earlier while maintaining accuracy. However, the passive decision-making involved in the these static exit placement creates a dilemma: fine-grained placement may cause high performance and energy overhead due to frequent exit layer execution, while coarse-grained placement may miss early exit opportunities. Moreover, common energy-saving techniques like adjusting processor configurations are not applicable once inference begins. To overcome these challenges and improve computation and energy efficiency, we propose LEAP, a software-hardware co-design approach. On the software side, LEAP proactively predicts exit points at runtime, reducing computation by enabling early exits without requiring every pre-placed exit layer to be executed. On the hardware side, LEAP adjusts processor settings—such as frequency and voltage—based on single or multiple predicted exits to optimize energy consumption while adhering to latency requirements. Extensive experimental results show that LEAP significantly improves efficiency. Compared to standard inference, LEAP reduces computation by up to 76.4% and saves up to 83.2% in energy. Compared to state-of-the-art early exit methods, LEAP achieves up to 27.9% less computation and 57.1% more energy savings, while maintaining similar accuracy and latency. Yingtao Shen, Xiangjie Li, Yehan Ma, Weidong Cao 0001, An Zou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | From ICs to Device: A Survey on Hardware Tampering Detection via Power Delivery Network and Signal TraceabstractProtecting the integrity of hardware against invasive tampering within the supply chain is critical to ensuring overall system resilience, reliability, and trustworthiness. As electronic systems become increasingly complex and globally distributed, the risk of malicious modifications or unauthorized alterations to hardware components continues to grow. In this survey, we provide a comprehensive exploration of detection and mitigation strategies for invasive hardware tampering across multiple levels of the hardware stack. We consider threats at various granularities–from individual on-board integrated circuits (ICs), to Printed Circuit Board Assemblies (PCBAs), and up to complete end-user devices. Our focus centers on three primary categories of invasive tampering: hardware Trojans, counterfeit components, and physical manipulations. A key emphasis of this survey is on hardware security techniques that leverage alterations in electrical characteristics induced by tampering. These include changes in power delivery network (PDN), signal trace, and other low-level electrical pathways. Such variations often serve as sensitive indicators of physical intrusions or modifications and are particularly useful for monitoring hardware integrity throughout its lifecycle–from manufacturing and deployment to maintenance and eventual decommissioning. We examine both golden-reference-based and golden-free detection approaches, highlighting their operational principles, design tradeoffs, and applicability in different threat scenarios. Furthermore, we survey evaluation methodologies and metrics used to assess the effectiveness, scalability, and robustness of these techniques. This article aims at providing a unified and up-to-date overview of detection research framework based on electrical characteristic variation, offering critical insights for researchers and practitioners working to safeguard hardware systems against invasive tampering and supply chain threats. Minqing Sun, Lanqi Ding, Huifeng Zhu, Yier Jin, An Zou |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | DenSparSA: A Balanced Systolic Array Approach for Dense and Sparse Matrix MultiplicationabstractNumerous studies have proposed hardware architectures to accelerate sparse matrix multiplication, but these approaches often incur substantial area and power overhead, significantly compromising their usage in dense scenarios. On the other hand, systolic arrays deliver high efficiency for dense matrix operations, but their application to sparse matrices remains challenging. An ideal design should process both dense and sparse matrices with high efficiency to satisfy performance and versatility requirements.In this paper, we introduce DenSparSA, a balanced systolic array centralized architecture that can execute sparse matrix computations with minimal overhead to original dense matrix computations. DenSparSA supports both single-side and dual-side unstructured sparse matrix multiplications with high efficiency. At the same time, the additional hardware required for managing sparsity is compact and decoupled from the conventional systolic array, allowing for minimal power overhead when switched back to dense matrix operations via circuit gating. The proposed design is implemented with Nangate 45 nm. Implementation results show that DenSparSA achieves a speedup ranging from $1.9 \times$ to $22 \times$ compared to the classic systolic array for sparse workloads, while maintaining relatively low area and power overhead. For dense workloads, the power overhead can be reduced to $\mathbf{1 2 \%}$ for BF16 and 5% for FP32. Compared with existing solutions for sparse acceleration, DenSparSA delivers competitive ($0.82 \times-1.32 \times$) efficiency in sparse scenarios and $1.17 \times-2.28 \times$ better efficiency for dense scenarios, indicating a better balance between both situations. Tianrui Ma, An Zou |
DAC | 5 |
| 2025 | RICH: Heterogeneous Computing for Real-Time Intelligent ControlabstractOver the past years, intelligent control tasks, such as deep neural networks (DNNs), have demonstrated significant potential in control systems. However, deploying intelligent control policies on heterogeneous computing platforms presents open challenges. These challenges extend beyond the apparent conflict between intensive computation and timing constraints and further encompass the interactions between task executions and complicated control performance. To address these challenges, this paper introduces RICH, a general and end-to-end approach to facilitate intelligent control tasks on heterogeneous computing architectures. RICH incorporates both offline Control-Oriented Computation and Resource Mapping (CCRM) and runtime Most Remaining Accelerator Segment Number First Scheduling (MRAF). Given the control tasks, the CCRM starts with balancing the computation workloads and processor resources with the goal of optimizing overall control performance. Subse-quently, the MRAF employs segment-level real-time scheduling to ensure the timely execution of tasks. Extensive experiments on the robotic arms applications (by hardware-in-the-loop simulator) demonstrate that the RICH can work as a general and end-to-end approach. These experiments reveal significant improvements in control performance, with enhancements of 50.7% observed for intelligent control applications deployed on heterogeneous computing platforms. Jintao Chen 0001, Yuankai Xu, Yinchen Ni, An Zou, Yehan Ma |
DATE | 4 |
| 2025 | RTHeter: Simulating Real-Time Scheduling of Multiple Tasks on Heterogeneous ArchitecturesabstractThe rising popularity of AI applications is driving the adoption of heterogeneous computing architectures to handle complex computations. However, as these heterogeneous architectures grow more complex, optimizing the scheduling of multiple tasks and meeting strict timing constraints becomes significantly challenging. Current studies on real-time scheduling on heterogeneous processors lack agile and flexible simulation tools that can quickly adapt to varying system settings, leading to inefficiencies in system design. Additionally, the high costs associated with evaluating real-time performance in terms of human and facility efforts further complicate the development process. To address these challenges, this paper introduces a comprehensive hierarchical simulating approach and a corresponding simulator designed for flexible heterogeneous computing platforms. The simulator supports ideal or practical, off-the-shelf or customizable heterogeneous architectures, upon which the simulator can execute both parallel and dependent tasks. Utilizing this simulator, we present two case studies that were time-consuming previously but are now easily achieved by the proposed simulator. The first case study reveals the possibility of using policy-based reinforcement learning to explore novel scheduling strategies; the second explores the dominant processors within heterogeneous architectures, providing insights for optimizing the heterogeneous architecture design. Yinchen Ni, Jiace Zhu, Yier Jin, An Zou |
DATE | 4 |
| 2025 | SSMDVFS: Microsecond-Scale DVFS on GPGPUs with Supervised and Self-Calibrated MLabstractOver the past decade, as GPUs have evolved to achieve higher computational performance, their power density has also accelerated. Consequently, improving energy efficiency and reducing power consumption has become critically important. Dynamic voltage and frequency scaling (DVFS) is an effective technique for enhancing energy efficiency. With the advent of integrated voltage regulators, DVFS can now operate on microsecond$(\boldsymbol{\mu}\mathbf{s})$timescales. However, developing a practical and effective strategy to guide rapid DVFS remains a significant challenge. This paper proposes a supervised and self-calibrated machine learning framework (SSMDVFS) to guide microsecond-scale GPU voltage and frequency scaling. This framework features an end-to-end design that encompasses data generation, neural network model design, training, compression, and final runtime calibration. Unlike analytical models, which struggle to accurately represent GPU architectures, and reinforcement learning approaches, which can be challenging to converge during runtime, the SSMDVFS offers a practical solution for guiding microsecond-scale voltage and frequency scaling. Experimental results demonstrate that the proposed framework improves energy-delay product (EDP) by 11.09% and outperforms analytical models and reinforcement learning approaches by 13.17% and 36.80 %, respectively. Minqing Sun, Yingtao Shen, Wei Yan 0005, Qinfen Hao, An Zou |
DATE | 6 |
| 2025 | RT-VirtIO: Towards the Real-Time Performance of VirtIO in a Two-Tier Computing ArchitectureabstractWith the popularity of virtualization technology, ensuring reliable I/O operations with timing constraints in virtual environments becomes increasingly critical. Timing-predictable virtual I/O enhances the responsiveness and efficiency of virtualized systems, facilitating their seamless integration into time-critical applications such as industrial automation and robotics. Its significance lies in meeting rigorous performance standards, minimizing latency, and consistently delivering predictable I/O performance. As a result, virtual machines can effectively support mission-critical and time-sensitive workloads. However, due to the complicated system architecture, the I/O operations in virtualization face competition from tasks within the same virtual machine and those in different virtual machines who share the same host machine. This study presents RT-VirtIO, a practical approach to provide predictable real-time I/O operations. RT-VirtIO addresses the challenges associated with lengthy data paths and complex resource management. Through early-stage characterization, this study identifies key factors contributing to poor I/O real-time performance and then builds an analytical model and a learning-based data-driven model to predict the tail I/O latency. Leveraging these two models, RT-VirtIO effectively captures these dynamics, enabling the development of a general and applicable optimization framework. Experimental results demonstrate that RT-VirtIO significantly improves real-time performance in virtual environments (by 20.07% ~ 30.90%) without necessitating hardware modifications, which exhibit promising applicability across a broader range of scenarios. Siwei Ye, Minqing Sun, Huifeng Zhu, Yier Jin, An Zou |
DATE | 5 |
| 2025 | Mesh Network Scheduling Based on Cyber-Physical Sensitivity for Wireless Control SystemsabstractWireless control systems (WCSs) are gaining rapid development in industrial automation. Compared to the star topology, mesh networks offer greater compatibility for large-scale applications that require high reliability, scalability, and extended coverage. In WCSs, multiple control loops share the multi-hop mesh network, leading to non-negligible and long- span communication latency in critical flows, which can severely degrade the overall control performance. Additionally, the criticality of each control flow largely depends on the features of the physical plant dynamics and the mesh network configuration, which is essential to properly and exactly represent. Moreover, the online scheduling and reconfiguration for large-scale mesh network for WCSs also pose unique challenges. In this paper, we propose a mesh network scheduling mechanism based on cyber-physical sensitivity. Firstly, we model each control loop as a switched system to represent the impact of arbitrary and fluctuating communication latency. Second, we propose a novel online criticality indicator, cyber-physical sensitivity (CP-Sensi), which accurately reflects the criticality of each control flow by synthesizing the switched model, runtime physical states, and network conditions. Finally, we design a CP-Sensi-based scheduling mechanism and an efficient piggyback-based network reconfiguration protocol tailored for mesh networks. Extensive studies with 12 control loops demonstrate that the proposed CP-Sensi and online mesh network scheduling achieve superior control performance compared to state-of-the-art approaches. Ruijie Fu, An Zou, Cailian Chen, Xin-Ping Guan, Yehan Ma |
RTAS | 2 |
| 2025 | HARD: Hardening Real-Time Scheduling and Analysis for Accelerator Enabled ComputingabstractDespite the advancements in supporting artificial intelligence, accelerator-enabled computing architectures still struggle to meet strict timing constraints due to the complex interactions between CPU cores and accelerators. Although various scheduling and response-time analysis techniques have been developed, a significant gap remains between the conservative hard real-time schedulability (i.e., worst-case response times) and the average measured schedulability on real systems. This pessimism significantly limits the deployment of hard real-time tasks on accelerator-enabled computing platforms. To address this, we propose HARD, a real-time scheduling approach that integrates scheduling strategies, response time analysis, and practical scheduler designs for general accelerator-enabled computing platforms. Benefiting the subtask level segmented characteristics that are ignored by classic schedulers, the proposed HARD can significantly improve the theoretically guaranteed hard real-time schedulability. Extensive experiments on off-the-shelf Intel CPUs and NVIDIA GPUs show that HARD outperforms state-of-the-art scheduling and analysis approaches, delivering a 11.3% improvement in hard real-time schedulability and a remarkable 45.1 % reduction in pessimism. Yinchen Ni, Tianrui Ma, Jintao Chen 0001, Chongye Yang, Siwei Ye, Yuankai Xu, Yier Jin, An Zou |
RTAS | 8 |
| 2025 | MATCH: Real-Time Scheduling of Multiple and Parallel Data Copies in Heterogeneous ArchitecturesabstractIn recent years, multiple data copies become popular in heterogeneous computing architectures. They enable parallel data transfer among diverse processing units. Tasks executed on such heterogeneous architectures often exhibit heightened re-source competitions and intricate task dependencies, posing challenges in meeting strict timing constraints. Due to the dominant roles of data copies in the heterogeneous architecture, effective scheduling and tight response time analysis could contribute to the timing performance of the entire heterogeneous computing system. In this work, we introduce MATCH, which offers realtime scheduling and end-to-end response time analysis for the multiple parallel data copies that are popular in mainstream heterogeneous architectures. We first identify the aggravated resource competition and task dependency from multiple data copies and comprehensive task execution patterns. Then, we provide a real-time scheduling strategy and cross-granularity schedulability analysis to deal with resource competition and task dependency. Extensive evaluation demonstrates that efficient scheduling and analysis on multiple parallel data copies can significantly improve the schedulability by 55.5%-144.4%. Additionally, experiments conducted on various scales of heterogeneous systems demonstrate that MATCH can significantly reduce pessimism in response time analysis by up to 22.8%-57.5%. Importantly, the proposed approach is compatible with existing scheduling approaches that do not consider multiple parallel data copies and are readily applied to off-the-shelf heterogeneous computing systems. Yinchen Ni, Yuankai Xu, Jintao Chen 0001, Jing Li 0025, Christopher D. Gill, Xuan Zhang 0001, Yier Jin, An Zou |
RTAS | 8 |
| 2025 | FALCON: FPGA Accelerated Real-Time Intelligent Controller for Autonomous SystemsabstractThe growing complexity and stringent real-time demands of autonomous systems, such as self-driving cars and drones, have driven the adoption of intelligent control methods based on deep neural networks (DNNs). While these methods offer improved control performance over traditional modelbased approaches, they also pose significant computational challenges, particularly for resource-constrained platforms. FieldProgrammable Gate Arrays (FPGAs) offer an attractive solution due to their energy efficiency and customizable architecture. In this work, we propose FALCON, an innovative approach for designing real-time intelligent controllers for autonomous systems using FPGA accelerators. Our approach begins with designing DNN-based intelligent controllers with varying levels of complexity and accuracy on the FPGA platform. Then, a performance function is proposed to capture the interplay among controller complexity, computational behavior, physical system characteristics, and overall control performance. Based on this performance function, we develop an algorithm-hardware codesign framework to determine the optimal control complexity, hardware configuration, and resource allocation. Finally, a case study on the co-design of intelligent controllers and FPGAbased overlay processors, together with a hardware-in-the-loop simulator, is conducted to demonstrate the advantages of the proposed methods. Compared to benchmarking controllers on other platforms, FALCON's optimized intelligent controller using FPGA accelerators shows competitive control performance with superior real-time capability and power efficiency. FALCON's optimization reduces the worst-case response time (WCRT) by up to 46.52%, improves the control performance by$1.93 \times$compared to the default setup. For performance per power efficiency, FALCON achieves a$3.67 \times$improvement compared to the DNN intelligent controller on TX2 and a remarkable$30.78 \times$improvement compared to traditional MPC on CPU. Siwei Ye, Jintao Chen 0001, Yehan Ma, An Zou |
RTSS | 4 |
| 2025 | Real-Time Scheduling and Analysis of Fixed-Priority Tasks on a Basic Heterogeneous Architecture With Multiple CPUs and Many PEsabstractWhile accelerator-based heterogeneous architectures have gained traction in accelerating AI tasks, effectively managing them with stringent timing constraints remains a challenge. Although many scheduling and response time analysis approaches are proposed for multi-core or heterogeneous multi-core (i.e., big.LITTLE cores) processors, direct application of them to accelerator-based heterogeneous architectures with multiple CPUs and numerous processing units (PEs) often results in significant pessimism. This paper introduces real-time scheduling and comprehensive response time analysis from unit-level micro view to job-level macro view, for general accelerator-based heterogeneous architectures, greatly enhancing schedulability and utilization rates. We begin by establishing a general task execution pattern on heterogeneous architectures that integrates multiple CPU cores and various PEs. Subsequently, we present a real-time scheduling strategy and corresponding response time analysis based on this task execution pattern from micro to macro views. Through extensive experiments conducted on GEMM and AI workloads, our proposed scheduling and response time analysis significantly outperforms state-of-the-art scheduling algorithms, improving schedulability by 10.3% to 52.9%. Furthermore, experiments on NVIDIA GPU systems indicate a potential pessimism reduction of up to 30.7%. As we target general heterogeneous architectures, our approach can be readily applied to off-the-shelf accelerator-based heterogeneous computing systems, ensuring adherence to deadlines and enhancing schedulability. Yuankai Xu, Yinchen Ni, Tiancheng He, Yier Jin, An Zou |
IEEE Trans. Computers | 6 |
| 2024 | ONE-SA: Enabling Nonlinear Operations in Systolic Arrays For Efficient and Flexible Neural Network Inference
Yinchen Ni, An Zou |
DATE | 5 |
| 2024 | SCENIC: Capability and Scheduling Co-Design for Intelligent Controller on Heterogeneous PlatformsabstractModern control systems, including robotics, drones, and autonomous vehicles, are increasingly incorporating intelligent controllers such as deep neural networks (DNNs) supported by heterogeneous processors. However, unlike conventional control algorithms on homogeneous platforms, the design and runtime execution of intelligent control tasks on heterogeneous computing platforms pose more rigorous demands and substantial challenges. These challenges encompass not only inherent conflicts between algorithm complexity and accuracy but also the couplings and trade-offs among run-time execution latency, end-to-end system performance, and reliability with timing constraints. To address these challenges, this paper introduces an end-to-end capability and scheduling co-design approach to efficiently design intelligent control tasks on heterogeneous computing architectures. We first introduce a novel and general control capability function, which bridges the control performance with the complexity of the intelligent controller, computation latency, and the properties of the physical plants. Subsequently, we formulate a comprehensive optimization problem to properly design algorithm capability and assign limited heterogeneous computational resources from offline heterogeneous resource allocation to run-time execution. Finally, we present a case study on the intelligent control of autonomous quadcopters (with the hardware-in-the-loop simulator built on Microsoft AirSim), and the extensive experiments demonstrate the superiority of the capability and scheduling co-design in terms of overall system performance compared with state-of-the-art design approaches. Jintao Chen 0001, An Zou, Yuankai Xu, Yehan Ma |
RTSS | 2 |
| 2024 | Performance Optimization and Stability Guarantees for Multi-tier Real-Time Control SystemsabstractModern control systems are embracing multi-tier architectures integrating end devices and edge servers. However, due to the distinct control performance demands associated with each control task, it is a formidable challenge to optimize the control performance of multiple control tasks subject to stringent computation resource constraints while guaranteeing stability. Moreover, inherent contradictions exist in the timing aspect between the stability guarantee, which relies on offline analysis, and the run-time control performance, which should be enhanced online. It is essential to bridge the gap between the real-time scheduling of control tasks and their actual control performance. In this paper, we propose a novel real-time scheduling approach for multi-tier control systems, which leverages end devices for executing real-time control tasks and edge devices for runtime coordination. Specifically, we first introduce a new datadriven value function, called time/state/utility functions (TSUF), for modeling control system performance. TSUF captures not only timing but also the dynamic states of the physical plants. Subsequently, we propose value-based control scheduling (VCS), which is a multi-granularity scheduling mechanism based on our TSUF value function. VCS distinguishes the scheduling of stability jobs for ensuring system stability and performance jobs for optimizing real-time control performance based on run-time physical states. Finally, through realistic case studies involving multiple control loops, we demonstrate the advantages of VCS over existing scheduling approaches in terms of both control and real-time performance. Yehan Ma, Ruijie Fu, An Zou, Jing Li 0025, Cailian Chen, Chenyang Lu 0001, Xin-Ping Guan |
RTSS | 3 |
| 2024 | Smart Sensing and Communication Co-Design for IIoT-Based Control SystemsabstractIndustrial Internet of Things (IIoT)-based control is growing rapidly, such as smart factories and industrial automation. Sensing and transmitting physical state measurements is the first step and the prerequisite for IIoT-based control. However, sensor interference (e.g., electromagnetic interference on sensing, temperature, and humidity variations in the field) and network interference (e.g., metal obstacles and background noises) may destroy the control performance by interfering with sensing and communication processes. Most of the present upstream “fixed sensors-networking-state estimation” approaches cannot effectively deal with sensor and network interferences due to the fixed measurements/estimation and network resource limitations. To optimize the performance of IIoT-based control, we propose a smart sensing and communication co-design (SSCC) framework to select more potential sensors and establish the corresponding network scheduling. SSCC consists of a smart estimator (SE) and a sensing communication mode switching (SCMS) agent. The SE detects sensor interference and obtains resilient state estimation based on collaborative sensing. SCMS agent dynamically switches sensor selections and network configurations (routing and transmission number) in an integrated manner based on the network and plant states by solving a performance optimization problem. We propose a lightweight SCMS approach by searching a predefined mode table. We perform simulations integrating TOSSIM and MATLAB/Simulink, and semi-physical experiments on a real wireless sensor-actuator network composed of TelosB nodes. The results show that the SSCC framework can effectively improve the control performance and enhance network energy efficiency under various types of interference by dynamically selecting sensors and allocating network resources. Ruijie Fu, Jintao Chen 0001, Yutong Lin, An Zou, Cailian Chen, Xin-Ping Guan, Yehan Ma |
IEEE Internet Things J. | 4 |
| 2024 | Comprehensive Optimal Network Scheduling Strategies for Wireless Control SystemsabstractAlthough wireless control is one of the key technologies for future industries, most wireless networks are only used for monitoring. When wireless networks are applied to transmit control commands, the uncertain link qualities and limited network resources may destroy the performance of multi-loop control systems. Hence, it is critical to allocate these resources to optimize the control performance as the network condition changes and plants evolve. This article presents comprehensive optimal scheduling strategies for wireless control systems based on adaptive dynamic programming. First, we propose an effective adaptive dynamic programming scheduling (ADPS) strategy to solve the optimal scheduling problem based on the single-step control performance at runtime while significantly reducing computational complexity. Moreover, to overcome the “short-sightedness” of single-step performance prediction, we extend ADPS to ADPS-m ( m ulti-step prediction), which optimizes multi-step performance by incorporating a longer-horizon evolution of the plants. Furthermore, we propose ADPS-H ( H eterogeneous flow scheduling) to support heterogeneous flows with different data rates and sizes and ADPS-H-m ( m ulti-step prediction for H eterogeneous flow scheduling), which schedules heterogeneous flows in a longer prediction horizon. We prove that all these scheduling strategies can achieve optimality and stability under mild assumptions. Extensive experiments integrating TOSSIM and MATLAB/Simulink are performed to evaluate all of the proposed methods in case studies of four- and ten-loop control systems. The simulation results demonstrate that these strategies can effectively improve the control performance at lower computing costs under both cyber and physical disturbances. Under the noise level of \(-\) 76 dBm, for the four-loop case, ADPS achieves the same control performance as the linear programming while saving 99.5% of the execution time. ADPS-m further improves the control performance by up to 27.0% compared with ADPS at the prediction horizon of 3, and ADPS-H-m improves the performance by up to 32.3% and 8.4% compared with round-robin and ADPS-H, respectively. The ten-loop case indicates the effectiveness and scalability of the proposed approaches. Ruijie Fu, Lancong Guo, An Zou, Cailian Chen, Xin-Ping Guan, Yehan Ma |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2023 | Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceabstractBy adding exiting layers to the deep learning networks, early exit can terminate the inference earlier with accurate results. However, the passive decision-making of whether to exit or continue the next layer has to go through every pre-placed exiting layer until it exits. In addition, it is hard to adjust the configurations of the computing platforms alongside the inference proceeds. By incorporating a low-cost prediction engine, we propose a Predictive Exit framework for computation- and energy-efficient deep learning applications. Predictive Exit can forecast where the network will exit (i.e., establish the number of remaining layers to finish the inference), which effectively reduces the network computation cost by exiting on time without running every pre-placed exiting layer. Moreover, according to the number of remaining layers, proper computing configurations (i.e., frequency and voltage) are selected to execute the network to further save energy. Extensive experimental results demonstrate that Predictive Exit achieves up to 96.2% computation reduction and 72.9% energy-saving compared with classic deep learning networks; and 12.8% computation reduction and 37.6% energy-saving compared with the early exit under state-of-the-art exiting strategies, given the same inference accuracy and latency. Xiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu, Yingtao Shen, Yehan Ma, An Zou |
AAAI | 7 |
| 2023 | EENet: Energy Efficient Neural Networks with Run-time Power ManagementabstractDeep learning approaches, such as convolution neural networks (CNNs), have achieved tremendous success in versatile applications. However, one of the challenges to deploy the deep learning models on resource-constrained systems is its huge energy cost. As a dynamic inference approach, early exit adds exiting layers to the networks, which can terminate the inference earlier with accurate results to save energy. The current passive decision-making for energy regulation of early exit cannot adapt to ongoing inference status, varying inference workloads, and timing constraints, let alone guide the reasonable configuration of the computing platforms alongside the inference proceeds for potential energy saving. In this paper, we propose an Energy Efficient Neural Networks (EENet), which introduces a plug-in module to the state-of-the-art networks by incorporating run-time power management. Within each inference, we establish prediction of where the network will exit and adjust computing configurations (i.e., frequency and voltage) accordingly over a small timescale. Considering multiple inferences over a large timescale, we provide frequency and voltage calibration advice, given inference workloads and timing constraints. Finally, the dynamic voltage and frequency scaling (DVFS) governor configures voltage and frequency to execute the network according to the prediction and calibration. Extensive experimental results demonstrate that EENet achieves up to 63.8% energy-saving compared with classic deep learning networks and 21.5% energy-saving compared with the early exit under state-of-the-art exiting strategies, together with improved timing performance. Xiangjie Li, Yingtao Shen, An Zou, Yehan Ma |
DAC | 3 |
| 2023 | Energy Efficient Real-Time Scheduling on Heterogeneous Architectures with Self-SuspensionabstractIt is witnessed that heterogeneous architectures, such as GPUs, TPUs, and FPGAs, have made complex algorithms practical in the last decade. Despite multiple efforts to study the scheduling of these parallel and complex tasks on heterogeneous architectures, the power and energy consumption of the platforms have yet to be well managed under real-time task deadlines. To establish high schedulability in heterogeneous architectures, many scheduling strategies and models, such as multi-segment selfsuspension (MSSS), have been proposed by pioneer researchers. However, directly applying this model to heterogeneous architectures with multiple CPUs and many processing elements (PEs) suffers aggravated power consumption due to the pessimism in the scheduling algorithm and the tolerance margin in the worst-case execution time (WCET) model. Therefore, this paper presents an energy-efficient real-time scheduling approach called EESchedule, which works on heterogeneous architectures with guaranteed schedulability and improved power efficiency. In EESchedule, we build a general task execution model for the general heterogeneous architectures integrating multiple CPUs and many PEs. Then, an energy-efficient real-time scheduling strategy is introduced. Next, the response time and corresponding schedulability analysis are presented for EESchedule. Finally, extensive experiments on heterogeneous NVIDIA Jetson TX2 embedded systems and GPU servers with the Intel i9-10900x CPU and RTX 3080 GPU demonstrate that the EESchedule could achieve the same schedulability with 16.8%-40.7% and 39.0%-48.2% reduced power and energy consumption in comparison with state-of-the-art scheduling algorithms. Yuankai Xu, Jing Li 0025, Yehan Ma, Yier Jin, Christopher D. Gill, Xuan Zhang 0001, An Zou |
ISLPED | 9 |
| 2023 | F-LEMMA: Fast Learning-Based Energy Management for Multi-/Many-Core ProcessorsabstractOver the last two decades, as microprocessors have evolved to achieve higher computational performance, their power density has also increased at an accelerated rate. Improving energy efficiency and reducing power consumption are therefore critically important to modern computing systems. One effective technique for improving energy efficiency is dynamic voltage and frequency scaling (DVFS). With the emergence of integrated voltage regulators (IVRs), the speed of DVFS can reach microsecond ($\mu \text{s}$) timescales. However, a practical and effective strategy to guide fast DVFS remains a challenge. In this article, we propose F-LEMMA: a fast, learning-based, hierarchical DVFS framework consisting of a global power allocator in the kernel space, a reinforcement learning-based power management scheme at the architecture level, and a swift controller at the digital circuit level. This hierarchical approach leverages computation at the system and architecture levels with the short response time of the swift controller to achieve effective and rapid$\mu \text{s}$-level power management supported by the IVR. Our experimental results demonstrate that F-LEMMA can achieve significant energy savings (35.2%) across a broad range of workloads. Conservatively compared with existing state-of-the-art DVFS-based power management schemes that can only operate at millisecond timescales, F-LEMMA can provide notable (up to 11%) energy-delay product (EDP) improvements across benchmarks. Compared with state-of-the-art nonlearning-based power management, our method has a universally positive effect on evaluated benchmarks, proving its adaptability. An Zou, Yehan Ma, Karthik Garimella, Christopher D. Gill, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | RTGPU: Real-Time GPU Scheduling of Hard Deadline Parallel Tasks With Fine-Grain UtilizationabstractMany emerging cyber-physical systems, such as autonomous vehicles and robots, rely heavily on artificial intelligence and machine learning algorithms to perform important system operations. Since these highly parallel applications are computationally intensive, they need to be accelerated by graphics processing units (GPUs) to meet stringent timing constraints. However, despite the wide adoption of GPUs, efficiently scheduling multiple GPU applications while providing rigorous real-time guarantees remains challenging. Each GPU application has multiple CPU execution and memory copy segments, with GPU kernels running on different hardware resources. Because of the complicated interactions between heterogeneous segments of parallel tasks, high schedulability is hard to achieve with conventional approaches. This paper proposes RTGPU, which combines fine-grain GPU partitioning on the system-side with a novel scheduling algorithm on the theory-side. We start by building a model for CPU and memory copy segments. Leveraging persistent threads, we then implement fine-grained GPU partitioning with improved performance through interleaved execution. To reap the benefits of fine-grained GPU partitioning and schedule multiple parallel GPU applications, we propose a novel real-time scheduling algorithm based on federated scheduling and grid search with uniprocessor fixed-priority scheduling. Our approach provides real-time guarantees to meet hard deadlines and achieves over 11% improvement in system throughput and up to 57% schedulability improvement compared with previous work. We validate and evaluate RTGPU on NVIDIA GPU systems. Our system-side techniques can be applied on mainstream GPUs, and the proposed scheduling theory can be used in general heterogeneous computing platforms which have a similar task execution pattern. An Zou, Jing Li 0025, Christopher D. Gill, Xuan Zhang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | SHAPE: Scheduling of Fixed-Priority Tasks on Heterogeneous Architectures with Multiple CPUs and Many PEsabstractDespite being employed in burgeoning efforts to accelerate artificial intelligence, heterogeneous architectures have yet to be well managed with strict timing constraints. As a classic task model, multi-segment self-suspension (MSSS) has been proposed for general I/O-intensive systems and computation offloading. However, directly applying this model to heterogeneous architectures with multiple CPUs and many processing units (PEs) suffers tremendous pessimism. In this paper, we present a real-time scheduling approach, SHAPE, for general heterogeneous architectures with significant schedulability and improved utilization rate. We start with building the general task execution pattern on a heterogeneous architecture integrating multiple CPU cores and many PEs such as GPU streaming multiprocessors and FPGA IP cores. A real-time scheduling strategy and corresponding schedulability analysis are presented following the task execution pattern. Compared with state-of-the-art scheduling algorithms through comprehensive experiments on unified and versatile tasks, SHAPE improves the schedulability by 11.1% - 100%. Moreover, experiments performed on the NVIDIA GPU systems further indicate up to 70.9% of pessimism reduction can be achieved by the proposed scheduling. Since we target general heterogeneous architectures, SHAPE can be directly applied to off-the-shelf heterogeneous computing systems with guaranteed deadlines and improved schedulability. Yuankai Xu, Tiancheng He, Yehan Ma, Yier Jin, An Zou |
ICCAD | 6 |
| 2021 | System-level Early-stage Modeling and Evaluation of IVR-assisted Processor Power Delivery System
An Zou, Huifeng Zhu, Jingwen Leng, Xin He 0011, Vijay Janapa Reddi, Christopher D. Gill, Xuan Zhang 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2020 | Real-Time Scheduling upon a Host-Centric Acceleration Architecture with Data OffloadingabstractChallenging scheduling problems arise in the implementation of cyber-physical systems upon heterogeneous platforms with (serial) data offloading and (parallel) computation. In this paper, we adapt techniques from scheduling theory to model, analyze, and derive scheduling algorithms for real-time workloads on such platforms. We characterize the performance of the proposed algorithms, both analytically via the approximation ratio metric and experimentally through simulation experiments upon synthetic workloads that are justified via a case study on a CPU-GPU platform. The evaluation exposes some divergence between the analytical characterization and experimental one; recommendations that seek to balance such divergent characterizations are made regarding the choice of algorithmic approaches. Jinghao Sun, Jing Li 0025, Zhishan Guo, An Zou, Xuan Zhang 0001, Kunal Agrawal 0001, Sanjoy Baruah |
RTAS | 4 |
| 2020 | Voltage-Stacked Power Delivery Systems: Reliability, Efficiency, and Power ManagementabstractIn today's manycore processors, the energy loss of more than 20% may result from inherent inefficiencies of conventional power delivery system (PDS) design. By stacking multiple voltage domains in series to lower the step-down conversion ratio of the off-chip voltage regulator module (VRM) and reduce the energy loss along the path of the power delivery network (PDN), voltage stacking (VS) offers a novel alternative power delivery technique to fundamentally improve power delivery efficiency (PDE). However, VS suffers from aggravated supply voltage noise from the current imbalance, which hinders its adoption. In this article, we investigate practical VS implementation in manycore processors to improve PDE and achieve reliable performance, while maintaining compatibility with advanced power management techniques. We first present the system configuration of a voltage-stacked manycore processor. We then systematically characterize supply voltage noise in VS, identify global, and residual differential currents as its dominant contributors, and calculate the possible worst supply voltage noise. We next propose a hybrid voltage regulation solution, based on a charge-recycling off-chip voltage regulator and distributed integrated voltage regulators, to mitigate supply voltage noise effectively. We also study the compatibility of VS with higher-level power management techniques. Finally, the performance of a voltage-stacked GPU system is comprehensively evaluated. The simulation results show that our approach can achieve 93.5% PDE, reducing the power loss by 13.6% compared to conventional single-layer PDS. An Zou, Jingwen Leng, Xin He 0011, Yazhou Zu, Christopher D. Gill, Vijay Janapa Reddi, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Efficient and reliable power delivery in voltage-stacked manycore system with hybrid charge-recycling regulatorsabstractVoltage stacking (VS) fundamentally improves power delivery efficiency (PDE) by series-stacking multiple voltage domains to eliminate explicit step-down voltage conversion and reduce energy loss along the power delivery path. However, it suffers from aggravated supply noise, preventing its adoption in mainstream computing systems. In this paper, we investigate a practical approach to enabling efficient and reliable power delivery in voltage-stacked manycore systems that can ensure worst-case supply noise reliability without excessive costly over-design. We start by developing an analytical model to capture the essential noise behaviors in VS. It allows us to identify dominant noise contributor and derive the worst-case conditions. With this in-depth understanding, we propose a hybrid voltage regulation solution to effectively mitigate noise with worst-case guarantees. When evaluated with real-world benchmarks, our solution can achieve 93.8% power delivery efficiency, an improvement of 13.9% over the conventional baseline. An Zou, Jingwen Leng, Xin He 0011, Yazhou Zu, Vijay Janapa Reddi, Xuan Zhang 0001 |
DAC | 1 |
| 2018 | Voltage-Stacked GPUs: A Control Theory Driven Cross-Layer Solution for Practical Voltage Stacking in GPUsabstractMore than 20% of the available energy is lost in "the last centimeter" from the PCB board to the microprocessor chip due to inherent inefficiencies of power delivery subsystems (PDSs) in today's computing systems. By series-stacking multiple voltage domains to eliminate explicit voltage conversion and reduce loss along the power delivery path, voltage stacking (VS) is a novel configuration that can improve power delivery efficiency (PDE). However, VS suffers from aggravated levels of supply noise caused by current imbalance between the stacking layers, preventing its practical adoption in mainstream computing systems. Throughput-centric manycore architectures such as GPUs intrinsically exhibit more balanced workloads, yet suffer from lower PDE, making them ideal platforms to implement voltage stacking. In this paper, we present a cross-layer approach to practical voltage stacking implementation in GPUs. It combines circuit-level voltage regulation using distributed charge-recycling integrated voltage regulators (CR-IVRs) with architecture-level voltage smoothing guided by control theory. Our proposed voltage-stacked GPUs can eliminate 61.5% of total PDS energy loss and achieve 92.3% system-level power delivery efficiency, a 12.3% improvement over the conventional single-layer based PDS. Compared to the circuit-only solution, the cross-layer approach significantly reduces the implementation cost of voltage stacking (88% reduction in area overhead) without compromising supply reliability under worst-case scenarios and across a wide range of real-world benchmarks. In addition, we demonstrate that the cross-layer solution not only complements on-chip CR-IVRs to transparently manage current imbalance and restore stable layer voltages, but also serves as a seamless interface to accommodate higher-level power optimization techniques, traditionally thought to be incompatible with a VS configuration. An Zou, Jingwen Leng, Xin He 0011, Yazhou Zu, Christopher D. Gill, Vijay Janapa Reddi, Xuan Zhang 0001 |
MICRO | 1 |
| 2017 | Ivory: Early-Stage Design Space Exploration Tool for Integrated Voltage RegulatorsabstractDespite being employed in burgeoning efforts to improve power delivery efficiency, integrated voltage regulators (IVRs) have yet to be evaluated in a rigorous, systematic, or quantitative manner. To fulfill this need, we present Ivory, a high-level design space exploration tool capable of providing accurate conversion efficiency, static performance characteristics, and dynamic transient responses of an IVR-enabled power delivery subsystem (PDS), enabling rapid trade-off exploration at early design stage, approximately 1000x faster than SPICE simulation. We demonstrate and validate Ivory with a wide spectrum of IVR topologies. In addition, we present a case study using Ivory to reveal the optimal PDS configurations, with underlying power break-downs and area overheads for the GPU manycore architecture, which has yet to embrace IVRs. An Zou, Jingwen Leng, Yazhou Zu, Tao Tong, Vijay Janapa Reddi, David Brooks 0001, Gu-Yeon Wei, Xuan Zhang 0001 |
DAC | 1 |