EDBT 2026 Demo / reviewers in the wild / expert
Biao Hu 0001
dblp:123/6590-1
· DBLP profile ↗
40ranked-venue papers
26as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 12 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-author · 2 since 2021Computer networks · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BOLT-PM: An Adaptive Bayesian Optimization Framework for Latency-Sensitive Task and Power Management in AIoT DevicesabstractThe integration of AI tasks into Artificial Intelligence of Things (AIoT) devices has gained significant attention for its ability to reduce transmission latency and enhance data privacy compared to traditional cloud-based solutions. However, heterogeneous AIoT platforms face the challenge of balancing stringent latency constraints with energy efficiency under dynamic AI workloads. To address this issue, we propose BOLT-PM, a Bayesian Optimization framework for Latency-sensitive Task and Power Management, which dynamically allocates computing resources and optimizes power consumption in AIoT devices. We design a multi-branch neural network to accurately predict task execution times under varying resource constraints, enabling precise and adaptive performance estimation. We further enhance the Bayesian optimization process by embedding a variational autoencoder (VAE) to construct a smooth latent representation of the search space, thereby accelerating convergence and improving optimization stability. We implement BOLT-PM as a lightweight, platform-agnostic runtime framework that continuously adapts to workload fluctuations, maintaining latency guarantees while minimizing energy consumption. We evaluate BOLT-PM on several commercial AIoT boards, including RK3588, Jetson TX2, and Raspberry Pi, and compare it against classical heuristic methods and state-of-the-art approaches. Experimental results show that BOLT-PM achieves substantial energy savings while ensuring low-latency AI task execution, demonstrating its effectiveness as a robust and energy-efficient solution for power-aware AIoT applications. Biao Hu 0001, Chenyu Cai, Xincheng Yang, Mingguo Zhao |
IEEE Internet Things J. | 1 |
| 2026 | Pushing Physical Limits and Uncovering Motion Templates of Spine-Based Quadruped Locomotion via Reinforcement LearningabstractFlexible spines are critical to the remarkable agility and speed of animals. Translating this biological advantage to quadruped robots presents a significant control challenge, particularly in coordinating the spine and limbs for maximal velocity. In this work, we utilize reinforcement learning (RL) to develop high-speed locomotion for a bioinspired mouse robot with a lateral flexible spine. The resulting controller achieves motor performance that demonstrably surpasses non-spined and model-based methods. More importantly, our analysis reveals the principles behind this performance: the emergence of two distinct motion templates. For high-speed walking, the robot learns a “whip-like” spinal oscillation to increase leg swing frequency, while for agile turning, it adopts a dynamic “bend-and-straighten” pattern. These findings demonstrate the capability of RL to not only generate high-performance controllers but also to produce emergent strategies that, upon analysis, reveal underlying principles of high-speed, spine-driven locomotion. Zhenshan Bing, Yulong Xiao, Yuhong Huang, Long Cheng 0007, Biao Hu 0001, Gang Chen 0023, Yang Gao 0001, Fuchun Sun 0001, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Robotics | 6 |
| 2025 | An Adaptive ROS2 Node Deployment Framework in Mobile Edge-Robot SystemsabstractMobile edge computing is an emerging computing paradigm that enhances the computational capabilities of mobile devices by offloading intensive tasks to edge servers. In robotic systems, MEC can significantly reduce response times and improve user experience. However, as robots move through their environments, factors such as the distance to edge servers and physical obstructions fluctuate, leading to variations in communication bandwidth and, consequently, communication delays. To address these challenges, this paper proposes an adaptive computing node offloading framework (ARDF) designed to optimize the dynamic deployment of Robot Operating System 2 (ROS2) nodes in robot-edge environments. The framework enables developers to flexibly deploy robotic computing tasks based on varying computational and network conditions. We validate its effectiveness through experiments involving robotic arm control, 3D detection applications, and numerically simulated ROS2 tasks under different bandwidth conditions. The results demonstrate that the framework significantly improves response times for ROS2 applications, even under fluctuating computational loads and network constraints. The code for the framework ARDF can be found on GitHub1. Xincheng Yang, Biao Hu 0001 |
IROS | 2 |
| 2025 | Mixed-Criticality Scheduling Toward Real-Time Applications in a Vehicular Edge Computing SystemabstractABSTRACT Scheduling applications in vehicular edge computing (VEC) systems poses significant challenges due to strict timing constraints and varying levels of criticality. This paper presents a three‐stage scheduling framework designed to efficiently manage the execution of mixed‐criticality applications. The proposed method introduces scheduling policies that reduce the complexity of scheduling dual‐criticality DAG (Directed Acyclic Graph) applications on servers by transforming them into equivalent uniprocessor scheduling problems. To further enhance performance, a population‐based evolutionary algorithm is employed to optimize virtual machine configurations on each server, while a game‐theoretic approach assigns DAG applications to servers. Experimental results show that the proposed scheme outperforms both state‐of‐the‐art dynamic programming (DP) and particle swarm optimization (PSO) methods. The proposed MCS approach achieves a strong balance between scheduling quality and computational efficiency, with an of 0.87, an 80% success rate, and a low computation time (310 s), making it well‐suited for real‐time edge systems. Compared to other methods like PSO+, DP, and OneVM, MCS offers near‐optimal performance while avoiding the high computational cost and scalability limitations faced by those alternatives. Biao Hu 0001, Xincheng Yang |
Concurr. Comput. Pract. Exp. | 1 |
| 2025 | Bayesian Optimization-Based Time-Sensitive and Power-Efficient DNN Task Partitioning in a Dynamic IoT Computing SystemabstractDeploying deep neural networks (DNNs) on resource-constrained Internet of Things (IoT) devices is challenging due to limited processing power, energy constraints, and stringent latency requirements. This article introduces a new method to split DNN tasks in IoT systems, focusing on saving power and reducing delays. We use Gaussian process regression (GPR) to predict how long tasks will take under different conditions. We also use a simple linear regression model to estimate how much power IoT devices use based on their CPU usage. These predictive models are integrated into a Bayesian optimization framework to determine the optimal DNN task partitioning point, balancing latency and energy efficiency. The system adjusts to changes in the network and device conditions, ensuring it works well in different situations. Experiments on a heterogeneous IoT testbed demonstrate that GPR accurately predicts execution latencies, and the linear regression model provides reliable power consumption estimates. The Bayesian optimization algorithm efficiently explores the tradeoff space, offering low power consumption and high latency satisfaction rates. Our code is shared for public usehttps://github.com/nucleusbiao/Time-Sensitive-and-Power-Efficient-DNN-Task-Partitioning. Biao Hu 0001, Qianru Wang, Xincheng Yang |
IEEE Internet Things J. | 1 |
| 2025 | Coordinating Computational Capacity for Adaptive Federated Learning in Heterogeneous Edge Computing SystemsabstractWith the rapid growth of IoT technology and the rise of smart devices, edge computing, particularly federated learning (FL), has gained importance for preserving user data privacy. However, FL faces challenges like non-independent identically distributed data and device heterogeneity, leading to model disparities and reduced precision. Our research proposes a novel adaptive FL framework specifically engineered to synchronize computational capacities within heterogeneous edge computing landscapes. Building upon the proof of convergence boundaries for local aggregation model, this algorithm adapts the number of iterations for local updates by considering the resource consumption relationship between local aggregation model and the local updated model by various clients. This method exhibit adaptability within an environment where disparities in edge device computational capacities exist, effectively balancing computational prowess among diverse devices and enhancing the output performance of federated learning Experiments on MNIST and PlantVillage datasets show that in heterogeneous environments, our algorithm outperforms existing methods, improving the loss function by at least 16.87% and the convergence speed by at least 2 times, in various environments (MobileNet, AlexNet). Kechang Yang, Biao Hu 0001, Mingguo Zhao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Workload-Aware Scheduling of Real-Time Jobs in Cloud Computing to Minimize Energy ConsumptionabstractCloud computing is a powerful paradigm that can provide high-quality computation services to customers. Because its energy consumption has a large effect on its service price, this study investigates how to minimize the energy consumption while achieving adequate response times for requested computations. Such a problem is formulated as a nonlinear integer program. By deriving a state transition equation, this problem is transformed into an unconventional 0–1 knapsack problem, and dynamic programming is then used to solve it. In addition to this solution, we develop an energy-efficient job accommodation scheme that can manage dynamic jobs with varying frequencies throughout a day. Unlike existing studies that abruptly switch off old virtual machines and create new ones for upcoming jobs, this scheme tries to accommodate them with current virtual machines, and new virtual machines are not created unless necessary. Conversely, when the workload declines, jobs on energy-inefficient servers are moved to other servers, such that some energy-inefficient servers can be switched off to save energy. This scheme adjusts the computing power adaptively and smoothly without lowering the system’s quality of service. Experimental results demonstrate that the proposed solution outperforms a particle swarm optimizer and two other heuristics in terms of accommodating jobs and saving energy consumption. Biao Hu 0001, Yinbin Shi, Gang Chen 0023, Zhengcai Cao, MengChu Zhou |
IEEE Internet Things J. | 1 |
| 2023 | Rotation-Invariant Descriptors Learned with Circulant Convolution Neural NetworksabstractExtracting local features for accurate correspondences between image pairs is an essential basis for various computer vision tasks. Recent works have shown that deep neural networks (DNNs) have demonstrated promising performance in challenging environments. However, these state-of-the-art DNN-based approaches are not well suited for the scenario with geometry rotations due to their intrinsic deficiencies of square kernel structure. That is, square kernel structures in standard DNNs cannot fully identify the essentials of the rotations in geometry. To address this problem, we present RICNN, a novel deep learning framework that encodes invariance against the rotations in geometry explicitly into convolutional neural networks. Rather than using the square-shaped kernel structure, RICNN adopts sectorshaped convolutional kernels to achieve encoding invariance in all rotations. With the explicitness of such rotation encoding, RICNN enables the transfer of perspective DNN models to obtain rotation-invariant descriptions. Furthermore, we propose a novel multi-level hinge triplet loss function to strengthen the matching constraints against geometry rotations. Comprehensive experiments demonstrate the strong generalization ability of the RICNN descriptor on the HPatches dataset. Toward the rotation invariance evaluation, our method shows state-of-the-art results. Wenwei Lin, Chonghao Zhong, Xunpei Sun, Haitao Meng, Gang Chen 0023, Biao Hu 0001, Zonghua Gu 0001 |
ICTAI | 6 |
| 2023 | Online energy-efficient scheduling of DAG tasks on heterogeneous embedded platforms
Biao Hu 0001, Xincheng Yang, Mingguo Zhao |
J. Syst. Archit. | 1 |
| 2023 | A Hybrid Scheduling Framework for Mixed Real-Time Tasks in an Automotive System With Vehicular NetworkabstractAs vehicles integrate more and more autonomous driving functionalities, it becomes more and more important to use vehicular networks to fully guarantee the safety and real-time performance of on-board computing tasks. Current studies on vehicular networks pay much attention to the performance improvement of network communication and resource allocation, while ignoring the fact that automotive on-board computing tasks play a significant role in vehicle safety and need to be elegantly handled in vehicular networks. In this paper, we propose a hybrid scheduling framework for meeting all hard real-time task deadlines while minimizing soft real-time task deadline misses. In particular, the proposed scheduler is composed of some local schedulers and a global scheduler, where the former guarantees the schedulability of all hard real-time tasks, and the latter decides the assignment of soft real-time jobs dynamically online. Depending on the remaining processing capability of a vehicular network, arrival jobs are either assigned for further processing or discarded. An approach combining the utilization-based schedulability test and demand-supply analysis is proposed to effectively assign tasks to processors offline. To meet as many soft real-time task deadlines as possible, tasks' demand and supply bounds are computed at runtime, such that online execution information can be used to compensate for the scheduling loss with the worst-case assumption made by the offline test. Experimental results demonstrate that, compared to a genetic algorithm, our proposed approach needs far less computation to assign more tasks offline. The online scheduling also saves many soft real-time jobs that were to be dropped by the offline algorithm. Biao Hu 0001, Yinbin Shi, Zhengcai Cao, MengChu Zhou |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Adaptive Energy-Minimized Scheduling of Real-Time Applications in Vehicular Edge ComputingabstractVehicular edge computing is a promising new computing paradigm that has lower service latency and higher bandwidth than cloud computing. However, the geographical dispersion of edge computing resources and the high dynamics of vehicles pose many challenges to its service provision. Aiming to minimize the energy consumption of vehicular edge computing servers, this article presents an adaptive scheduling approach for handling dynamic real-time computing requests. An auction-bid scheme is developed for deciding the roadside unit (RSU) to respond to the computing request, where the computing request is auctioned and the RSU with the least energy consumption gets the bid. This scheme works in a decentralized model that effectively reduces its implementation complexity. To process the computing request modeled as a directed acyclic graph (DAG) application, the upward rank value is used to decompose a DAG into individual tasks, and a deadline-aware queue jump algorithm is proposed to assign them to servers' queues in a specific RSU. A group scheduling scheme is developed to assign several applications as a group, for the purpose of searching for a better schedule. Extensive experiments are carried out to compare our proposed approach to some other heuristic and state-of-the-art approaches, and the results confirm the benefits of our proposed approach in terms of minimizing system energy consumption and providing a quick response to the computing request. Biao Hu 0001, Yinbin Shi, Zhengcai Cao |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Workload-Aware Scheduling of Multiple- Criticality Real-Time Applications in Vehicular Edge Computing SystemabstractIn this article, we study the problem of designing an adaptive scheduling scheme for dynamic multiple-criticality real-time applications in vehicular edge computing systems. This scheduling problem is formulated as a mixed-integer nonlinear problem. We propose a workload-aware scheduling approach that not only guarantees the applicaitons' mixed-criticality schedulability but also adaptively manages their execution depending on their released frequencies at runtime. In particular, we first present the response time analysis for multiple-criticality applications in the edge computing system with different computing capability servers. Then, we derive a state-transition equation that makes the dynamic programming applicable to building an excellent schedule for one specific criticality-level mode. Such an approach is extended to the system's different criticality-level modes. For the purpose of increasing the quality-of-service toward low-critical applications, we leave low-critical applications to execute as much as possible at runtime, depending on their predicted frequencies. Extensive experimental results show the superiority of our proposed approaches in terms of improving the schedulability success rate and reducing the number of suspended applications online. Biao Hu 0001, Zhilei Yan, Mingguo Zhao |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Energy-Minimized Scheduling of Intermittent Real-Time Tasks in a CPU-GPU Cloud Computing PlatformabstractDue to the flexibility, availability, and scalability of cloud computing services, more and more users seek solutions via cloud computing techniques. A cloud computing platform often consists of a large number of infrastructures, and its energy consumption is a big problem. In this article, we study how to minimize the energy consumption of a cloud computing platform when handling some intermittent real-time tasks. Unlike previous works that abstract users’ submitted tasks as single computation jobs and process them using CPU, this work proposes using CPU and GPU to process intermittent real-time tasks that occur at irregular intervals and their released computation jobs must be completed within required time limits. The energy consumption minimization problem is formulated as an integer nonlinear programming problem that needs to decide on a task assignment plan and a specific resource allocation plan. To effectively solve this problem, we define a state that represents the optimal solution for a given set of tasks with a given amount of resources, as well as a value function that represents the value of a state. In this way, we derive a state-transition equation and develop a dynamic programming method to solve the problem. This method is also extended to handle tasks whose arrival time is dynamic and unpredictable. Experiments show that the proposed algorithm can effectively reduce energy consumption, while its computation time is quite low compared to some other greedy methods. Biao Hu 0001, Xincheng Yang, Mingguo Zhao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | A Slope-Adaptive Navigation Approach for Ground Mobile RobotsabstractThe 2-dimensional cost map has been widely used for the navigation of ground mobile robot. Although it is effective when the ground is flat, it becomes clumsy and ineffective when the ground has some slopes, where such slopes are often misjudged as the forbidden area by the cost map. For this reason, we propose a slope-adaptive navigation approach based on multilayer cost map in this paper. Instead of taking the point cloud of slope as obstacles, we actively construct a multi-layer cost map that takes slope information into the map in the stage of building environment map. A slope detection algorithm is developed to switch the cost map during the robot navigation. The slope is then considered as a passable road, only with extra cost. In the case that the slope leads the robot to a new floor, we adopt the Aruco code to switch the map information, such that the navigation can still keep working. Both simulation and real-world experimental results demonstrate the high effectiveness of our proposed approach. Biao Hu 0001, Mingyue Cui, Zhengcai Cao |
SMC | 1 |
| 2022 | Scheduling Real-Time Parallel Applications in Cloud to Minimize Energy ConsumptionabstractCloud computing has become an important paradigm in which scalable resources such as CPU, memory, disk and IO devices can be provided to users to remotely process their applications. In a cloud computing platform, energy consumption accounts for a significant cost portion. This article thus aims to present an energy-efficient scheduling algorithm for processing a user application with a real-time requirement. This problem is formulated as a non-linear mixed integer programming problem. We start with providing an optimal closed-form solution to its relaxation problem that aims to minimize the energy consumption without considering real-time requirements. To meet real-time requirements, we propose how to adjust task placement and resource allocation by making a good tradeoff between energy consumption and task execution time. Lastly, we find two equivalent optimal resource allocation strategies once task placement has been done. We then propose to adjust the start time of task execution such that an application’s completion time can be further shortened. Experimental results on two real-case enchmarks and extensive synthetic applications demonstrate that our proposed method finds a schedule that generally has 30 and 20 percent less energy consumption than enhancement heterogeneous earliest finish time (E-HEFT) and genetic algorithm, respectively. Besides, the proposed method has a higher rate to successfully find a feasible schedule than them, and its computation time is close to E-HEFT’s, but far less than the genetic algorithm's. Biao Hu 0001, Zhengcai Cao, MengChu Zhou |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | Safety-Guaranteed and Development Cost- Minimized Scheduling of DAG Functionality in an Automotive SystemabstractIt is important to sufficiently guarantee an automotive system’s safety, because otherwise terrible consequences may happen. Generally the safety in an automotive system includes two aspects: reliability and timeliness. Previous studies have proposed many approaches to how to improve them. However, few of them consider the development cost along with their improvement. In this study, we aim to propose a method that can build a safety-guaranteed and development cost-minimized schedule for functionality modeled as a directed acyclic graph running on an automotive system. Unlike previous studies that tightly couple the development cost minimization with other requirements together, we start by building a schedule with the minimum development cost by ignoring safety requirement. Then, reliability and real-time requirements are subsequently taken into consideration. Together with automotive safety integrity level decomposition options provided by International Standard called ISO 26262, the decomposition is evaluated for each task to improve its safety, and tasks are then successively chosen to adjust the schedule, such that its safety can be maximized with incurring the least extra development cost. This procedure continues until a schedule that meets safety requirement is built. Experiments on a real-life automotive benchmark and extensive synthetic functionality demonstrate that our proposed heuristics outperform the state-of-the-art heuristic algorithm, and a typical intelligent optimization algorithm. Biao Hu 0001, Shengjie Xu 0003, Zhengcai Cao, MengChu Zhou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Energy-minimized Scheduling of Real-time Parallel Workflows on Heterogeneous Distributed Computing SystemsabstractToday's large-scale parallel workflows are often processed on heterogeneous distributed computing platforms. From an economic perspective, computing resource providers should minimize the cost while offering high service quality. It has become well-recognized that energy consumption accounts for a large part of a computing system's total cost, and timeliness and reliability are two important service indicators. This work studies the problem of scheduling a parallel workflow that minimizes the system energy consumption under the constraints of response time and reliability. We first mathematically formulate this problem as a Non-linear Mixed Integer Programming problem. Since this problem is hard to solve directly, we present some highly-efficient heuristic solutions. Specifically, we first develop an algorithm that minimizes the schedule length while meeting reliability requirement, on top of which we propose a processor-merging algorithm and a slack time reclamation algorithm using a dynamic voltage frequency scaling (DVFS) technique to reduce energy consumption. The processor-merging algorithm tries to turn off some energy-inefficient processors such that energy consumption can be minimized. The DVFS technique is applied to scale down the processor frequency at both processor and task levels to reduce energy consumption. Experimental results on two real-life workflows and extensive synthetic parallel workflows demonstrate their effectiveness. Biao Hu 0001, Zhengcai Cao, MengChu Zhou |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | Probability-based Path Planning for Multi-Robot Systems with Stochastic Behavior in a Grid MapabstractFor the multi-robot path planning on a grid map, the widely adopted robot model assumes that its motion is deterministic once a path has been decided. However, this assumption is not quite realistic because some interferences such as noise, friction and inaccurate control input could disturb the robot motion, leading to a stochastic behavior. In this paper, we tackle the problem of planning a multi-robot path based on the robot probabilistic motion model. At the beginning, we model the robot action with several probability distributions, where the basic actions include going forward, turning left/right, going backward, and wait. We then extend A-star algorithm incorporating these actions such that an optimal path can be planned for a single robot. Based on this result, we apply conflict-based search to optimally plan path for a multi-robot system. Because probability calculation demands too much computation we simplify the conflict detection scheme and make it applicable for online practice. Biao Hu 0001, Zhengcai Cao |
SMC | 1 |
| 2021 | Rapid Detection of Blind Roads and Crosswalks by Using a Lightweight Semantic Segmentation NetworkabstractAchieving the high accuracy of blind roads and crosswalks recognition is important for blind guiding equipment to help blind people sense the surrounding environment. A lightweight semantic segmentation network is proposed to quickly and accurately segment blind roads and crosswalks in a complex road environment. Specifically, a lightweight network with depthwise separable convolution as a component is used as a basic module to reduce the number of parameters of the model and increase the speed of semantic segmentation. In order to ensure the segmentation accuracy of the network, we use a densely connected atrous spatial pyramid pooling module to extract feature information of different angles and context feature modules to enhance the effectiveness of different levels of feature information fusion. To verify the effectiveness of the proposed method, we collect and produce a data set from a real environment, which contains two objects of blind roads and crosswalks1. Experimental results demonstrate that, compared to some state-of-the-art approaches, the proposed approach greatly improves the segmentation speed, while achieving better or similar accuracy, which shows that the proposed approach provides a better basis for the application of devices for guiding the blind.1https://github.com/qweawq/Blind-road-and-crosswalk-dataset Zhengcai Cao, Biao Hu 0001, MengChu Zhou |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Heterogeneous Multi-Robot Path Planning Based on Probabilistic Motion ModelabstractAn important problem in multi-robot system is how to coordinate robot's motion such that each robot can complete its task without collision. Previous approaches assume that robots' motion is deterministic once their paths have been planned, which is however not realistic because random interferences such as noise, friction and inaccurate control input in real-life system could disturb the robot motion, leading to a stochastic behavior. In this paper, we take this stochastic behavior into account when planning the path for a heterogeneous multi-robot system. We assume that the motion time of a robot from a location to another can be modeled as a probability distribution. Every robot has its own probability distribution of motion time between any two neighbor locations. We develop a conflict-detection scheme for this model and propose using the conflict-based search algorithm via probability calculation to find the optimal path that minimizes the entire motion time. We also simplify this conflict detection such that our proposed approach is applicable online for a large-scale system. Experimental results demonstrate the high effectiveness of our proposed approaches. Biao Hu 0001, Zhengcai Cao |
SMC | 1 |
| 2020 | Minimizing Resource Consumption Cost of DAG Applications With Reliability Requirement on Heterogeneous Processor SystemsabstractResource consumption cost minimization is important in embedded systems due to their limited resource and high computation load. Previous approaches apply either meta-heuristic algorithms such as genetic algorithm or simple heuristics to minimize the resource consumption cost, which, however, are computationally expensive and sometimes ineffective. In this article, we propose how to schedule a directed acyclic graph application with reliability requirement in a heterogeneous embedded system. The scheduling aim is to minimize system resource consumption cost. We start by finding a quasi-optimal solution that minimizes the resource consumption cost in ignorance of reliability. Then, we introduce an indicator that price the tradeoff between reliability and resource consumption cost, based on which we build a scheme that progressively tunes this schedule toward improving its reliability. We also explore an approach that updates this indicator in a lightweight way. Compared to several state-of-the-art approaches such as MRCRG, experimental results demonstrate that our approach constantly outperforms them, and its superiority becomes more significant with the enhancement of reliability requirement. Biao Hu 0001, Zhengcai Cao |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Minimizing Task Completion Time of Prioritized Motion Planning in Multi-Robot SystemsabstractPrioritized motion planning is an effective approach to decouple the motion coordination in multi-robot systems, in which priority assignment plays an important role on its performance. Previous approaches apply either a randomized search or a simple rule to prioritize robot motions, which are often computationally expensive and sometimes ineffective. In this paper, we present an effective approach to schedule robots movement with the aim of minimizing their task completion time. This approach first plans a time-optimal motion for each robot individually, then roughly prioritizes robots in the descending order of their motion-time, i.e., a robot is prioritized the highest if it has the largest motion-time, and vice versa. Then, we propose a heuristic to fine-tune robot priorities under the principle that the tuned priority shortens the task completion time. We also develop an asynchronous decentralized algorithm to accelerate the priority-tuning process. Compared to previous priority assignment approaches, our approach needs less computation time to generate a shorter task completion time in a multi-robot system. Biao Hu 0001, Zhengcai Cao |
SMC | 1 |
| 2019 | Real-time gesture recognition based on feature recalibration network with multi-scale information
Zhengcai Cao, Biao Hu 0001, Meng Zhou 0006, Qinglin Li |
Neurocomputing | 3 |
| 2019 | FFOB: efficient online mode-switch procrastination in mixed-criticality systems
Biao Hu 0001, Lothar Thiele, Pengcheng Huang 0001, Kai Huang 0001, Christoph Griesbeck, Alois C. Knoll |
Real Time Syst. | 1 |
| 2019 | High-Speed Scene Flow on Embedded Commercial Off-the-Shelf SystemsabstractScene flow is an essential part of a stereo-based perception system for autonomous driving and mobile robotics. As in most of these platforms, the computing resource is limited but the computing requirement is high, embedded and parallelized algorithms are of vital importance for real-time tasks. This paper develops a cross-platform embedded scene flow algorithm by using an OpenCL (Open Computing Language) programming. Meanwhile, we propose a method to achieve a good performance by using a novel coarse-grained software pipeline for the embedded stream application. Experimental results show that the proposed algorithm can boost the average processing speed to 50 fps for different commercial off-the-shelf (COTS) hardware, including desktop graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and mobile phone platforms. For certain GPUs, the peak frame rates can also reach 1000 fps. By comparing the efficiency among the serial platform, we illustrate that with the help of OpenCL programming, COTS platforms can provide enough computing resources for the stereo-based perception algorithm. Long Chen 0005, Mingyue Cui, Feihu Zhang, Biao Hu 0001, Kai Huang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | Implementing and Parallelizing Real-time Lane Detection on Heterogeneous PlatformsabstractLane detection is a cardinal functionality in state-of-the-art Advanced Driver Assistant Systems (ADAS). However, it is still not straightforward to fulfill the real-time performance demand of processing High Definition (HD) images with high robustness and scalability. To address this problem, we propose an improved lane detection algorithm based on top-view image transformation and two-stage RANdom SAmple Consensus (RANSAC) model fitting. By virtue of off-line affine homography matrix adaption to bound an adaptive Region Of Interest (ROI) for subsequent on-line Warp Perspective Mapping (WPM) transformation, the algorithm can analyze arbitrary on-road videos and generate adaptive ROI without priori knowledge about camera parameter. To ensure the scalability, we present a comprehensive parallel design of the application in a heterogeneous system consisting of multi-core CPU, GPU and FPGA. We show in detail how the potentially parallel task loads are implemented and optimized so that they can be mapped to the most suitable processor so as to achieve optimal performance. Experimental results reveal that our improved algorithm can robustly process the video streams with a higher accuracy. Moreover, the heterogeneous executions are capable of processing HD 1920×1080 images with runtime performance of 81.6 fps and 47.9 fps, respectively, on an AMD FirePro W7100 GPU and a Terasic Arria 10 FPGA. Xiebing Wang, Christopher Kiwus, Canhao Wu, Biao Hu 0001, Kai Huang 0001, Alois C. Knoll |
ASAP | 4 |
| 2018 | Scheduling and shaping of complex task activations for mixed-criticality systemsabstractIn this paper, we present a new schedulability test that can cope with complex activation patterns for mixed-criticality systems. Under this analysis we proceed to present a shaping approach that can adaptively make use of system slack to improve the quality of service to less critical tasks. Compared with the state-of-the-art scheduling analysis, our scheduling analysis is more effective in handling the case that activation events can be backlogged and task deadlines can be arbitrary; and the shaping approach furthermore reduces the dropped jobs of less critical tasks without jeopardizing the guarantee to critical tasks. Extensive simulations and real-life deployment in Raspberry Pi 3 board confirm the effectiveness of our proposed schedulability test and shaping approach. Biao Hu 0001, Kai Huang 0001 |
ASP-DAC | 1 |
| 2018 | An Efficient and Time-Optimal Trajectory Generation Approach for Waypoints Under Kinematic Constraints and Error BoundsabstractThis paper presents an approach to generate the time-optimal trajectory for a robot manipulator under certain kinematic constraints such as joint position, velocity, acceleration, and jerk limits. This problem of generating a trajectory that takes the minimum time to pass through specified waypoints is formulated as a nonlinear constraint optimization problem. Unlike prior approaches that model the motion of consecutive waypoints as a Cubic Spline, we model this motion with a seven-segment acceleration profile, as this trajectory results in a shorter overall motion time while staying within the bounds of the robot manipulator's constraints. The optimization bottleneck lies in the complexity that increases exponentially with the number of waypoints. To make the optimization scale well with the number of waypoints, we propose an approach that has linear complexity. This approach first divides all waypoints to consecutive batches, each with an overlap of two waypoints. The overlapping waypoints then act as a bridge to concatenate the optimization results of two consecutive batches. The whole trajectory is effectively optimized by successively optimizing every batch. We conduct experiments on practical scenarios and trajectories generated by motion planners to evaluate the effectiveness of our proposed approach over existing state-of-the-art approaches. Jianjie Lin, Nikhil Somani, Biao Hu 0001, Markus Rickert 0001, Alois C. Knoll |
IROS | 3 |
| 2018 | Adas on Cots with OpenCL: A Case Study with Lane DetectionabstractThe concept of autonomous cars is driving a boost for car electronics and the size of automotive electronics market is foreseen to double by 2025. How to benefit from this boost is an interesting question. This article presents a case study to test the feasibility of using OpenCL as the programming language and Cots components as the underlying computing platforms for Adas development. For representative Adas applications, a scalable lane detection is developed that can tune the trade-off between detection accuracy and speed. Our OpenCL implementation is tested on 14 video streams from different data-sets with different road scenarios on 5 Cots platforms. We demonstrate that the Cots platforms can provide more than sufficient computing power for the lane detection in the meanwhile our OpenCL implementation can exploit the massive parallelism provided by the Cots platforms. Kai Huang 0001, Biao Hu 0001, Long Chen 0005, Alois C. Knoll, Zhihua Wang 0001 |
IEEE Trans. Computers | 2 |
| 2018 | EDF-VD Scheduling of Flexible Mixed-Criticality System With Multiple-Shot TransitionsabstractThe existing mixed-criticality (MC) real-time task models assume that once any high-criticality task overruns, all high-criticality jobs execute up to their most pessimistic WCET estimations simultaneously in a one-shot manner. This is very pessimistic in the sense of unnecessary resource overbooking. In this paper, we propose a more generalized mixed-critical real-time task model, called flexible MC model with multiple-shot transitions (FMC-MST), to address this problem. In FMC-MST, high-criticality tasks can transit multiple intermediate levels to handle less pessimistic overruns independently and to nonuniformly scale the deadline on each level. We develop a run-time schedulability analysis for FMC-MST under EDF-VD scheduling, in which a better tradeoff between the penalties of low-criticality tasks and the overruns of high-criticality tasks is achieved to improve the service quality of low-criticality tasks. We also develop a resource optimization technique to find resource-efficient level-insertion configurations for FMC-MST task systems under MC timing constraints. Experiments demonstrate the effectiveness of FMC-MST compared with the state-of-the-art techniques. Gang Chen 0023, Nan Guan, Biao Hu 0001, Wang Yi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Online workload monitoring with the feedback of actual execution time for real-time systemsabstractGuaranteeing the system workload within design bounds is a basic requirement for a real-time system. Design-time bounds are usually based on worst-case activation patterns and worst-case execution time. While using the worst-case assumptions for online monitoring can guarantee the system safety, it also introduces unexplored slacks due to tasks consuming less than their worst-case execution times. In this paper, we introduce a monitoring scheme with the feedback of actual execution time for real-time systems. By using this runtime feedback instead of offline assumptions, this monitoring scheme can accept events that are considered as violations offline, and thereby improve the system utilization. In the experiments of both MATLAB simulation and MicroC/OS-II running in a softcore processor implemented on an FPGA, different probability distributions of actual execution time are used in analyzing how much the benefit can be gained from the feedback scheme. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
DATE | 1 |
| 2016 | Minimizing peak temperature for pipelined hard real-time systems
Long Cheng 0007, Kai Huang 0001, Gang Chen 0023, Biao Hu 0001, Alois C. Knoll |
DATE | 4 |
| 2016 | A scalable lane detection algorithm on COTSs with OpenCL
Kai Huang 0001, Biao Hu 0001, Jan Botsch, Nikhil Madduri, Alois C. Knoll |
DATE | 2 |
| 2016 | On-the-fly fast overrun budgeting for mixed-criticality systemsabstractIn mixed-criticality scheduling, the widely assumed mode-switch scheme assumes that both high- and low-criticality tasks are schedulable when no tasks overrun (normal mode) and all high-criticality tasks are schedulable even when they overrun (critical mode, where low-criticality tasks are abandoned/degraded). However, this scheme triggers a mode-switch immediately after any task overruns, which can be abrupt and pessimistic. In this paper, we tackle dual-criticality systems scheduled by earliest-deadline-first, and propose light-weight mode-switch schemes that are effective in keeping the system "away" from the critical mode. Our main idea is to perform overrun budgeting for all tasks as a whole, by monitoring task executions and updating a common overrun budget. This way, the overrun budget is shared among all tasks, and adaptively replenished leveraging run-time information; consequently, mode-switch can be postponed as much as possible. Experimental results demonstrate that the proposed mode-switch schemes outperform existing solutions to a large extent, in reducing the abandoned jobs and mode-switch frequencies, as well as in increasing the time ratio that all tasks are scheduled in the system. Biao Hu 0001, Kai Huang 0001, Pengcheng Huang 0001, Lothar Thiele, Alois C. Knoll |
EMSOFT | 1 |
| 2016 | Evaluation and Improvements of Runtime Monitoring Methods for Real-Time Event StreamsabstractRuntime monitoring is of great importance as a safeguard to guarantee the correctness of system runtime behaviors. Two state-of-the-art methods, dynamic counters and l -repetitive function, were recently developed to tackle the runtime monitoring for real-time systems. While both are reported to be efficient in monitoring arbitrary events, the monitoring performance between them has not yet been evaluated. This article evaluates both methods in depth, to identify their strengths and weaknesses. New methods are proposed to efficiently monitor the many-to-one connections that are abstracted as AND and OR components on multiple inputs. Representative scenarios are used as our case studies to quantitatively demonstrate the evaluations. Both methods are implemented in hardware F pga . The timing overhead and resource usages of implementing the two methods are evaluated. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Adaptive Workload Management in Mixed-Criticality Systems
Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2015 | Evaluation of runtime monitoring methods for real-time event streamsabstractRuntime monitoring is of great importance as a safe guard to guarantee the correctness of system runtime behaviors. Two new methods, i.e., dynamic counters and l-repetitive function, are recently developed to tackle the runtime monitoring for hard real-time systems. This paper investigates in depth these two newly developed runtime monitoring methods, trying to evaluate and identify their strengths and weaknesses. Representative scenarios are used as our case studies to quantitatively demonstrate our comparisons. We also provide FPGA implementations and resource usages of both methods. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Alois C. Knoll |
ASP-DAC | 1 |
| 2015 | Adaptive runtime shaping for mixed-criticality systemsabstractThis paper investigates runtime shaping for mixed-criticality systems to increase the system QoS. Unlike the previous work in the literature that enforces an offline workload bound, an adaptively shaping approach is proposed where the incoming workload of the low-critical tasks is regulated by the actual demand of the high-critical tasks. This actual demand is adaptively updated using the historical arrival information of the high-critical tasks and thus can maximize the runtime QoS of low-critical tasks. To reduce the online overheads of computing the workload demand, a lightweight scheme with the complexity of O(n log(m)) is developed. Experiments are also provided to demonstrate the effectiveness and efficiency of our approach. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
EMSOFT | 1 |
| 2014 | Abstract: Shared L2 Cache Management in Multicore Real-Time SystemabstractIn multicore system, shared cache interference has been recognized as one of the major factors that degrade the average performance as well as predictability of system. How to manage the shared cache in order to optimize the system performance while guaranteeing the system predictability is still an open issue. State-of-the-art techniques on this topic use page coloring to partition the shared cache at OS level. In this paper, we present a shared cache management scheme for multicore system. This shared cache management scheme supports way-based cache partitioning at hardware level, building task-level time-triggered reconfigurable-cache multicore system. We evaluated the proposed scheme w.r.t. different numbers of cores and cache modules and prototyped the constructed MPSoCs on FPGA. Gang Chen 0023, Biao Hu 0001, Kai Huang 0001, Alois C. Knoll, Di Liu 0002 |
FCCM | 2 |
| 2012 | The optimization of spring stiffness for passive dynamic walkerabstractOn a passive dynamic walker, we find that an appropriate spring placed on the mass center of two legs can greatly improve the walker's walking performance, which includes its walking speed, disturbance rejection ability and step length. However, the extent to which spring influences these properties is unknown and how to choose an appropriate spring stiffness to achieve optimal walking performance still needs to be discussed. In this paper, we present a study on the effect of spring on these properties by constructing a synthesize index P to assess the walking performance. Through numerical simulation on three models, we find the walker with spring has a better walking performance over a pure passive dynamic walker and the extension spring on the mass center of two legs can improve a walker's overall walking performance much better. Biao Hu 0001, Mingguo Zhao |
IROS | 1 |