Gang Chen 0023

dblp:67/6383-23 · DBLP profile ↗
← Back
80ranked-venue papers
14as first author
56since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 51 · 13 first-author · 31 since 2021Artificial intelligence and machine learning · 16 · 16 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring Surround-View Fisheye Camera 3D Object Detection
abstract
In this work, we explore the technical feasibility of implementing end-to-end 3D object detection (3DOD) with surround-view fisheye camera system. Specifically, we first investigate the performance drop incurred when transferring classic pinhole-based 3D object detectors to fisheye imagery. To mitigate this, we then develop two methods that incorporate the unique geometry of fisheye images into mainstream detection frameworks: one based on the bird's-eye-view (BEV) paradigm, named FisheyeBEVDet, and the other on the query-based paradigm, named FisheyePETR. Both methods adopt spherical spatial representations to effectively capture fisheye geometry. In light of the lack of dedicated evaluation benchmarks, we release Fisheye3DOD, a new open dataset synthesized using CARLA and featuring both standard pinhole and fisheye camera arrays. Experiments on Fisheye3DOD demonstrate that our fisheye-compatible modeling improves accuracy by up to 6.2% compared to baseline methods.
Changcai Li, Wenwei Lin, Zuoxun Hou, Gang Chen 0023, Wei Zhang 0009, Wei-Shi Zheng 0001
AAAI4
2026 DynamicRTL: RTL Representation Learning for Dynamic Circuit Behavior
abstract
There is a growing body of work on using Graph Neural Networks (GNNs) to learn representations of circuits, focusing primarily on their static characteristics. However, these models fail to capture circuit runtime behavior, which is crucial for tasks like circuit verification and optimization. To address this limitation, we introduce DR-GNN (DynamicRTL-GNN), a novel approach that learns RTL circuit representations by incorporating both static structures and multi-cycle execution behaviors. DR-GNN leverages an operator-level Control Data Flow Graph (CDFG) to represent Register Transfer Level (RTL) circuits, enabling the model to capture dynamic dependencies and runtime execution. To train and evaluate DR-GNN, we build the first comprehensive dynamic circuit dataset, comprising over 6,300 Verilog designs and 63,000 simulation traces. Our results demonstrate that DR-GNN outperforms existing models in branch hit prediction and toggle rate prediction. Furthermore, its learned representations transfer effectively to related dynamic circuit tasks, achieving strong performance in power estimation and assertion prediction.
Yunhao Zhou, Yi Liu 0081, Zhengyuan Shi, Lingwei Yan, Gang Chen 0023, Qiang Xu 0001, Guojie Luo
AAAI10
2026 RheoSparse: Exploring Finer Grained Structured Sparsity for Small Language Models
abstract
Small Language Models (SLMs) are designed for efficient on-device deployment, but compressing them without significant accuracy loss remains challenging. Structured sparsity methods like N:M pruning, which removes N parameters out of every M, are widely used to improve hardware efficiency. However, their rigid patterns, such as the commonly adopted 4:8 format, often degrade SLM performance, since these models are more sensitive to parameter removal than larger ones. We observe that finer-grained patterns like N:64 can better preserve accuracy under the same sparsity budget, yet current inference systems do not efficiently support them, especially during token generation. Furthermore, applying such fine-grained sparsity uniformly across all layers is suboptimal, as different layers respond differently to pruning. To address this, we propose RheoSparse. First, we use coarse-to-fine evolutionary search to assign sparsity levels across layers under a global budget. Second, we design a highly-optimized Sparse Matrix-Vector Multiplication (SpMV) kernel that efficiently supports arbitrary structured sparsity patterns during token generation. For example, on Qwen2.5-1.5B, RheoSparse reduces perplexity (PPL) by 33.09% and improves downstream task performance by 9.3% compared to 4:8 sparsity, while maintaining the same parameter count. Furthermore, our SpMV kernel outperforms the state-of-the-art sparse kernel SpInfer by up to 49.6%.
Jianing Zheng, Gang Chen 0023
DATE2
2026 An efficient and low-latency attention model for event denoising
Gang Chen 0023
Neural Networks2
2026 Rotation-invariant representation learning by sector convolution neural networks
Wenwei Lin, Xunpei Sun, Chonghao Zhong, Haitao Meng, Gang Chen 0023, Bingxian Zhang, Zonghua Gu 0001
Pattern Recognit.5
2026 Terafly: A Multinode FPGA-Based Accelerator Design for Efficient Cooperative Inference in LLMs
abstract
In this paper, we propose Terafly, a multi-node accelerator design tailored for efficient Large Language Model (LLM) deployment and inference. Conventional accelerator architectures struggle to effectively handle both the prefill and decode stages during inference. To address this limitation, we introduce a hybrid spatial-temporal architecture that combines the high-throughput advantages of spatial architectures with the flexibility of temporal architectures, enabling it to accommodate the diverse inference patterns of LLMs. In addition, we propose a generation framework to streamline the customization of our LLM-friendly design for various deployment scenarios. Within this framework, users can specify their requirements such as model type, target platform, and performance goals. The framework then generates multiple accelerator nodes and maps them to distinct Super Logic Regions (SLRs) within a single FPGA, enabling cooperative inference under a model parallelism scheme. Through experiments, our generated accelerator can be easily deployed on both Alveo U250 and U50lv cards, serving models ranging from OPT-350M to OPT-1.3B under various performance settings. Notably, when running OPT-1.3B using the generated dual-node accelerator on a single Alveo U50lv card, we achieve an average 1.1x speed-up and a 3.4x improvement in energy efficiency compared to the Nvidia A100 GPU.
Jianing Zheng, Gang Chen 0023, Libo Huang 0002, Xin Lou 0001, Wei-Shi Zheng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 Pushing Physical Limits and Uncovering Motion Templates of Spine-Based Quadruped Locomotion via Reinforcement Learning
abstract
Flexible spines are critical to the remarkable agility and speed of animals. Translating this biological advantage to quadruped robots presents a significant control challenge, particularly in coordinating the spine and limbs for maximal velocity. In this work, we utilize reinforcement learning (RL) to develop high-speed locomotion for a bioinspired mouse robot with a lateral flexible spine. The resulting controller achieves motor performance that demonstrably surpasses non-spined and model-based methods. More importantly, our analysis reveals the principles behind this performance: the emergence of two distinct motion templates. For high-speed walking, the robot learns a “whip-like” spinal oscillation to increase leg swing frequency, while for agile turning, it adopts a dynamic “bend-and-straighten” pattern. These findings demonstrate the capability of RL to not only generate high-performance controllers but also to produce emergent strategies that, upon analysis, reveal underlying principles of high-speed, spine-driven locomotion.
Zhenshan Bing, Yulong Xiao, Yuhong Huang, Long Cheng 0007, Biao Hu 0001, Gang Chen 0023, Yang Gao 0001, Fuchun Sun 0001, Kai Huang 0001, Alois C. Knoll
IEEE Trans. Robotics7
2025 M2Flow: A Motion Information Fusion Framework for Enhanced Unsupervised Optical Flow Estimation in Autonomous Driving
abstract
Estimating optical flow in occluded regions is a crucial challenge in unsupervised settings. In this work, we introduce M2Flow, a novel framework for unsupervised optical flow estimation that integrates motion information from multiple frames to address occlusions. By modeling inter-frame motion information and employing Motion Information Propagation (MIP) module, M2Flow effectively propagates and integrates motion information across frames, while concurrently estimating bidirectional optical flows for multiple frames. In addition, to handle occlusions across multiple frames, we provide two augmentation modules specifically designed for our multi-frame model to further refine optical flow. The experiments on KITTI and Sintel datasets demonstrate that M2Flow outperforms other state-of-the-art unsupervised approaches, especially in solving occlusions.
Xunpei Sun, Gang Chen 0023, Zuoxun Hou
AAAI2
2025 RISC-TAE: Instruction Set Extension for Transformer Model Acceleration
abstract
This paper proposes RISC-TAE(Transformer Acceleration Engine), a RISC-V instruction set extension with microarchitectural co-design, to meet the performance and energy-efficiency requirements of Transformer models in edge scenes.The design integrates operator-specialized computing units (GEMM/Softmax/GELU) with hardware-managed dataflow orchestration through custom RISC-V instructions, effectively resolving energy efficiency bottlenecks and memory access fragmentation in existing solutions. Experiments demonstrate that RISC-TAE achieves 23.35× and 8.01× speedups over scalar processor CVA6 and vector processor ARA respectively for BERT inference, while outperforming RISC-VTF by 1.4×, providing a scalable solution for edge Transformer deployment.
Yanping Shao, Zhouquan Liu, Junbo Tie, Gang Chen 0023, Libo Huang 0002
CASES7
2025 OmniStereo: Real-time Omnidireactional Depth Estimation with Multiview Fisheye Cameras
abstract
Fast and reliable omnidirectional 3D sensing is essential to many applications such as autonomous driving, robotics and drone navigation. While many well-recognized methods have been developed to produce high-quality omnidirectional 3D information, they are too slow for real-time computation, limiting their feasibility in practical applications. Motivated by these shortcomings, we propose an efficient omnidirectional depth sensing framework, called OmniStereo, which generates high-quality 3D information in real-time. Unlike prior works, OmniStereo employs Cassini projection to simplify the photometric matching and introduces a lightweight stereo matching network to minimize computational overhead. Additionally, OmniStereo proposes a novel fusion method to handle depth discontinuities and invalid pixels complemented by a refinement module to reduce mapping-introduced errors and recover fine details. As a result, OmniStereo achieves state-of-the-art (SOTA) accuracy, surpassing the second-best method over 32% in MAE, while maintaining real-time efficiency. It operates more than 16.5× faster than the second-best method in accuracy on TITAN RTX, achieving 12.3 FPS on embedded device Jetson AGX Orin, underscoring its suitability for real-world deployment. The code is available at https://github.com/DengJiaxi1/OmniStereo.
Jiaxi Deng, Yushen Wang, Haitao Meng, Zuoxun Hou, Yi Chang 0002, Gang Chen 0023
CVPR6
2025 Late Breaking Results: AFS: Improving Accuracy of Quantized Mamba via Aggressive Forgetting Strategy
abstract
Mamba overcomes the quadratic complexity problem inherent in Transformer models while maintaining comparable contextual modeling capabilities. However, Mamba-based foundation models encounter challenges in achieving efficient inference on resource-constrained devices, primarily due to their considerable size. Model compression techniques, such as linear quantization, offer a viable solution to this problem. Nevertheless, the introduction of significant outliers during Mamba's past-state forgetting process can lead to a notable decrease in accuracy when employing linear quantization. To overcome these challenges, this paper introduces the Aggressive Forgetting Strategy (AFS), an innovative and efficient algorithm designed to mitigate the quantization issues caused by outliers in the state forgetting mechanism. AFS incorporates a computation-free approach for handling outliers, facilitating both efficient and accurate linear quantization for Mamba, which is essential for applications in resource-constrained scenarios. By leveraging the AFS strategy, Mamba can perform more efficient inference, while significantly improving the accuracy by up to 21.5× compared to conventional methods.
Zhouquan Liu, Libo Huang 0002, Ling Yang 0008, Gang Chen 0023, Yongwen Wang
DATE4
2025 LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
abstract
In this paper, we propose LoopLynx, a scalable dataflow architecture for efficient LLM inference that optimizes FPGA usage through a hybrid spatial-temporal design. The design of LoopLynx incorporates a hybrid temporal-spatial architecture, where computationally intensive operators are implemented as large dataflow kernels. This achieves high throughput similar to spatial architecture, and organizing and reusing these kernels in a temporal way together enhances FPGA peak performance. Furthermore, to overcome the resource limitations of a single device, we provide a multi-FPGA distributed architecture that overlaps and hides all data transfers so that the distributed accelerators are fully utilized. By doing so, LoopLynx can be effectively scaled to multiple devices to further explore model parallelism for large-scale LLM inference. Evaluation of GPT-2 model demonstrates that LoopLynx can achieve comparable performance to state-of-the-art single FPGA-based accelerations. In addition, compared to Nvidia A100, our accelerator with a dual-FPGA configuration delivers a 2.52x speed-up in inference latency while consuming only 48.1% of the energy.
Jianing Zheng, Gang Chen 0023
DATE2
2025 SONet: Towards Practical Online Neural Network for Enhancing Hard-to-Predict Branches
Zhenxuan Xiong, Libo Huang 0002, Ling Yang 0008, Hui Guo 0004, Songwen Pei, Gang Chen 0023, Yongwen Wang
Euro-Par (2)8
2025 Low-Cost Approximate Floating-Point Multiplier Design Based on SSA and Sparse Processing
Gang Chen 0023, Qianmin Yang, Yongwen Wang, Libo Huang 0002
ICA3PP (6)3
2025 Adaptive Receptive Field Convolution for Top-view Fisheye Images Segmentation
abstract
Top-view fisheye cameras are cost-efficient devices used for omnidirectional perception. However, the wide field of view (FOV) of these cameras causes significant image distortion, while the top-view setup introduces rotational symmetry, resulting in the degradation of performance of the standard convolution neural network when processing fisheye images. To address this problem, we present ARFC, a novel method called adaptive receptive field convolution, specifically designed to extract rotation- and scale-equivariant representations from top-view fisheye images. Unlike traditional orientation-static convolutions, ARFC incorporates an adaptive rotating kernel (ARK) to separate rotation from distortion and capture rotational equivariant features. Additionally, a multi-scale fusion module (MSFM) is implemented to combine scale-distorted features. Experimental evaluations conducted on THEODORE for segmentation tasks illustrate the superior performance of ARFC compared to the current state-of-the-art methods. Code is available at: https://github.com/LinMenwill/ARFC.git.
Wenwei Lin, Gang Chen 0023, Changcai Li
ICASSP2
2025 FastPoint: Super Lightweight Keypoint Detection, Description and Depth Estimation Framework
abstract
Keypoint detection, description and depth estimation are fundamental parts for many computer vision tasks, like image matching and simultaneous localization and mapping (SLAM) systems. Early methods are based on human heuristics, extracted handcrafted feature from image pixel, which may lead to unstable and confusable results. Recent years learned-based methods have achieved remarkable results in this field. However, most of these methods cannot ensure real-time application, and the lightweight deployment of network models remains challenging. Additionally, the black-box nature of neural network operations lacks interpretability, which constrains further application expansion. This paper proposed a super lightweight framework, FastPoint, used shallow binary neural network combine with handcrafted differentiable module FastLayer, which outputs key-points location, descriptor and depth estimation simultaneously. Our FastPoint is trained on data synthetically created form KITTI Raw dataset and evaluated on HPatches. The results demonstrate that our method achieves competitive performance compared to existing approaches while significantly reducing computational complexity.
Chonghao Zhong, Haitao Meng, Wenwei Lin, Gang Chen 0023, Alois C. Knoll
ICTAI4
2025 Pixel-wise Rain Error Distillation Framework for Robust Sparse Matching in Rainy Scene
abstract
Sparse local feature matching has seen significant progress with the rise of deep learning. Unfortunately, the performance could suffer from significant degradation under rainy conditions. Most of the existing image-level methods focus on pixel-level similarity rather than feature discrimination, limiting their effectiveness for matching tasks. Meanwhile, contrastive learning offers a promising feature-level solution to domain discrepancies but often overlooks pixel-wise contrast and explicit geometrical constraints, resulting in unstable feature detection and description in rainy conditions. To address these issues, we propose an effective contrastive learning framework, called PRED, for local feature matching under rainy conditions, which leverages the inherent differences between geometrical and rainy discrepancies as an additional constraint for deraining contrastive learning, while preserving geometrical invariance in feature detection and description. To achieve this, the PRED framework integrates an Error-Aware Rainy Distillation (EARD) module and Pixel-wise Selective Contrast (PSC) module to explore mutual properties of pixel-wise geometrical- and rainy-invariant local features between clean and rainy images. Extensive experiments on the rainy-distorted HPatches and MegaDepth datasets demonstrate the effectiveness of the PRED framework, achieving state-of-the-art performance in sparse local feature matching under rainy conditions.
Wenwei Lin, Pinhuan Wei, Gang Chen 0023
IJCNN3
2025 Brief Announcement: LCTree: A Fast Hardware BVH Constructor for Real-Time Ray Tracing
abstract
Unlike traditional rasterization rendering, ray tracing is a groundbreaking technology that has revolutionized the realistic rendering of images, marking a significant leap forward. However, achieving real-time ray tracing in dynamic scene applications remains a challenging task. This difficulty arises primarily from the substantial technical bottlenecks related to the frequent need for reconstructing or incrementally updating acceleration structures essential for efficient ray calculations.
Run Yan, Su Yin, Hui Guo 0004, Yongwen Wang, Gang Chen 0023, Nong Xiao 0001, Libo Huang 0002
SPAA5
2025 A variable-gain fixed-time convergent neurodynamic network for time-variant quadratic programming under unknown noises
Biao Song, Tinghe Hong, Weibing Li, Gang Chen 0023, Yongping Pan 0001, Kai Huang 0001
Neurocomputing4
2025 Adverse Weather Optical Flow: Cumulative Homogeneous-Heterogeneous Adaptation
abstract
Optical flow has made great progress in clean scenes, while suffers degradation under adverse weather due to the violation of the brightness constancy and gradient continuity assumptions of optical flow. Typically, existing methods mainly adopt domain adaptation to transfer motion knowledge from clean to degraded domain through one-stage adaptation. However, this direct adaptation is ineffective, since there exists a large gap due to adverse weather and scene style between clean and real degraded domains. Moreover, even within the degraded domain itself, static weather (e.g., fog) and dynamic weather (e.g., rain) have different impacts on optical flow. To address above issues, we explore synthetic degraded domain as an intermediate bridge between clean and real degraded domains, and propose a cumulative homogeneous-heterogeneous adaptation framework for real adverse weather optical flow. Specifically, for clean-degraded transfer, our key insight is that static weather possesses the depth-association homogeneous feature which does not change the intrinsic motion of the scene, while dynamic weather additionally introduces the heterogeneous feature which results in a significant boundary discrepancy in warp errors between clean and degraded domains. For synthetic-real transfer, we figure out that cost volume correlation shares a similar statistical histogram between synthetic and real degraded domains, benefiting to holistically aligning the homogeneous correlation distribution for synthetic-real knowledge distillation. Under this unified framework, the proposed method can progressively and explicitly transfer knowledge from clean scenes to real adverse weather. In addition, we further collect a real adverse weather dataset with manually annotated optical flow labels and perform extensive experiments to verify the superiority of the proposed method.
Hanyu Zhou, Yi Chang 0002, Zhiwei Shi 0001, Wending Yan, Gang Chen 0023, Yonghong Tian 0001, Luxin Yan
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Fault-tolerant DAG Scheduling with Runtime Reconfiguration on Multicore Real-Time Systems
abstract
Fault tolerance and real-time performance are two essential goals for directed acyclic graph (DAG) scheduling. However, the redundant tasks to compensate for faults can significantly prolong the completion time of a DAG, i.e., the makespan. In addition, the unpredictable runtime failure status, i.e., a fault may or may not occur during task execution, results in a huge difference between the offline schedule and the actual execution. Existing list scheduling methods can support fault-tolerant execution with all redundant tasks taken into account before runtime. However, such methods cannot effectively reduce the actual makespan as conditional execution of redundant tasks is not considered during scheduling. To address the above issues, this paper proposes a fault-tolerant DAG scheduling method with a runtime reconfiguration facility to optimize the actual makespan. First, the fault-aware makespan is optimized by a fine-grained offline scheduling method considering the worst-case scenario and the runtime flexibility. Then, the runtime reconfiguration mechanism safely moves the influential nodes ahead using the additional time interval to minimize the actual makespan. The experimental results indicate that the proposed method outperforms the state-of-the-art (SOTA) methods in terms of schedulability and actual makespan.
Yuanhai Zhang, Shuai Zhao 0004, Gang Chen 0023, Kai Huang 0001
ASAP3
2024 QuickTree: A Fast Hardware BVH Construction Engine
abstract
Ray tracing has emerged as a powerful technique for generating visually stunning and realistic images compared to rasterization. With the continuous advancements in computer hardware, modern GPUs have integrated specialized ray tracing acceleration units to enhance rendering capabilities further. However, achieving realtime ray tracing presents a challenge in dynamic scenes, where spatial data structures used for accelerated rendering must be reconstructed or updated when there are changes in the scene primitives. This paper introduces QuickTree, a novel Bounding Volume Hierarchy (BVH) construction engine based on the linear BVH (LBVH) optimization algorithm. QuickTree addresses the challenge of dynamic scenes support by employing a highly parallel and pipelined system design. This innovative approach ensures fast construction speed. QuickTree demonstrates significant performance improvements. Compared to the currently fastest MergeTree, it has increased construction speed by 10% and reduced area by 45% compared to RayCore, which has the smallest chip area.
Yin Su, Hui Guo 0004, Run Yan, Yongwen Wang, Nong Xiao 0001, Gang Chen 0023, Libo Huang 0002
CF7
2024 Real-time Stereo-based 3D Object Detection for Streaming Perception
abstract
The ability to promptly respond to environmental changes is crucial for the perception system of autonomous driving. Recently, a new task called streaming perception was proposed. It jointly evaluate the latency and accuracy into a single metric for video online perception. In this work, we introduce StreamDSGN, the first real-time stereo-based 3D object detection framework designed for streaming perception. StreamDSGN is an end-to-end framework that directly predicts the 3D properties of objects in the next moment by leveraging historical information, thereby alleviating the accuracy degradation of streaming perception. Further, StreamDSGN applies three strategies to enhance the perception accuracy: (1) A feature-flow-based fusion method, which generates a pseudo-next feature at the current moment to address the misalignment issue between feature and ground truth. (2) An extra regression loss for explicit supervision of object motion consistency in consecutive frames. (3) A large kernel backbone with a large receptive field for effectively capturing long-range spatial contextual features caused by changes in object positions. Experiments on the KITTI Tracking dataset show that, compared with the strong baseline, StreamDSGN significantly improves the streaming average precision by up to 4.33%. Our code is available at https://github.com/weiyangdaren/streamDSGN-pytorch.
Changcai Li, Zonghua Gu 0001, Gang Chen 0023, Libo Huang 0002, Wei Zhang 0092
NeurIPS3
2024 ROTA-I/O: Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
abstract
In safety-critical systems, timing accuracy is the key to achieving precise I/O control. To meet such strict timing requirements, dedicated hardware assistance has recently been investigated and developed. However, these solutions are often fragile, due to unforeseen timing defects. In this paper, we propose a robust and timing-accurate I/O co-processor, which manages I/O tasks using Execution Time Servers (ETSs) and a two-level scheduler. The ETSs limit the impact of timing defects between tasks, and the scheduler prioritises ETSs based on their importance, offering a robust and configurable scheduling infrastructure. Based on the hardware design, we present an ETS-based timing-accurate I/O schedule, with the ETS parameters configured to further enhance robustness against timing defects. Experiments show the proposed I/O control method outperforms the state-of-the-art method in terms of timing accuracy and robustness without introducing significant overhead.
Zhe Jiang 0004, Shuai Zhao 0004, Xin Si, Gang Chen 0023, Nan Guan
RTSS5
2024 Workload-Aware Scheduling of Real-Time Jobs in Cloud Computing to Minimize Energy Consumption
abstract
Cloud computing is a powerful paradigm that can provide high-quality computation services to customers. Because its energy consumption has a large effect on its service price, this study investigates how to minimize the energy consumption while achieving adequate response times for requested computations. Such a problem is formulated as a nonlinear integer program. By deriving a state transition equation, this problem is transformed into an unconventional 0–1 knapsack problem, and dynamic programming is then used to solve it. In addition to this solution, we develop an energy-efficient job accommodation scheme that can manage dynamic jobs with varying frequencies throughout a day. Unlike existing studies that abruptly switch off old virtual machines and create new ones for upcoming jobs, this scheme tries to accommodate them with current virtual machines, and new virtual machines are not created unless necessary. Conversely, when the workload declines, jobs on energy-inefficient servers are moved to other servers, such that some energy-inefficient servers can be switched off to save energy. This scheme adjusts the computing power adaptively and smoothly without lowering the system’s quality of service. Experimental results demonstrate that the proposed solution outperforms a particle swarm optimizer and two other heuristics in terms of accommodating jobs and saving energy consumption.
Biao Hu 0001, Yinbin Shi, Gang Chen 0023, Zhengcai Cao, MengChu Zhou
IEEE Internet Things J.3
2024 Timing-accurate scheduling and allocation for parallel I/O operations in real-time systems
Yuanhai Zhang, Shuai Zhao 0004, Gang Chen 0023, Haoyu Luo, Kai Huang 0001
J. Syst. Archit.3
2024 A Low-Cost Floating-Point Dot-Product-Dual-Accumulate Architecture for HPC-Enabled AI
abstract
The dot-product$\sum _{i=1}^{N} A_{i}\times B_{i}$is one of the most frequently used operations for a wide variety of high-performance computing (HPC) and artificial intelligence (AI) applications. However, for large-scale algorithms, such as acrshort GEMM and acrshort FFT, independent additions are necessary to accumulate the results of length-limited dot-product in order to form the final result, thus increasing latency and overhead. Hence, we proposed a dot-product-dual-accumulate (DPDAC) architecture capable of performing$\left({\sum _{i=1}^{N=1,2,4} A_{i}\times B_{i} + \sum _{j=1}^{M=1,2} C_{j}}\right)$on a wide range of formats. The proposed architecture supports both single-path and dual-path execution. The single path is designed for performing acrshort DP acrshort FMA or DPDAC of lower formats, while dual-path supports parallel operations for single-precision (SP) addition and 2-term SP or acrshort TF32 dot-product or 4-term acrshort HP or BF16 dot-product. Moreover, numerical precision conversion is also supported by the proposed architecture, allowing for the conversion of numbers to higher or lower formats. The proposed DPDAC has been demonstrated to significantly reduce the overhead in comparison to discrete designs that utilize multiple single-mode acrshort FP units to achieve the same functionalities. Furthermore, when compared to the state-of-the-art multiple-precision designs, the proposed architecture has been shown to support a wide range of formats and a greater variety of operations with lower costs.
Hongbing Tan, Libo Huang 0002, Hui Guo 0004, Qianming Yang, Li Shen 0007, Gang Chen 0023, Liquan Xiao, Nong Xiao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 An Efficient Single Image De-Raining Model With Decoupled Deep Networks
abstract
Single image de-raining is an emerging paradigm for many outdoor computer vision applications since rain streaks can significantly degrade the visibility and render the function compromised. The introduction of deep learning (DL) has brought about substantial advancement on de-raining methods. However, most existing DL-based methods use single homogeneous network architecture to generate de-rained images in a general image restoration manner, ignoring the discrepancy between rain location detection and rain intensity estimation. We find that this discrepancy would cause feature interference and representation ability degradation problems which significantly affect de-raining performance. In this paper, we propose a novel heterogeneous de-raining architecture aiming to decouple rain location detection and rain intensity estimation (DLINet). For these two subtasks, we provide dedicated network structures according to their differential properties to meet their respective performance requirements. To coordinate the decoupled subnetworks, we develop a high-order collaborative network learning the dynamic inter-layer interactions between rain location and intensity. To effectively supervise the decoupled subnetworks during training, we propose a novel training strategy that imposes task-oriented supervision using the label learned via joint training. Extensive experiments on synthetic datasets and real-world rainy scenes demonstrate that the proposed method has great advantages over existing state-of-the-art methods.
Wencheng Li, Gang Chen 0023, Yi Chang 0002
IEEE Trans. Image Process.2
2023 Unsupervised Hierarchical Domain Adaptation for Adverse Weather Optical Flow
abstract
Optical flow estimation has made great progress, but usually suffers from degradation under adverse weather. Although semi/full-supervised methods have made good attempts, the domain shift between the synthetic and real adverse weather images would deteriorate their performance. To alleviate this issue, our start point is to unsupervisedly transfer the knowledge from source clean domain to target degraded domain. Our key insight is that adverse weather does not change the intrinsic optical flow of the scene, but causes a significant difference for the warp error between clean and degraded images. In this work, we propose the first unsupervised framework for adverse weather optical flow via hierarchical motion-boundary adaptation. Specifically, we first employ image translation to construct the transformation relationship between clean and degraded domains. In motion adaptation, we utilize the flow consistency knowledge to align the cross-domain optical flows into a motion-invariance common space, where the optical flow from clean weather is used as the guidance-knowledge to obtain a preliminary optical flow for adverse weather. Furthermore, we leverage the warp error inconsistency which measures the motion misalignment of the boundary between the clean and degraded domains, and propose a joint intra- and inter-scene boundary contrastive adaptation to refine the motion boundary. The hierarchical motion and boundary adaptation jointly promotes optical flow in a unified framework. Extensive quantitative and qualitative experiments have been performed to verify the superiority of the proposed method.
Hanyu Zhou, Yi Chang 0002, Gang Chen 0023, Luxin Yan
AAAI3
2023 ViraEye: An Energy-Efficient Stereo Vision Accelerator with Binary Neural Network in 55 nm CMOS
abstract
This paper presents the ViraEye chip, an energy-efficient stereo vision accelerator based on the binary neural network (BNN) to achieve high-quality and real-time stereo estimation. This stereo vision accelerator is designed as an end-to-end full pipeline architecture where all processing procedures, including stereo rectification, BNNs, cost aggregation and post-processing, are implemented on the ViraEye chip. ViraEye allows for top level pipelining between accelerator and image sensors, and no external CPUs or GPUs are required. The accelerator is implemented using SMIC 55nm CMOS technology and achieves top-performing processing speed in terms of million disparity estimations per second (MDE/s) metric among the existing ASIC in the open literature.
Gang Chen 0023, Qian Huang 0004
ASP-DAC2
2023 EDS-SLAM: An Energy-Efficient Accelerator for Real-Time Dense Stereo SLAM with Learned Feature Matching
abstract
Simultaneous Localization And Mapping (SLAM) is a fundamental component used in many applications, such as robotic navigation and augmented reality. To achieve on-line navigation on resource-constrained platforms, a fair amount of research has been devoted to develop real-time SLAM accelerators. However, most existing accelerators have the following limitations: (1) relying on hand-crafted visual features, which cannot offer robust feature association in complex environments; (2) only recovering sparse points of the observed scene, which are not adequate for high-level tasks such as obstacle avoidance and path planning. In this paper, we present EDS-SLAM, an energy-efficient and robust architecture for real-time dense stereo SLAM system based on learned feature matching to achieve real-time visual localization and dense mapping reconstruction. To achieve both resource efficiency and perception robustness, EDS-SLAM utilizes a share-used binary neural network (BNN) architecture to perform not only key-points association for computing robust camera poses, but also stereo matching for obtaining high-quality mapping. In EDS-SLAM, we design a set of dedicated, well optimized accelerators to achieve real-time implementations for the pipelines of stereo matching, feature extraction and key-point association. The effectiveness of EDS-SLAM is evaluated through comprehensive experiments. Compared to the state-of-the-art dense SLAM system with software-based implementation, EDS-SLAM achieves comparable accuracy on the localization but higher precision dense 3D mapping with up to$2.03\times$speed-up as well as up to$23.1\times$energy efficiency improvement, respectively.
Qian Huang 0004, Gaoxing Shang, Gang Chen 0023
ICCAD4
2023 Towards Effective Training of Robust Spiking Recurrent Neural Networks Under General Input Noise via Provable Analysis
abstract
Recently, bio-inspired spiking neural networks (SNN) with recurrent structures (SRNN) have received increasingly more attention due to their appealing properties for energy-efficiently solving time-domain classification tasks. SRNN s are often executed in noisy environments on resource-constrained devices which can however greatly compromise its accuracy. Thus, one fundamental question that remains unanswered is whether a formal analysis under the general input noise disturbances can be obtained to guarantee the robustness of SRNNs. Several studies have shown great promises by optimizing the bound over adverse perturbations based on Lipschitz continuity theorem, but most of these theoretical analysis are confined to convolutional neural networks (CNN). In this work, we take a further step towards robust SRNN training via provable robustness analysis over input noise perturbations. We show that it is feasible to establish bound analysis for evaluating noise sensitivity for SRNN by using the relation between the input current and the membrane potential change magnitude across a time window. Inspired by the theoretical analysis, we next propose a targeted penalty term in the objective function for training robust SRNN. Experimental results show that our solution outperforms the more complicated state-of-the-art methods on the commonly tested Fashion MNIST and CIFAR-IO image classification datasets.
Wendong Zheng, Yu Zhou 0029, Gang Chen 0023, Zonghua Gu 0001
ICCAD3
2023 ER3D: An Efficient Real-time 3D Object Detection Framework for Autonomous Driving
abstract
3D object detection is a vital computer vision task in mobile robotics and autonomous driving. However, most existing methods have exclusively focused on achieving high accuracy, leading to complex and bulky systems that can not be deployed in a real-time manner. In this paper, we propose the ER3D (Efficient and Real-time 3D) object detection framework, which takes stereo images as input and predicts 3D bounding boxes. Instead of using the complex network architecture, we leverage a fast-but-inaccurate method of semi-global matching (SGM) for depth estimation. To eliminate the accuracy degradation in 3D detection caused by inaccurate depth estimation, we introduce decoupled regression head and 3D distance-consistency loU loss to boost the accuracy performance of the 3D detector with a small computing overhead. ER3D achieves both high-precision and real-time performance to enable practical applications of 3D object detection systems on robotic systems. Extensive experiments with the comparison of the state of the arts demonstrate the superior practicability of ER3D, which achieves comparable detection accuracy with significant leadership on inference efficiency.
Haitao Meng, Changcai Li, Gang Chen 0023, Zonghua Gu 0001, Alois C. Knoll
ICPADS3
2023 FastOmniMVS: Real-time Omnidirectional Depth Estimation from Multiview Fisheye Images
abstract
Omnidirectional depth sensing, which can build 3-D scene structure of the objects of interest in all directions without any blind regions, is essential for ensuring the safety of robotic systems in complex environments. Recent researches attempt to use end-to-end deep neural network (DNN) models for omnidirectional depth estimation from multiview fisheye images. However, most of these existing DNN-based researches cannot achieve real-time performance. In this paper, we propose, FastOmniMVS, a lightweight end-to-end deep neural network (DNN) to achieve real-time omnidirectional depth estimation. FastOmniMVS consists of three parts: a feature extraction module that extracts features from the input fisheye images, spherical sweeping that maps the features to the spherical space to form a 4D cost volume, and finally, the depth is computed by a 3D encoder-decoder. Additionally, we conduct Quantization-Aware Training (QAT) for this network and deploy the quantized model on Jetson Orin. Through extensive experiments, we demonstrate that FastOmniMVS achieves high computational efficiency while maintaining performance close to the state-of-the-art (SOTA) accuracy. It can achieve a real-time inference speed of 35 frames per second (FPS) on the Jetson Orin device.
Yushen Wang, Jiaxi Deng, Haitao Meng, Gang Chen 0023
ICPADS5
2023 Rotation-Invariant Descriptors Learned with Circulant Convolution Neural Networks
abstract
Extracting local features for accurate correspondences between image pairs is an essential basis for various computer vision tasks. Recent works have shown that deep neural networks (DNNs) have demonstrated promising performance in challenging environments. However, these state-of-the-art DNN-based approaches are not well suited for the scenario with geometry rotations due to their intrinsic deficiencies of square kernel structure. That is, square kernel structures in standard DNNs cannot fully identify the essentials of the rotations in geometry. To address this problem, we present RICNN, a novel deep learning framework that encodes invariance against the rotations in geometry explicitly into convolutional neural networks. Rather than using the square-shaped kernel structure, RICNN adopts sectorshaped convolutional kernels to achieve encoding invariance in all rotations. With the explicitness of such rotation encoding, RICNN enables the transfer of perspective DNN models to obtain rotation-invariant descriptions. Furthermore, we propose a novel multi-level hinge triplet loss function to strengthen the matching constraints against geometry rotations. Comprehensive experiments demonstrate the strong generalization ability of the RICNN descriptor on the HPatches dataset. Toward the rotation invariance evaluation, our method shows state-of-the-art results.
Wenwei Lin, Chonghao Zhong, Xunpei Sun, Haitao Meng, Gang Chen 0023, Biao Hu 0001, Zonghua Gu 0001
ICTAI5
2023 Dense Depth Estimation for Monocular Endoscope Robot with an Adaptive Baseline
abstract
Depth information is useful to surgeons and surgical assistance systems. However, it is a challenging task to estimate the depth of various surgical scenes based on a monocular endoscope. We propose a depth estimation approach for a monocular endoscope with a stereo matching algorithm. The monocular endoscope is moved horizontally by a robotic endoscope holder to simulate a stereo vision system. The main challenge is how to obtain a proper baseline for better depth information generation as the depth range of a surgical scene is unknown beforehand. We design a baseline evaluation and selection algorithm to search for suitable baselines for surgical scenes with different depth ranges. Experimental results show that our approach improves the average accuracy of different depth scenarios by 10.8% when the error range is 2mm.
Rihui Song, Zhidong Tan, Hongli Liang, Yehua Ling, Gang Chen 0023, Kai Huang 0001, Jin Gong
SMC5
2023 LRP-based network pruning and policy distillation of robust and non-robust DRL agents for embedded systems
abstract
Summary Reinforcement learning (RL) is an effective approach to developing control policies by maximizing the agent's reward. Deep reinforcement learning uses deep neural networks (DNNs) for function approximation in RL, and has achieved tremendous success in recent years. Large DNNs often incur significant memory size and computational overheads, which may impede their deployment into resource‐constrained embedded systems. For deployment of a trained RL agent on embedded systems, it is necessary to compress the policy network of the RL agent to improve its memory and computation efficiency. In this article, we perform model compression of the policy network of an RL agent by leveraging the relevance scores computed by layer‐wise relevance propagation (LRP), a technique for Explainable AI (XAI), to rank and prune the convolutional filters in the policy network, combined with fine‐tuning with policy distillation. Performance evaluation based on several Atari games indicates that our proposed approach is effective in reducing model size and inference time of RL agents. We also consider robust RL agents trained with RADIAL‐RL versus standard RL agents, and show that a robust RL agent can achieve better performance (higher average reward) after pruning than a standard RL agent for different attack strengths and pruning rates.
Siyu Luan, Zonghua Gu 0001, Qingling Zhao, Gang Chen 0023
Concurr. Comput. Pract. Exp.5
2023 A GPU-accelerated real-time human voice separation framework for mobile phones
Gang Chen 0023, Zhaoheng Zhou, Shengyu He, Wang Yi 0001
J. Syst. Archit.1
2023 An efficient real-time accelerator for high-accuracy DNN-based optical flow estimation in FPGA
Yuanxing Yan, Yehua Ling, Kai Huang 0001, Gang Chen 0023
J. Syst. Archit.4
2023 FTSC: Fault-tolerant scheduling and control co-design for distributed real-time system
Yuanhai Zhang, Zijin Xu, Nan Guan, Shuai Zhao 0004, Gang Chen 0023, Kai Huang 0001
J. Syst. Archit.6
2023 A robust and real-time DNN-based multi-baseline stereo accelerator in FPGAs
Yehua Ling, Haitao Meng, Gang Chen 0023
J. Syst. Archit.5
2023 A Hybrid Spiking Neurons Embedded LSTM Network for Multivariate Time Series Learning Under Concept-Drift Environment
abstract
Complicated temporal patterns can provide important information for accurate time series forecasting. Existing long short-term memory (LSTM) model with attention mechanism have achieved significant performance. However, the exponential decay of long-term memory of LSTM has not be resolved yet in these efforts, remaining a longstanding open problem in recurrent nature. This problem exhibits a bottleneck which restricts the performance of existing studies. Recently, spiking neural networks (SNNs) have shown high efficiency in capturing temporal patterns via the surrogate gradient (SG) method to resolve this issue. However, the concept-drift environment makes it impossible to pre-set the variance into the standard SG method due to time-varying data distribution. In this paper, we propose a novel adaptive and hybrid spiking (AHS) module embedded LSTM, collaborating with two attention mechanisms (called HSN-LSTM) to resolve above-mentioned problems. First, the AHS module is analyzed theoretically can remain long-term memory. Moreover, our smooth SG method avoids pre-setting of variance, which is not sensitive in the above scenarios. Besides, we use the negative log-likelihood function to adjust the attention score for alleviating the negative impact from the concept-drift. Experiment results show the HSN-LSTM outperformed the state-of-the-art models on several multivariate time series datasets.
Wendong Zheng, Putian Zhao, Gang Chen 0023, Yonghong Tian 0001
IEEE Trans. Knowl. Data Eng.3
2022 FlowAcc: Real-Time High-Accuracy DNN-based Optical Flow Accelerator in FPGA
abstract
Recently, accelerator architectures have been designed to use deep neural networks (DNNs) to accelerate computer vision tasks, possessing the advantages of both accuracy and speed. Optical flow accelerator is however not among these architectures that DNNs have been successfully deployed. Existing hardware accelerators for optical flow estimation are all designed for classic methods and generally perform poorly in estimated accuracy. In this paper, we present FlowAcc, a dedicated hardware accelerator for DNN-based optical flow estimation, adopting a pipelined hardware design for real-time processing of image streams. We design an efficient multiplexing binary neural network (BNN) architecture for pyramidal feature extraction to significantly reduce the hardware cost and make it independent of the pyramid level number. Furthermore, efficient hamming distance calculation and competent flow regularization are utilized for hierarchical optical flow estimation to greatly improve the system efficiency. Comprehensive experimental results demonstrate that FlowAcc achieves state-of-the-art estimation accuracy and real-time performance on the Middlebury dataset when compared with the existing optical flow accelerators.
Yehua Ling, Yuanxing Yan, Kai Huang 0001, Gang Chen 0023
DATE4
2022 Ultra-Flow: An Ultra-fast and High-quality Optical Flow Accelerator with Deep Feature Matching on FPGA
abstract
Dense and accurate optical flow estimation is an important requirement for dynamic scene perception in autonomous systems. However, most of the existing FPGA accelerators are based on classic methods, which cannot deal with large displacements of moving objects in ultra-fast scenes. In this paper, we present Ultra-Flow, an ultra-fast pipelined architecture for efficient optical flow estimation and refinement. Ultra-Flow utilizes binary neural networks to generate the robust feature map, on which hierarchical matching is directly performed. Therefore, multiple usages of neural networks at hierarchical levels can be avoided to achieve hardware efficiency in Ultra-Flow. Optimizations, including local flow regularization and enhanced matching, are further used to improve the throughput and refine the optical flow to obtain higher accuracy. Evaluation results show that, compared to state-of-the-art FPGA accelerators, Ultra-Flow achieves leading accuracy in the Middlebury sequences at ultra-fast processing speed up to 687.92 frames/s for 640 × 480 pixel images.
Yehua Ling, Yuanxing Yan, Kai Huang 0001, Gang Chen 0023
FPL4
2022 De-snowing LiDAR Point Clouds With Intensity and Spatial-Temporal Features
abstract
Point clouds from 3D light detection and ranging (LiDAR) are widely used. Noise caused by falling snow reduces the availability of point clouds. Due to the sparseness of LiDAR point clouds and the fact that the snow point clouds are easily affected by multi factors such as wind or snowfall conditions, it is difficult to accurately remove the snow while preserving the details of the point clouds. To solve the problem, this paper presents a de-snowing approach combining the intensity and spatial-temporal features. An intensity-based filter firstly removes the snow. Then a repairing method restores the non-snow points based on the spatial-temporal features. Experimental results demonstrate that our approach outperforms existing work in the literature and performs the least damage to the point clouds in different snowfall scenarios.
Boyang Li 0009, Jieling Li, Gang Chen 0023, Hejun Wu, Kai Huang 0001
ICRA3
2022 LRP-based Policy Pruning and Distillation of Reinforcement Learning Agents for Embedded Systems
abstract
Reinforcement Learning (RL) is an effective approach to developing control policies by maximizing the agent’s reward. Deep Reinforcement Learning (DRL) uses Deep Neural Networks (DNNs) for function approximation in RL, and has achieved tremendous success in recent years. Large DNNs often incur significant memory size and computational overheads, which greatly impedes their deployment into resource-constrained embedded systems. For deployment of a trained RL agent on embedded systems, it is necessary to compress the Policy Network of the RL agent to improve its memory and computation efficiency. In this paper, we perform model compression of the Policy Network of an RL agent by leveraging the relevance scores computed by Layer-wise Relevance Propagation (LRP), a technique for Explainable AI (XAI), to rank and prune the convolutional filters in the Policy Network, combined with fine-tuning with Policy Distillation. Performance evaluation based on several Atari games indicates that our proposed approach is effective in reducing model size and inference time of RL agents.
Siyu Luan, Zonghua Gu 0001, Qingling Zhao, Gang Chen 0023
ISORC5
2022 Lite-Stereo: A Resource-Efficient Hardware Accelerator for Real-Time High-Quality Stereo Estimation Using Binary Neural Network
abstract
Stereo estimation plays a key role in many autonomous systems, such as robotics and self-driving cars. Recent work on StereoEngine, an FPGA-based accelerator for deep neural network (DNN)-based stereo estimation, has been demonstrated as a promising solution to achieve both real-time and high accuracy performance for depth sensing. However, this solution still suffers from over-utilizing the hardware resource of FPGAs. In this article, we present Lite-Stereo, a resource-efficient DNN-based stereo vision accelerator to improve the hardware efficiency for StereoEngine running on a resource-constrained FPGA. To achieve this, we design a set of optimized hardware architectures for resource-demanding bottleneck modules. In order to balance the gap between the processing speed and resource efficiency, the process elements in binary neural network modules are shared within and across modules. In addition, we provide reusing strategies on path aggregation and neighbor calculation to improve the resource efficiency of the semi-global matching module. Evaluation results demonstrate that Lite-Stereo reduces the hardware cost of ALUTs and RAM bits by 60% and 29%, respectively, without compromising the accuracy and energy efficiency compared with StereoEngine.
Yehua Ling, Haitao Meng, Kai Huang 0001, Gang Chen 0023
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 An Accurate GRU-Based Power Time-Series Prediction Approach With Selective State Updating and Stochastic Optimization
abstract
Accurate power time-series prediction is an important application for building new industrialized smart cities. The gated recurrent units (GRUs) models have been successfully employed to learn temporal information for power time-series prediction, demonstrating its effectiveness. However, from a statistical perspective, these existing models are geometrically ergodic with short-term memory that causes the learned temporal information to be quickly forgotten. Meanwhile, these existing approaches completely ignore the temporal dependencies between the gradient flow in the optimization algorithm, which greatly limits the prediction accuracy. To resolve these issues, we propose a novel GRU model coupling two new mechanisms of selective state updating and adaptive mixed gradient optimization (GRU-SSU-AMG) to improve the accuracy of prediction. Specifically, a tensor discriminator is used for adaptively determining whether hidden state information needs to be updated at each time step for learning the extremely fluctuating information in the proposed selective GRU (SGRU). In addition, an adaptive mixed gradient (AdaMG) optimization method that mixes the moment estimations is proposed to further improve the capability of learning the temporal dependencies information. The effectiveness of the GRU-SSU-AMG has been extensively evaluated on five different real-world datasets. The experimental results show that the GRU-SSU-AMG achieves significant accuracy improvement compared with the state-of-the-art approaches.
Wendong Zheng, Gang Chen 0023
IEEE Trans. Cybern.2
2021 Understanding the Property of Long Term Memory for the LSTM with Attention Mechanism
abstract
Recent trends of incorporating LSTM network with different attention mechanisms in time series forecasting have led researchers to consider the attention module as an essential component. While existing studies revealed the effectiveness of attention mechanism with some visualization experiments, the underlying rationale behind their outstanding performance on learning long-term dependencies remains hitherto obscure. In this paper, we aim to elaborate on this fundamental question by conducting a thorough investigation of the memory property for LSTM network with attention mechanism. We present a theoretical analysis of LSTM integrated with attention mechanism, and demonstrate that it is capable of generating an adaptive decay rate which dynamically controls the memory decay according to the obtained attention score. In particular, our theory shows that attention mechanism brings significantly slower decays than the exponential decay rate of a standard LSTM. Experimental results on four real-world time series datasets demonstrate the superiority of the attention mechanism for maintaining long-term memory when compared to the state-of-the-art methods, and further corroborate our theoretical analysis.
Wendong Zheng, Putian Zhao, Kai Huang 0001, Gang Chen 0023
CIKM4
2021 A GPU -accelerated Deep Stereo- LiDAR Fusion for Real-time High-precision Dense Depth Sensing
abstract
Active LiDAR and stereo vision are the most commonly used depth sensing techniques in autonomous vehicles. Each of them alone has weaknesses in terms of density and reliability and thus cannot perform well on all practical scenarios. Recent works use deep neural networks (DNNs) to exploit their complementary properties, achieving a superior depth-sensing. However, these state-of-the-art solutions are not satisfactory on real-time responsiveness due to the high computational complexities of DNNs. In this paper, we present FastFusion, a fast deep stereo-LiDAR fusion framework for real-time high-precision depth estimation. FastFusion provides an efficient two-stage fusion strategy that leverages binary neural network to integrate stereo-LiDAR information as input and use cross-based LiDAR trust aggregation to further fuse the sparse LiDAR measurements in the back-end of stereo matching. More importantly, we present a GPU-based acceleration framework for providing a low latency implementation of FastFusion, gaining both accuracy improvement and real-time responsiveness. In the experiments, we demonstrate the effectiveness and practicability of FastFusion, which obtains a significant speedup over state-of-the-art baselines while achieving comparable accuracy on depth sensing.
Haitao Meng, Chonghao Zhong, Jianfeng Gu 0001, Gang Chen 0023
DATE4
2021 Brief Industry Paper: SylixOS: A Secure and Compatible RTOS with Constant Scheduling on SMP
abstract
Real time operating system (RTOS) is an operating system which is designed to execute tasks in a guaranteed time. In this paper, we introduce SylixOS - a powerful and open source RTOS. As an industrial RTOS, SylixOS has been used in aerospace, military and many more embedded devices. We introduce the development path and some advanced features of SylixOS. In addition, SylixOS also supports the multiprocess. We discuss how SMP scheduling algorithms can be applied in SylixOS. Performance evaluations of SylixOS are presented to assess the real-time efficiency of SylixOS.
Yuanhai Zhang, Jinxing Jiao, Guizhou Xu, Gang Chen 0023
RTAS5
2021 Peak temperature analysis and optimization for pipelined hard real-time systems
Long Cheng 0007, Kai Huang 0001, Liang Mi, Gang Chen 0023, Alois C. Knoll
Inf. Sci.4
2021 An efficient GPU-accelerated inference engine for binary neural network on mobile phones
Shengyu He, Haitao Meng, Zhaoheng Zhou, Kai Huang 0001, Gang Chen 0023
J. Syst. Archit.6
2021 Hardware accelerator for an accurate local stereo matching algorithm using binary neural network
Yehua Ling, Haitao Meng, Gang Chen 0023
J. Syst. Archit.5
2021 Efficient runtime slack management for EDF-VD-based mixed-criticality scheduling
Junjie Yang 0001, Guangyi Xu 0004, Gang Chen 0023, Nan Guan, Kai Huang 0001
J. Syst. Archit.3
2021 Efficient and Effective Dimension Control in Automotive Applications
abstract
In automotive industry, the production line for assembling mechanical parts of vehicles must place and weld hundreds of components on the right positions of the platform. The accuracy of deploying the components has great impact on the quality and performance of the produced vehicle. To ensure the assembly accuracy, a critical task in the production process is the so-called dimension quality control. The current state of practice in automotive industries is mainly based on a manual process where experienced engineers use production data to identify accuracy problems and suggest solutions for corrections on fixture adjustment in the assembly line. It is an extremely inefficient process, which typically takes the engineers around ten days for one batch of vehicles and a year to achieve the required assembly accuracy for final production. In this article, we present an automatic technique for dimension control. We formulate the dimension control problem as a constraint programming problem and present a refinement method to prune the exploration space. Our technique can not only identify the wrongly deployed parts leading to dimensional defects, but also provide high-quality fixture adjustment decisions. Experiments conducted on industrial production data from BMW Brilliance Automotive demonstrate the significantly improved efficiency and effectiveness of dimension control in automotive industries with our approach.
Gang Chen 0023, Mingsong Lv, Wang Yi 0001, Xue (Steve) Liu, Hao Chen 0087, Bo Zhu 0006
IEEE Trans. Ind. Informatics2
2020 PhoneBit: Efficient GPU-Accelerated Binary Neural Network Inference Engine for Mobile Phones
abstract
Over the last years, a great success of deep neural networks (DNNs) has been witnessed in computer vision and other fields. However, performance and power constraints make it still challenging to deploy DNNs on mobile devices due to their high computational complexity. Binary neural networks (BNNs) have been demonstrated as a promising solution to achieve this goal by using bit-wise operations to replace most arithmetic operations. Currently, existing GPU-accelerated implementations of BNNs are only tailored for desktop platforms. Due to architecture differences, mere porting of such implementations to mobile devices yields suboptimal performance or is impossible in some cases. In this paper, we propose PhoneBit, a GPU-accelerated BNN inference engine for Android-based mobile devices that fully exploits the computing power of BNNs on mobile GPUs. PhoneBit provides a set of operator-level optimizations including locality-friendly data layout, bit packing with vectorization and layers integration for efficient binary convolution. We also provide a detailed implementation and parallelization optimization for PhoneBit to optimally utilize the memory bandwidth and computing power of mobile GPUs. We evaluate PhoneBit with AlexNet, YOLOv2 Tiny and VGG16 with their binary version. Our experiment results show that PhoneBit can achieve significant speedup and energy efficiency compared with state-of-the-art frameworks for mobile devices.
Gang Chen 0023, Shengyu He, Haitao Meng, Kai Huang 0001
DATE1
2020 Fault-tolerant real-time tasks scheduling with dynamic fault handling
Gang Chen 0023, Nan Guan, Kai Huang 0001, Wang Yi 0001
J. Syst. Archit.1
2020 StereoEngine: An FPGA-Based Accelerator for Real-Time High-Quality Stereo Estimation With Binary Neural Network
abstract
Stereo estimation is essential to many applications such as mobile autonomous robots, most of which ask for real-time response, high energy, and storage efficiency. Deep neural networks (DNNs) have shown to yield significant gains in improving accuracy. However, these DNN-based algorithms are challenging to be deployed on energy and resource-constrained devices due to the high computational complexities of DNNs. In this article, we present StereoEngine, a fully pipelined end-to-end stereo vision accelerator that computes accurate dense depth in a real-time and energy-efficient manner. An efficient stereo algorithm is developed and optimized for a high-quality hardware-friendly implementation, that leverages binary neural network (BNN) to learn discriminative binary descriptors to improve the disparity. The design of StereoEngine is a standalone DNN-based stereo vision system where all processing procedures are implemented on a hardware platform. The effectiveness of StereoEngine is evaluated by comprehensive experiments. Compared with software-based implementations on the highend and embedded Nvidia GPUs, StereoEngine achieves up to 3×, 13×, and 50× speedups, as well as up to 211×, 58×, and 73× energy efficiency improvement, respectively. Furthermore, StereoEngine achieves leading accuracy when compared to state-of-the-art hardware implementations on the challenging KITTI dataset.
Gang Chen 0023, Yehua Ling, Haitao Meng, Shengyu He, Kai Huang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 GPU-Accelerated Real-Time Stereo Estimation With Binary Neural Network
abstract
Depth estimation from stereo images is essential to many applications such as robotics and autonomous vehicles, most of which ask for the real-time response, high energy and storage efficiency. Recent work has shown deep neural networks (DNN) perform extremely well for stereo estimation. However, these state-of-the-art DNN based algorithms are challenging to be deployed into real-world applications due to the high computational complexities of DNNs. Most of them are too slow for real-time inference and require several seconds of GPU computation to process image frames. In this article, we address the problem of fast stereo estimation and propose an efficient and light-weighted stereo matching system, called StereoBit, to produce a disparity map in a real-time manner while achieving close to state-of-the-art accuracy. To achieve this goal, we propose a binary neural network to generate weighted Hamming distance for an efficient similarity join in stereo estimation. In addition, we propose a novel approximation approach to derive StereoBit network directly from the well-trained network with the cosine similarity. Our approximation strategies enable a significant speedup while maintaining almost the same accuracy compared to the network with the cosine similarity. Furthermore, we present an optimization framework for fully exploiting the computing power of StereoBit. The framework provides a significant speedup of stereo estimation routines, and at the same time, reduces the memory usage for storing parameters. The effectiveness of StereoBit is evaluated by comprehensive experiments. StereoBit can achieve 60 frames per second on an NVIDIA TITAN Xp GPU on KITTI 2012 benchmark while achieving 3-pixel non-occluded stereo error 3.56 percent.
Gang Chen 0023, Haitao Meng, Yucheng Liang, Kai Huang 0001
IEEE Trans. Parallel Distributed Syst.1
2018 Utilization-Based Scheduling of Flexible Mixed-Criticality Real-Time Tasks
abstract
Mixed-criticality models are an emerging paradigm for the design of real-time systems because of their significantly improved resource efficiency. However, formal mixed-criticality models have traditionally been characterized by two impractical assumptions: once any high-criticality task overruns, all low-criticality tasks are suspended and all other high-criticality tasks are assumed to exhibit high-criticality behaviors at the same time. In this paper, we propose a more realistic mixed-criticality model, called the flexible mixed-criticality (FMC) model, in which these two issues are addressed in a combined manner. In this new model, only the overrun task itself is assumed to exhibit high-criticality behavior, while other high-criticality tasks remain in the same mode as before. The guaranteed service levels of low-criticality tasks are gracefully degraded with the overruns of high-criticality tasks. We derive a utilization-based technique to analyze the schedulability of this new mixed-criticality model under EDF-VD scheduling. During run time, the proposed test condition serves an important criterion for dynamic service level tuning, by means of which the maximum available execution budget for low-criticality tasks can be directly determined with minimal overhead while guaranteeing mixed-criticality schedulability. Experiments demonstrate the effectiveness of the FMC scheme compared with state-of-the-art techniques.
Gang Chen 0023, Nan Guan, Di Liu 0002, Qingqiang He, Kai Huang 0001, Todor P. Stefanov, Wang Yi 0001
IEEE Trans. Computers1
2018 Scheduling Analysis of Imprecise Mixed-Criticality Real-Time Tasks
abstract
In this paper, we study the scheduling problem of the imprecise mixed-criticality model (IMC) under earliest deadline first with virtual deadline (EDF-VD) scheduling upon uniprocessor systems. Two schedulability tests are presented. The first test is a concise utilization-based test which can be applied to the implicit deadline IMC task set. The suboptimality of the proposed utilization-based test is evaluated via a widely-used scheduling metric, speedup factors. The second test is a more effective test but with higher complexity which is based on the concept of demand bound function (DBF). The proposed DBF-based test is more generic and can apply to constrained deadline IMC task set. Moreover, in order to address the high time cost of the existing deadline tuning algorithm, we propose a novel algorithm which significantly improve the efficiency of the deadline tuning procedure. Experimental results show the effectiveness of our proposed schedulability tests, confirm the theoretical suboptimality results with respect to speedup factor, and demonstrate the efficiency of our proposed algorithm over the existing deadline tunning algorithm. In addition, issues related to the implementation of the IMC model under EDF-VD are discussed.
Di Liu 0002, Nan Guan, Jelena Spasic, Gang Chen 0023, Songran Liu, Todor P. Stefanov, Wang Yi 0001
IEEE Trans. Computers4
2018 EDF-VD Scheduling of Flexible Mixed-Criticality System With Multiple-Shot Transitions
abstract
The existing mixed-criticality (MC) real-time task models assume that once any high-criticality task overruns, all high-criticality jobs execute up to their most pessimistic WCET estimations simultaneously in a one-shot manner. This is very pessimistic in the sense of unnecessary resource overbooking. In this paper, we propose a more generalized mixed-critical real-time task model, called flexible MC model with multiple-shot transitions (FMC-MST), to address this problem. In FMC-MST, high-criticality tasks can transit multiple intermediate levels to handle less pessimistic overruns independently and to nonuniformly scale the deadline on each level. We develop a run-time schedulability analysis for FMC-MST under EDF-VD scheduling, in which a better tradeoff between the penalties of low-criticality tasks and the overruns of high-criticality tasks is achieved to improve the service quality of low-criticality tasks. We also develop a resource optimization technique to find resource-efficient level-insertion configurations for FMC-MST task systems under MC timing constraints. Experiments demonstrate the effectiveness of FMC-MST compared with the state-of-the-art techniques.
Gang Chen 0023, Nan Guan, Biao Hu 0001, Wang Yi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 Online workload monitoring with the feedback of actual execution time for real-time systems
abstract
Guaranteeing the system workload within design bounds is a basic requirement for a real-time system. Design-time bounds are usually based on worst-case activation patterns and worst-case execution time. While using the worst-case assumptions for online monitoring can guarantee the system safety, it also introduces unexplored slacks due to tasks consuming less than their worst-case execution times. In this paper, we introduce a monitoring scheme with the feedback of actual execution time for real-time systems. By using this runtime feedback instead of offline assumptions, this monitoring scheme can accept events that are considered as violations offline, and thereby improve the system utilization. In the experiments of both MATLAB simulation and MicroC/OS-II running in a softcore processor implemented on an FPGA, different probability distributions of actual execution time are used in analyzing how much the benefit can be gained from the feedback scheme.
Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll
DATE3
2016 Minimizing peak temperature for pipelined hard real-time systems
Long Cheng 0007, Kai Huang 0001, Gang Chen 0023, Biao Hu 0001, Alois C. Knoll
DATE3
2016 EDF-VD Scheduling of Mixed-Criticality Systems with Degraded Quality Guarantees
abstract
This paper studies real-time scheduling of mixed-criticality systems where low-criticality tasks are still guaranteed some service in the high-criticality mode, with reduced execution budgets. First, we present a utilization-based schedulability test for such systems under EDF-VD scheduling. Second, we quantify the suboptimality of EDF-VD (with our test condition) in terms of speedup factors. In general, the speedup factor is a function with respect to the ratio between the amount of resource required by different types of tasks in different criticality modes, and reaches 4/3 in the worst case. Furthermore, we show that the proposed utilization-based schedulability test and speedup factor results apply to the elastic mixed-criticality model as well. Experiments show effectiveness of our proposed method and confirm the theoretical suboptimality results.
Di Liu 0002, Jelena Spasic, Nan Guan, Gang Chen 0023, Songran Liu, Todor P. Stefanov, Wang Yi 0001
RTSS4
2016 Evaluation and Improvements of Runtime Monitoring Methods for Real-Time Event Streams
abstract
Runtime monitoring is of great importance as a safeguard to guarantee the correctness of system runtime behaviors. Two state-of-the-art methods, dynamic counters and l -repetitive function, were recently developed to tackle the runtime monitoring for real-time systems. While both are reported to be efficient in monitoring arbitrary events, the monitoring performance between them has not yet been evaluated. This article evaluates both methods in depth, to identify their strengths and weaknesses. New methods are proposed to efficiently monitor the many-to-one connections that are abstracted as AND and OR components on multiple inputs. Representative scenarios are used as our case studies to quantitatively demonstrate the evaluations. Both methods are implemented in hardware F pga . The timing overhead and resource usages of implementing the two methods are evaluated.
Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll
ACM Trans. Embed. Comput. Syst.3
2016 Adaptive Workload Management in Mixed-Criticality Systems
Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll
ACM Trans. Embed. Comput. Syst.3
2015 Evaluation of runtime monitoring methods for real-time event streams
abstract
Runtime monitoring is of great importance as a safe guard to guarantee the correctness of system runtime behaviors. Two new methods, i.e., dynamic counters and l-repetitive function, are recently developed to tackle the runtime monitoring for hard real-time systems. This paper investigates in depth these two newly developed runtime monitoring methods, trying to evaluate and identify their strengths and weaknesses. Representative scenarios are used as our case studies to quantitatively demonstrate our comparisons. We also provide FPGA implementations and resource usages of both methods.
Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Alois C. Knoll
ASP-DAC3
2015 Adaptive runtime shaping for mixed-criticality systems
abstract
This paper investigates runtime shaping for mixed-criticality systems to increase the system QoS. Unlike the previous work in the literature that enforces an offline workload bound, an adaptively shaping approach is proposed where the incoming workload of the low-critical tasks is regulated by the actual demand of the high-critical tasks. This actual demand is adaptively updated using the historical arrival information of the high-critical tasks and thus can maximize the runtime QoS of low-critical tasks. To reduce the online overheads of computing the workload demand, a lightweight scheme with the complexity of O(n log(m)) is developed. Experiments are also provided to demonstrate the effectiveness and efficiency of our approach.
Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll
EMSOFT3
2015 Profiling and annotation combined method for multimedia application specific MPSoC performance estimation
abstract
Accurate and fast performance estimation is necessary to drive design space exploration and thus support important design decisions. Current techniques are either time consuming or not accurate enough. In this paper, we solve these problems by presenting a hybrid method for multimedia multiprocessor system-on-chip (MPSoC) performance estimation. A general coverage analysis tool GNU gcov is employed to profile the execution statistics during the native simulation. To tackle the complexity and keep the analysis and simulation manageable, the orthogonalization of communication and computation parts is adopted. The estimation result of the computation part is annotated to a transaction accurate model for further analysis, by which a gradual refinement of MPSoC performance estimation is supported. The implementation and its experimental results prove the feasibility and efficiency of the proposed method.
Kai Huang 0002, Siwen Xiu, Dandan Zheng 0001, Min Yu 0006, De Ma, Kai Huang 0001, Gang Chen 0023, Xiaolang Yan
Frontiers Inf. Technol. Electron. Eng.8
2015 Applying Pay-Burst-Only-Once Principle for Periodic Power Management in Hard Real-Time Pipelined Multiprocessor Systems
abstract
Pipelined computing is a promising paradigm for embedded system design. Designing a power management policy to reduce the power consumption of a pipelined system with nondeterministic workload is, however, nontrivial. In this article, we study the problem of energy minimization for coarse-grained pipelined systems under hard real-time constraints and propose new approaches based on an inverse use of the pay-burst-only-once principle. We formulate the problem by means of the resource demands of individual pipeline stages and propose two new approaches, a quadratic programming-based approach and fast heuristic, to solve the problem. In the quadratic programming approach, the problem is transformed into a standard quadratic programming with box constraint and then solved by a standard quadratic programming solver. Observing the problem is NP-hard, the fast heuristic is designed to solve the problem more efficiently. Our approach is scalable with respect to the numbers of pipeline stages. Simulation results using real-life applications are presented to demonstrate the effectiveness of our methods.
Gang Chen 0023, Kai Huang 0001, Christian Buckl, Alois C. Knoll
ACM Trans. Design Autom. Electr. Syst.1
2014 Resource optimization for CSDF-modeled streaming applications with latency constraints
abstract
In this paper, we study the problem of minimizing the number of processors required for scheduling latency-constrained streaming applications modeled as CSDF graphs, where the actors of a CSDF are executed as strictly periodic tasks. We formalize the problem and prove that due to the strict periodicity of actors the problem is an integer convex programming problem, that can be solved efficiently by using an existing convex programming solver. We evaluate our solution approach on a set of 13 real-life streaming applications modeled as CSDF graphs and demonstrate that it can reduce the number of processors in more than 52% of the conducted experiments in comparison to an existing approach.
Di Liu 0002, Jelena Spasic, Jiali Teddy Zhai, Todor P. Stefanov, Gang Chen 0023
DATE5
2014 Abstract: Shared L2 Cache Management in Multicore Real-Time System
abstract
In multicore system, shared cache interference has been recognized as one of the major factors that degrade the average performance as well as predictability of system. How to manage the shared cache in order to optimize the system performance while guaranteeing the system predictability is still an open issue. State-of-the-art techniques on this topic use page coloring to partition the shared cache at OS level. In this paper, we present a shared cache management scheme for multicore system. This shared cache management scheme supports way-based cache partitioning at hardware level, building task-level time-triggered reconfigurable-cache multicore system. We evaluated the proposed scheme w.r.t. different numbers of cores and cache modules and prototyped the constructed MPSoCs on FPGA.
Gang Chen 0023, Biao Hu 0001, Kai Huang 0001, Alois C. Knoll, Di Liu 0002
FCCM1
2014 Adaptive dynamic power management for hard real-time pipelined Multiprocessor Systems
abstract
Energy efficiency is a critical design concern for embedded systems. Dynamic power management (DPM) schemes in Multiprocessor System on Chips (MPSoCs) has been wildly used to explore the idleness of processors and dynamically reduce the energy consumption by putting idle processors to low-power states. In this paper, we explore how to effectively apply dynamic power management in adaptive manner to reduce leakage power consumption for coarse-grained pipelined systems under hard real-time requirements. At each adaptive point, a system transformation is proposed to model the pipeline system with unfinished events as multi-stream system. By using extended pay-burst-only-once principle, the service curves for corresponding stream can be computed as a constraint for a minimal resource demand and energy minimization problem can be formulated with respect to the resource demands at each adaptive point. One light-weight heuristic, called balance workload scheme (BWS), is proposed in this paper to solve the minimization problem. Simulation results using real-life applications are presented to demonstrate the effectiveness of our approach.
Gang Chen 0023, Kai Huang 0001, Alois C. Knoll
RTCSA1
2014 Energy optimization for real-time multiprocessor system-on-chip with optimal DVFS and DPM combination
abstract
Energy optimization is a critical design concern for embedded systems. Combining D VFS +D PM is considered as one preferable technique to reduce energy consumption. There have been optimal D VFS +D PM algorithms for periodic independent tasks running on uniprocessor in the literature. Optimal combination of D VFS and D PM for periodic dependent tasks on multicore systems is however not yet reported. The challenge of this problem is that the idle intervals of cores are not easy to model. In this article, a novel technique is proposed to directly model the idle intervals of individual cores such that both D VFS and D PM can be optimized at the same time. Based on this technique, the energy optimization problem is formulated by means of mixed integrated linear programming. We also present techniques to prune the exploration space of the formulation. Experimental results using real-world benchmarks demonstrate the effectiveness of our approach compared to existing approaches.
Gang Chen 0023, Kai Huang 0001, Alois C. Knoll
ACM Trans. Embed. Comput. Syst.1
2013 Cache partitioning and scheduling for energy optimization of real-time MPSoCs
abstract
Cache partitioning is a promising technique to reduce energy consumption of the cache subsystem for MPSoCs. Currently, most existing techniques focus primarily on static partition on core level. In this paper, we present a task-level approach and show that it outperforms core-level strategies. By taking the interference patterns of individual tasks into account, our approach generates optimal task-level cache partition schemes as well as feasible schedules at compilation time by means of a mixed integer linear programming formulation. We also present techniques to prune the exploration space of our formulation. Experimental results using real-world benchmarks demonstrate that our approach achieves 33% energy savings on average compared to core-based cache partition approaches.
Gang Chen 0023, Kai Huang 0001, Alois C. Knoll
ASAP1
2013 Energy optimization with worst-case deadline guarantee for pipelined multiprocessor systems
abstract
Pipelined computing is a promising paradigm for embedded system design. Designing the scheduling policy for a pipelined system is however more involved. In this paper, we study the problem of the energy minimization for coarse-grained pipelined systems under hard real-time constraints and propose a method based on an inverse use of the pay-burst-only-once principle. We formulate the problem by means of the resource demands of individual pipeline stages and solve it by quadratic programming. Our approach is scalable w.r.t the number of the pipeline stages. Simulation results using real-life applications as well as commercialized processors are presented to demonstrate the effectiveness of our method.
Gang Chen 0023, Kai Huang 0001, Christian Buckl, Alois C. Knoll
DATE1
2013 Effective Online Power Management with Adaptive Interplay of DVS and DPM for Embedded Real-Time System
abstract
Effective power management is an important design concern for modern embedded systems. In this paper, we present an effective framework to integrate both DVS and DPM to optimize the overall energy consumption. We propose an online algorithm to determine the optimal operating frequency and mode transition of a processor based on the runtime workload. Our algorithm runs in O(n) time, where n is the number of the events stored in the system buffer. A feasibility analysis is also presented, which serves as a criteria for setting the system buffer as well as runtime schedulabilty check. The evaluations with specifications of two commercial processors show that our algorithm is more energy-efficient compared to existing schemes in the literature.
Gang Chen 0023, Kai Huang 0001, Christian Buckl, Alois C. Knoll
DSD1
2012 Conforming the runtime inputs for hard real-time embedded systems
abstract
Timing is an important concern when designing an embedded system. While lots of researches on hard real-time systems focus on design-time analysis, monitoring the corresponding runtime behaviors are seldom investigated. In this paper, we investigate the conformity problem for runtime inputs of a hard real-time system. We adopt the widely used arrival curve model which captures the worst/best-cases event arrivals in the time interval domain and propose an algorithm to on-the-fly evaluate the conformity of the system input w.r.t. given arrival curves. The developed algorithm is lightweight in terms of both computation and memory overheads, which is particularly suitable for resource-constrained embedded systems. We also provide proofs and an Fpga implementation to demonstrate the effectiveness of our approach.
Kai Huang 0001, Gang Chen 0023, Christian Buckl, Alois C. Knoll
DAC2