Zonghui Li

dblp:91/10658 · DBLP profile ↗
← Back
30ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-0772-1738ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 5 first-author · 12 since 2021Computer networks · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EA-ONS: A Deterministic Time Synchronization Framework for 5G-TSN Integrated LEO Constellations
Zonghui Li, Chaoqun You, Yue Gao 0001
ICDCS2
2026 COS Scheduler: Conflict-Oriented Search Principle for Fast Slot-Free Time-Triggered Scheduling
Zonghui Li
RTAS2
2026 ADA: Asynchronous Deterministic Access for Time-Sensitive Networking
abstract
Time-Sensitive Networking (TSN) enables deterministic transmission by requiring end systems to support 802.1Qbv and 802.1AS standards, which makes practical deployment challenging, especially for asynchronous end systems (AES) such as various legacy devices already deployed in industry. We present a new deterministic transmission scheme for asynchronous time-triggered flows from AES. First, we design an Asynchronous Deterministic Access (ADA) mechanism with two components: “Async to Sync” and “Latency Adaptation.” These components allow an AES flow to wait for its scheduled period, ensuring a theoretical end-to-end (E2E) deterministic delay. Second, we propose an ADA scheduling model to optimize the scheduling periods for AES flows. However, latency adaptation may cause AES flows to conflict with other flows at the final hop switch. To address this problem, we propose a conflict resolution mechanism that calculates safety distances between flows to reduce delay jitter. Finally, we implement ADA in our switches and evaluate its performance in terms of E2E delay and jitter. ADA is shown to achieve E2E delay and jitter reductions of 57.1–83.5% and 60.7–85.4%, respectively, over SOTA methods.
Wenlin Zhu, Zonghui Li, Kang G. Shin
IEEE Trans. Netw.5
2025 Time-Triggered Flow Scheduling for Synchronization-Deviation-Tolerant TSN
abstract
Time-sensitive networking (TSN) has been widely used in industrial automation and automotive applications by precisely opening and closing the gates of packet queues. However, both clock drift and network congestion will likely result in timing misalignment when the clock synchronization protocol gPTP (802.1 AS) is used. This timing misalignment may, in turn, cause failure in forwarding packets at scheduled times and hence unexpected delay jitters, or even miss application deadlines. To address this acute problem, we propose a novel TSN scheduling algorithm, called SDT-TSN (Synchronization-Deviation-Tolerant TSN), to ensure that packets can still arrive on time and be transmitted deterministically even in the presence of inexact time synchronization. First, we formalize the linear constraint model of flow scheduling to maximize the tolerance of inexact time synchronization. Then, we propose an optimal algorithm based on SMT (Satisfiability Modulo Theories) and a fast heuristic algorithm to solve the packet scheduling problem under inexact network synchronization. SDT-TSN is the first to derive the maximum tolerable time-synchronization deviation. Finally, we evaluate SDT-TSN, demonstrating its capability of eliminating packet-forwarding failures due to the commonly-used/assumed constant time-synchronization deviation and increasing the tolerable synchronization deviation from 140µs to 480µs.
Jiamin Cui, Zonghui Li, Kang G. Shin
ICNP3
2025 SeFA: A Seed-Filter Adaptation Method for Robust Vision Services in IoT Devices
abstract
Using low-rank adaptation to fine-tuning pretrained neural network models has attracted widespread attention due to its advantages of low resource requirements, high precision, and no additional inference delay. However, most existing methods are designed for large language models based on Transformer structure and lack adaptation to convolutional neural networks (CNN) widely used on Internet of Things (IoT) devices. In addition, IoT devices are usually deployed outdoors and collect a large amount of data affected by the environment. Providing a highly robust model is a prerequisite for providing high-quality services. To this end, this paper proposes a highly robust seed-filter adaptation method (SeFA) for pre-trained CNNs. SeFA introduces an adaptation branch with the same structure as the backbone network. In the adaptation branch, some filters are first designated seed filters and grouped. Then, other filters are generated based on the grouped seed filters and nonlinear transformation functions (NLFs) with different hyperparameters. The parameters of the seed filters are updated with model training, and the hyperparameters of the NLFs are randomly initialized and frozen. Both grouping seed filters and configuring NLFs with nonlearnable hyperparameters can improve the robustness of the model. This is because grouping seed filters can generate diverse filters on demand without increasing the model's complexity, and the NLFs' rules can regularize the model. The key idea of this paper is to propose SeFA with flexible controllable learnable parameters, high robustness and adaptability to pretrained CNN models, to facilitate fine-tuning of pre-trained CNN models on resource-constrained IoT devices to provide highly robust visual services. Experimental results on the CIFAR-10, CIFAR-10-C, CIFAR-100, CIFAR-100-C, and Icons50 datasets demonstrate that the proposed SeFA outperforms other state-of-the-art methods. Specifically, based on the ResNet152, on the CIFAR-10-C dataset, the accuracy of our SeFA is about$+7 {\%}$higher than that of the full fine-tuning method.
Chuntao Ding, Longquan Zhang, Junna Zhang, Zonghui Li, Li Zhang 0004
ICWS4
2025 Schedulability-Driven Topology Optimization for EtherCAT-TSN Networks in Industrial Automation
abstract
The integration of EtherCAT and TSN has been proposed to enhance performance of EtherCAT networks in industrial automation. EtherCAT over TSN transforms a traditional EtherCAT ring into multiple shorter rings interconnected via TSN switches, enabling concurrent data transmission across segments and reducing cycle time. However, the use of multiple segments introduces contention for the master in the return path, potentially leading to scheduling failures and an increase in cycle time. We study the impact of network topology, i.e., the number of segments and slave node distribution, on schedulability, and formulate the topology optimization problem for the converged network based on schedulability analysis. We evaluate our methodology and optimal solution using SMT and Integer Programming (LIP) solvers, respectively. Numerical results demonstrate the effectiveness of our method, and our solution outperforms baselines.
Yi Duan, Hongyun Zheng, Zonghui Li, Yongxiang Zhao, Zhibo Pang
INDIN3
2025 Low Jitter Framework for the Converged Networks of CAN-FD and TSN
abstract
The convergence of Controller Area Network with Flexible Data-Rate (CAN-FD) and Time-Sensitive Networking (TSN) presents a critical pathway to enable deterministic cross-domain communication in next-generation intelligent vehicles. However, different transmission mechanisms are raising significant challenges in maintaining low jitter and guaranteed latency. This paper proposes a low jitter framework based on a CANFD-TSN gateway to address these limitations through three key innovations: 1) A CQF-based gateway architecture integrating cyclic queuing with deadline-aware traffic scheduling, 2) An ILP model optimizing queue switching cycles, and 3) A bidirectional phase alignment mechanism that compensates asymmetric queuing delays through gateway timestamp synchronization, achieving microsecond-level jitter suppression. Extensive OMNeT++ simulations demonstrate the framework’s effectiveness, 18% higher scheduling success rates compared to conventional methods (RCSF, SPs, EDF) under 160-flow scenarios, while reducing end-to-end jitter by 42% through alignment time compensation.
Fucheng Li, Chunxi Li, Zonghui Li, Zhibo Pang
INDIN3
2025 Time-Triggered Communication for Deterministic Ad Hoc Networks
abstract
Ad Hoc networks, as a flexible type of wireless sensor network, find wide applications in disaster relief, and industrial scenarios. In industrial applications, there is a growing demand for deterministic communication to support time-critical business flows. However, previous research mainly focused on aspects like routing and re-routing, and few studies have addressed the issue of ensuring determinism. This paper proposes a novel Time-Triggered Ad Hoc (TTA) network framework. It uses a Time Division Multiple Access (TDMA)-based time-triggered transmission mechanism to achieve self-organized, deterministic communication. The framework includes time synchronization and offline scheduling to optimize transmission performance. Experimental results show that the TTA framework outperforms traditional methods in terms of network capacity, latency, and jitter, demonstrating its effectiveness in solving the determinism problem in Ad Hoc networks.
Runqi Hu, Zonghui Li, Bo Ai 0001, Zhibo Pang
INDIN4
2025 DRM-CQF: Enhanced Deterministic Transmission between Profinet and TSN
abstract
Time-sensitive networking (TSN) is an important research direction for the transformation and upgrading of industrial internet infrastructure. In future industrial sites, TSN and traditional industrial networks will coexist in the same network, and this integration will be inevitable. Ensuring reliable and deterministic transmission of data flows in the converged network of Profinet and TSN will be a key research topic. This paper presents a compatible way for the Cyclic Queuing and Forwarding (CQF) queuing model of TSN and the Isochronous Real-Time (IRT) communication of Profinet. Firstly, we propose a Delay Reservation Mechanism based on CQF (DRM-CQF). This mechanism achieves reliable and deterministic transmission by delaying the sending time of cross-domain data flows in the Profinet and reserving transmission opportunities for cross-domain data flows in TSN. Secondly, we construct a mathematical optimization model based on DRM-CQF to schedule data flows in the converged network to seek the optimal schedule. Experimental results show that DRM-CQF can ensure the reliable transmission of cross-domain data flows in the Profinet and TSN converged network, and the end-to-end average delay is reduced by 49% compared with other CQF scheduling methods.
Chunxi Li, Zonghui Li, Zhibo Pang
INDIN3
2025 QoS-Guaranteed Joint User Clustering, Beam Selection, and Power Allocation for 5G mmWave MIMO-NOMA Transmission in Industrial Scenarios
abstract
With the advancement of 5G technologies, the integration of millimeterwave, multiple-input multiple-output and non-orthogonal multiple access (mmWave MIMO-NOMA) has attracted increasing attention in the field of the industrial internet of things (IIoT). This paper focuses on the three settings of 5G mmWave MIMO-NOMA downlink transmission system, i.e., user clustering, beam selection, and power allocation, for users with diverse quality of service (QoS) requirements in industrial scenarios. The goal is to maximize the weighted sum rate of different IIoT users by joint user clustering, beam selection and power allocation. We formulate the problem and divide it into integer programming and continuous programming sub-problems, which are respectively solved by a hierarchical clustering-based user clustering and QoS-based beam selection algorithm and a QoS-based deep deterministic policy gradient (DDPG) power allocation algorithm for sub-optimal solutions. Simulation results show that compared with baselines our proposed method can greatly improve the system weighted sum rate, while meeting QoS requirements of different users.
Xiaobing Zhong, Hongyun Zheng, Zonghui Li
INDIN4
2025 KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
abstract
Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, We introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on nine state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems.
Yongliang Wu, Zonghui Li, Xinting Hu, Xinyu Ye, Xianfang Zeng, Bernt Schiele, Ming-Hsuan Yang 0001, Xu Yang 0004
NeurIPS2
2025 Time Synchronization for 5G and TSN Integrated Networking
abstract
Emerging industrial applications involving robotic collaborative operations and mobile robots require a more reliable and precise wireless network for deterministic data transmission. To meet this demand, the 3rd Generation Partnership Project (3GPP) is promoting the integration of 5th Generation Mobile Communication Technology (5G) and Time-Sensitive Networking (TSN). Time synchronization is essential for deterministic data transmission. Based on the 3GPP’s vision of the 5G and TSN integrated networking with interoperability, we improve the time synchronization of TSN to conquer the multi-gNB competition, re-transmission, and mobility problems for the integrated 5G time synchronization. We implemented the improvement mechanisms and systematically validated the performance of 5G+TSN time synchronization. Based on the simulation in 500m x 500m industrial environments, the improved time synchronization achieved a precision of 1 microsecond with interoperability between 5G nodes and TSN nodes.
Zonghui Li, Bo Ai 0001
IEEE J. Sel. Areas Commun.2
2025 Deterministic Transmission for the Asynchronous Converged Networks of Profinet and TSN
abstract
With the rapid growth of Industry 4.0, time-sensitive networking (TSN) has emerged as the new infrastructure for future industrial Internet of Things (IoT) communication. Ensuring the compatibility between TSN and legacy networks is inevitable. The ideal compatibility is to achieve deterministic interconnection and interoperability without changes in hardware and communication protocols, in other words, only using standard devices with software management. This paper targets the ideal compatibility of TSN and Profinet Isochronous Real Time (IRT). First, we propose an inter-domain Multiple Transmission Opportunity Mechanism (MTOM) to enable the asynchronous converged network of TSN and Profinet. The mechanism reserves multiple transmission time slots for cross-domain data flows to reduce their end-to-end delay and jitter. Second, we formulate an asynchronous scheduling model (ASM) based on MTOM to coschedule flows in inter-and-intra domains. Finally, a case study is performed on a typical industrial network. The experiment results demonstrate that the proposed MTOM can only use standard devices to achieve deterministic transmission of Profinet and TSN converged networks. Compared with previous asynchronous converged networks, the delay and jitter are reduced by 86% and 80% on average, respectively.
Chunxi Li, Yongxiang Zhao, Zonghui Li
IEEE J. Sel. Areas Commun.4
2025 FastScheduler: Polynomial-Time Scheduling for Time-Triggered Flows in TSN
abstract
Time-Sensitive Networking (TSN) has emerged as a promising network paradigm for time-critical applications, such as industrial control, where flow scheduling is crucial to ensure low latency and determinism. As production flexibility demands increase, network topology and flow requirements may change, necessitating more efficient TSN scheduling algorithms to guarantee real-time and deterministic data transmission. In this work, we present FastScheduler, a polynomial-time, deterministic TSN scheduler, which can schedule thousands of Time-Triggered (TT) flows within arbitrary network topologies. The key innovations of FastScheduler include an Equivalent Reduction Technique to simplify the generic model while preserving the feasible scheduling space, a Deterministic Heuristic Strategy to ensure a consistent and reproducible scheduling process, and a Polynomial-Time Scheduling Algorithm to perform dynamic and real-time scheduling of periodic TT flows. Extensive experiments on various topologies show that FastScheduler can effectively simplify the model, reducing variables/constraints by 35%/62%, and schedule 1,000 TT flows in subsecond time. Furthermore, it runs 2/3 orders of magnitude faster and improves the schedulability by 12%/20% compared to heuristic/deep reinforcement learning-based methods. FastScheduler is well-suited for the dynamic requirements of industrial control networks.
Zonghui Li
IEEE Trans. Netw. Serv. Manag.5
2025 A 57.2 nW, 1.3-5 V VIN, -85 dB PSRR, 50 μs Start-Up Time, Bandgap Reference Circuit
abstract
This article presents a low-power bandgap reference (BGR) featuring high power supply rejection ratio (PSRR) and fast start-up capability, operating across a wide supply voltage range of 1.3–5 V. A novel prebiased pulse current injection technique is proposed in the start-up circuit, achieving a 1% settling time of$50~\mu $s and a$25\times $speed gain during start-up. To enhance supply noise immunity, the proposed BGR employs a preregulated (PR)-based amplifier that effectively decouples the reference voltage from supply voltage fluctuations. Fabricated in a 0.18-$\mu $m BCD process, the proposed reference occupies an active area of 0.0394 mm2. Under a 5 V supply, the circuit generates a 1.2 V reference voltage while consuming only 48 nA quiescent current. Operating down to a minimum supply voltage of 1.3 V, it maintains a low power consumption of 57.2 nW at room temperature. The reference exhibits an average temperature coefficient (TC) of 5.95 ppm/°C across a wide temperature range ($- 40~^{\circ }$C to$125~^{\circ }$C) and achieves an outstanding line sensitivity (LS) of 0.00308%/V over the 1.3–5 V supply range. Furthermore, the measured PSRR reaches −85 dB at 100 Hz.
Zonghui Li, Yani Li, Libo Qian, Zhangming Zhu
IEEE Trans. Very Large Scale Integr. Syst.1
2024 A Verification Framework for Time-Triggered Networks Based on Timed Colored Petri Net
abstract
Time-triggered (TT) network provides a low-cost service to meet the strong demand of modern industry networks for real-time communication. Both simulation and reachability analysis provide effective research methods for TT networks. This paper presents a verification framework for the TT network based on Timed Colored Petri Nets (TCPN). We propose an automatic formal modeling method for the behavior of message transmissions in the TT network. We harness timed multisets of TCPN to model and analyze time-triggered message transmission latencies. For a holistic system evaluation, we characterize the critical system properties such as boundedness and liveness. We substantiate the properties by the reachability analysis. We demonstrate the effectiveness through a case study by simulating the message transmission and reachability analysis in the state space. Finally, we analyze the influencing factors of the state space in the automotive scenario. The proposed automatic modeling method can effectively reduce the state space scale.
Wenjie Zhong, Jiantao Zhou 0002, Tao Sun 0002, Zonghui Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 A Deterministic Embedded End-System Tightly Coupled With TSN Schedule
abstract
Distributed real-time systems (DRTSs) composed of many embedded end-systems have been widely adopted in the industrial fields. Time-sensitive networking (TSN), as a promising communication infrastructure for DRTS, has shown great potential in industry and academia. TSN assumes that an end-system can release critical tasks to process critical packets strictly according to the prescheduled time. Unfortunately, two factors currently damage this assumption: 1) the jitter caused by system architecture during task release, task execution, and packet transmission; and 2) a TSN schedule result may exceed the execution capability of the end-system and cause conflicts. This article proposed deterministic chip (DetChip), a system-on-chip capable of deterministically implementing the TSN schedule result. DetChip supports time-triggered task release, time-predictable task execution and precise network transmission. Based on DetChip, this article first formalizes the execution capability of the end-system as end-system constraints (ECs). Existing TSN scheduling algorithms integrating ECs can solve conflicts by co-scheduling end-systems and the TSN network. Compared with previous works, DetChip only introduces few clock cycles jitter for critical task execution according to the TSN schedule result. The proposed ECs obtain$2\times$–$8\times$more conflict-free solutions for advanced scheduling algorithms with a linear increase in time overhead. Compared with the general end-system, DetChip can reduce$6\times$–$20\times$processing jitter to achieve better clock synchronization.
Chenglong Li 0007, Zonghui Li, Tao Li 0008, Cunlu Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Agglomerative Memory and Thread Scheduling for High-Performance Ray-Tracing on GPUs
abstract
Ray-tracing rendering has long been considered as a promising technology to enable a higher level of visual experience. The democratization of the ray-tracing rendering to consumer platforms, however, poses significant challenges to rendering hardware and software due to its highly irregular computing patterns. In fact, modern ray-tracing techniques typically depend on a tree-based acceleration structure to reduce the computing complexity of intersection testing of rays and graphics primitives. The traversal by a massive number of rays on a graphics processing unit (GPU) incurs a significant amount of irregular memory traffic, which turns out to be a major stumbling block for real-time performance. In this work, a scheduling mechanism, so-called Agglomerative Memory and Thread Scheduling, is proposed to unleash the inherence parallelism in the ray-tracing process on GPUs. It is associated with a tile-based ray-tracing framework in which the acceleration structure (i.e., KD-tree in this work) is partitioned into subtrees that can be completely loaded into the on-chip L1 cache inside a streaming multiprocessor. An effective scheduling mechanism collects threads with regard to the subtrees hit by their respective rays and regroup threads into warps for dispatching. In addition, subtrees are dynamically preloaded into the L1 cache of multiprocessors in an on-demand fashion. The proposed scheduler can be integrated on today’s high-end GPUs with only minor overhead. Microarchitecture simulation results prove that the proposed framework significantly improves memory efficiency and outperforms a traditional GPU microarchitecture by 47.4% for average.
Yufei Ni, Yangdong Deng, Zonghui Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Duplicacy: A New Generation of Cloud Backup Tool Based on Lock-Free Deduplication
abstract
The pervasive deployment of cloud services poses an ever-increasing demand for cross-client deduplication solutions to save network bandwidth, lower storage costs, and improve backup speeds. However, existing solutions typically depend on lock-based approaches relying on a centralized chunk database, which tends to hinder performance scalability. In this article, we present a new cross-client cloud backup solution, named Duplicacy, based on a Lock-Free Deduplication approach. Lock-Free Deduplication stores chunks to network or cloud storage using content hashes as file names. It then adopts a two-step fossil deletion algorithm to solve the hard problem of deleting unreferenced chunks in the presence of concurrent backups, without the need for any locks. Experiments demonstrate that Duplicacy enables significant performance improvement for backups over previous well-known backup tools. In addition, Duplicacy can work with many general-purpose network or cloud storage services which only support a basic set of file operations, and turn them into sophisticated deduplication-aware storage servers without server-side changes.
Zonghui Li, Gilbert Chen, Yangdong Deng
IEEE Trans. Cloud Comput.1
2021 Flow Scheduling for Conflict-Free Network Updates in Time-Sensitive Software-Defined Networks
abstract
The digital transformation of industry requires industrial control networks provide high flexibility and determinacy. Time-sensitive software-defined networking that combines time-sensitive networking and software-defined networking is a new network paradigm which provides both real-time transmission feature and network flexibility. During network updates, the transmission consistency needs to be maintained. However, previous mechanisms mostly target on the proper schedule transition, which cannot guarantee no frame loss and also introduces extra update overhead. The article proposes a novel flow schedule generation model which guarantees no frame loss during network updates even with the basic two-phase update mechanism and introduces no extra update overhead. Two algorithms are designed for the model to adapt to different application scenarios: the offline algorithm poses better schedulability, whereas the online one consumes less time with slightly decreased schedulability. The experiments on two real-world industrial networks demonstrate our mechanism achieves zero frame loss without extra update overhead compared to existing methods, and the online algorithm saves 40% execution time with at most 10% schedulability decrease when the bandwidth utilization is less than 50%.
Zaiyu Pang, Zonghui Li, Sukun Zhang, Yanfen Xu, Hai Wan, Xibin Zhao
IEEE Trans. Ind. Informatics3
2020 Time-Triggered Switch-Memory-Switch Architecture for Time-Sensitive Networking Switches
abstract
Time-sensitive networking (TSN) is a set of extended standards for the IEEE 802.3 Ethernet under development by the IEEE 802.1 TSN task group. TSN depends on two key components, scheduling and fault tolerance, to provide realtime and reliable transmission. There is a strong motivation to replace the widely used field-buses with TSNs in industrial networking applications. However, industrial network devices are typical application-specific embedded systems with limited memory resources. Time-sensitive (TS) transmission certainly prefers on-chip memory, which is even more scarce for embedded systems. As a result, it is critical for TSNs to develop memory-efficient switching techniques with scalable schedulability and elegant fault-tolerance support. This paper proposes a time-triggered switch-memory-switch (SMS) architecture for memory-efficient TSN switches. First, based on the SMS shared memory, our architecture makes it possible to statically schedule memory allocation with full utilization for TS traffic and the remaining memory for other traffic. Compared with perport memory, the shared memory achieves a ratio of (nn/n!) (≈ (en/√(2πn)), n → ∞), where n is the port number, in the feasible solution space under memory constraints and thus significantly improves scheduling memory ability and flexibility. Moreover, we develop a fault-tolerance scheme for reliable transmission. It facilitates a memory-efficient implementation of the popular multiline redundancy in industrial networks. The scheme is validated by five classes of memory conflicts and a case study on two-line redundancy.
Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Model-Based Adaptation of Mixed-Criticality Multiservice Systems for Extreme Physical Environments
abstract
An increasingly important trend in the design of industry-strength embedded systems is the integration of multiple services with varying criticality levels into a common computing platform. Such systems are characterized as mixed-criticality multiservice systems (MCMSs). An MCMS has to survive in rigorous environments posed by industry-level requirements. Such survival, however, is becoming continuously more challenging due to the growing system complexity and integrating more and more services. While existing works typically target reliability-driven design optimization to improve the system robustness rather than deal with the surviving problem of the system in extreme physical environments, this paper addresses the problem by enabling the service capability transitions of an MCMS to adapt to the environments. This paper proposes a service capability model to capture the importance of functional modules for the criticality of different services. A model-based service-capability transition mechanism is designed to automatically identify the maximum allowed service capability under a given physical environment. A case study of the proposed techniques was performed on an industrial Ethernet switch which is a typical MCMS, to validate the capability of adaptation to high and low temperatures. The experimental results demonstrate the significant potential of our approach to improve system survivability under extreme physical environments.
Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 An Enhanced Reconfiguration for Deterministic Transmission in Time-Triggered Networks
abstract
The emerging momentum of digital transformation of industry, i.e. Industry 4.0, poses strong demands for integrating industrial control networks, and Ethernet to enable the real-time Internet of Things (RT-IoT). Time-triggered (TT) networks provide a cost-efficient integrated solution while RT-IoT arouses the reconfiguration challenges: the network has to be flexible enough to adapt to changes and yet provides deterministic transmission persistently during network reconfiguration. Software defined network benefits the flexible industrial control by configuring the rules handling frames. However, previous reconfiguration mechanisms are mostly oriented to the context of data centers and wide area networks and thus do not consider the deterministic transmission in TT networks. This paper focuses on the reconfiguration (i.e., updates) for the deterministic transmission. To minimize the overhead during updates, namely the minimum number of loss frames and the minimum duration time of updates, we first establish an update theory based on the dependence relationship derived by the conflicts during updates. In addition then the reconfiguration problem is modeled with the dependence graph built by the relationship. On such a basis, we present a reconfiguration mechanism and its implementation to solve the problem. Finally, we evaluate the proposed reconfiguration mechanism in two real industrial network topologies. The experimental results demonstrate that compared with previous methods, our mechanism significantly reduces the number of loss frames and achieves zero loss in almost all cases.
Zonghui Li, Hai Wan, Zaiyu Pang, Qiubo Chen, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE/ACM Trans. Netw.1
2018 Work-in-Progress: A Flattened Priority Framework for Mixed-Criticality Real-Time Systems
abstract
Recent years witnessed a fast growing popularity of mixed-criticality real-time applications on smart devices. Priority schedulers are typically the central component to provide differential quality of service (QoS) for mixed-criticality tasks. The increasing deployment of such applications on smart devices, however, poses new challenges for the design of effective schedulers. First, the scheduling algorithms for mixed-criticality tasks are generally NP-complete. Second, the scheduling algorithms have to be effective enough under the limited computing resource of smart devices. This paper presents a work-in-progress report on a novel technique to design efficient and effective mixed-criticality schedulers. We propose a flattened priority framework to transform a non-priority scheduler into a priority one. The framework is typically a iterative framework based on feedback loops. Given an optimal nonpriority scheduler, for P priorities, the transformed scheduler converges with P iterations in the worst case. With the proposed framework, the design of priority schedulers is simplified into the design of non-priority schedulers. Such a simplification dramatically lowers the design effort and system complexity. A case study was performed on FPGA-based Industrial Ethernet switches. The proposed method achieves a 30%~50% reduction in the usage of look-up tables (LUTs) without performance loss.
Zonghui Li, Hai Wan, Yangdong Deng, Ming Gu 0001
RTAS1
2017 Path compression kd-trees with multi-layer parallel construction a case study on ray tracing
abstract
Kd-tree is a fundamental data structure with extensive applications in computer graphics. The performance of many interactive applications such as real-time ray tracing hinges on the construction and traversal efficiency of kd-trees. In recent years, there is a pressing demand for accelerating the construction process due to the fast-growing need of handling dynamic scenes. Existing construction algorithms typically follow a layer-by-layer scheme, which significantly limits the efficiency on the use of multi-core CPUs and GPUs. In this paper, we propose a concurrent multi-layer kd-tree construction algorithm to unleash the inherent parallelism. For a given scene, the algorithm uses Morton code to split its bounding box and orders primitives by Morton curve. A path compression procedure is then concurrently executed on all essential nodes that contain primitives to generate the hierarchy in the target kd-tree. All redundant nodes that have no primitives along the compression paths are collapsed to fast slip empty space. The fully parallel algorithmic scheme adapts variable primitives space and drastically shortens the construction time. A case study on ray tracing benchmarks demonstrates that our kd-tree construction method outperforms the state of art work by an average factor of over 10 and still enables high performance traversal.
Zonghui Li, Yangdong Deng, Ming Gu 0001
I3D1
2015 FastTree: a hardware KD-tree construction acceleration engine for real-time ray tracing
Yangdong Deng, Yufei Ni, Zonghui Li
DATE4
2015 RadixBoost: A hardware acceleration structure for scalable radix sort on graphic processors
abstract
In this paper, we propose RadixBoost, a hardware acceleration structure for scalable 32-bit integer radix sort on GPU. The whole structure is integrated into a GPU microarchitecture as a special functional unit and can be started by new instructions. Our design enables a significantly faster sorting procedure for general purpose GPU computing. The RadixBoost architecture was validated by an FPGA prototype integrated in FPGA-based GPU microarchitecture simulator, Fastlanes. An ASIC evaluation of RadixBoost was also performed. Our results proved that RadixBoost outperformed its GPU software equivalent by a factor of over 6 with an 1% and 3% increase in area and power respectively in cutting-edge Fermi GPU.
Shikai Li, Kuan Fang, Yufei Ni, Zonghui Li, Yangdong Deng
ISCAS5
2014 Fully parallel kd-tree construction for real-time ray tracing
abstract
This work proposes a fully parallel kd-tree construction algorithm, which depends on the Morton code to identify all candidate split planes and derive their exact positions in parallel. Our techniques drastically shorten construction process. Experimental results on a set of frequently used scenes prove that the proposed kd-tree construction algorithm outperforms a state-of-the-art algorithm kd-tree construction algorithm by over one order of magnitude.
Zonghui Li, Tong Wang 0036, Yangdong Deng
I3D1
2013 FastLanes: An FPGA accelerated GPU microarchitecture simulator
abstract
Graphic Processing Units (GPUs) have emerged as a new general purpose computing platform that attracts significant research efforts. Currently, GPU architecture research resorts to time-consuming software simulations to evaluate microarchitecture innovations. In this paper, we propose FastLanes, an FPGA based simulator for a generic GPU microarchitecture, to enable hardware-accelerated simulation. FastLanes consists of a function model and a timing model, both implemented on FPGA. The functional model implements the full functionality of a multiprocessor of GPU and emulates multiple multiprocessors via time-division multiplexing. We develop a hybrid implementation strategy in which certain GPU logic is directly mapped to FPGA while the other logic is simulated by reusing the same FPGA logic. A corresponding context shifting mechanism is proposed to store execution states of threads from FPGA to external on-board memory, and vice versa. Such a mechanism makes it possible to simulate hundreds of GPU cores on a single FPGA evaluation board. Driven by the functional simulation results, the timing model considers the detailed configuration of GPU microarchitecture to derive the performance evaluation. A compiler tool-chain is also developed to allow the execution of NVIDIA GPU binary on FastLanes. Experimental results prove that FastLanes outperforms its software equivalent by up to 2 orders of magnitude.
Kuan Fang, Yufei Ni, Jiayuan He 0005, Zonghui Li, Shuai Mu 0002, Yangdong Deng
ICCD4
2013 Design and optimization of multi-clocked embedded systems using formal technique
abstract
Today’s system-on-chip and distributed systems are commonly equipped with multiple clocks. The key challenge in designing such systems is that heterogenous control-oriented and data-oriented behaviors within one clock domain, and asynchronous communications between two clock domains have to be captured and evaluated in a single framework. In this paper, we propose to use timed automata and synchronous dataflow to capture the dynamic behaviors of multi-clock embedded systems. A timed automata and synchronous dataflow based modeling and analyzing framework is constructed to evaluate and optimize the performance of multiclock embedded systems. Data-oriented behaviors are captured by synchronous dataflow, while synchronous control-oriented behaviors are captured by timed automata, and inter clock-domain asynchronous communication can be modeled in an interface timed automaton or a synchronous dataflow module with the CSP mechanism. The behaviors of synchronous dataflow are interpreted by some equivalent timed automata to maintain the semantic consistency of the mixed model. Then, various functional properties can be simulated and verified within the framework. We apply this framework in the design process of a sub-system that is used in real world subway communication control system
Yu Jiang 0001, Zonghui Li, Hehua Zhang, Yangdong Deng, Ming Gu 0001, Jia-Guang Sun 0001
ESEC/SIGSOFT FSE2