Weitao Pan

dblp:181/4789 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
19since 2021 · last 2026
0000-0002-6388-5008ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 8 since 2021Computer networks · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing mixed-criticality scheduling in time-sensitive networks: Performance analysis and improved transmission window arrangement using sub-period partition
Yingge Feng, Li Zhen, Weitao Pan
Comput. Networks5
2026 Enhancing deterministic transmission in Time-Sensitive Networking: A Joint Guard Band Compression and Non-Disruptive Frame Preemption model
Jingzhuo Liu, Keyao Zhang, Qianxi Men, Li Zhen, Weitao Pan
J. Syst. Archit.6
2026 CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-Based CIM Architectures
abstract
Compute-in-memory (CIM) has emerged as a pivotal direction for accelerating workloads in the field of machine learning, such as Deep Neural Networks (DNNs). However, the effectively exploitation of sparsity in CIM systems presents numerous challenges, due to the inherent limitations in their rigid array structures. Designing sparse DNN dataflows and developing efficient mapping strategies also become more complex when accounting for diverse sparsity patterns and the flexibility of a multi-macro CIM structure. Despite these complexities, there is still an absence of a unified systematic view and modeling approach for diverse sparse DNN workloads in CIM systems. In this paper, we propose CIMinus, a framework dedicated to cost modeling for sparse DNN workloads on CIM architectures. It provides an in-depth energy consumption analysis at the level of individual components and an assessment of the overall workload latency. We validate CIMinus against contemporary CIM architectures and demonstrate its applicability in two use-cases. These cases provide valuable insights into both the impact of sparsity patterns and the effectiveness of mapping strategies, bridging the gap between theoretical design and practical implementation.
Yingjie Qi, Jianlei Yang 0001, Rubing Yang, Cenlin Duan, Xiaolin He, Ziyan He, Weitao Pan, Weisheng Zhao 0001
IEEE Trans. Computers7
2026 A Grouped Sorting Queue Supporting Dynamic Updates for Timer Management in High-Speed Network Interface Cards
Binghao Yue, Weitao Pan, Jiangyi Shi, Yue Hao 0001
IEEE Trans. Computers3
2025 Queuing Model with Setup and Multi-Vacations for Mixed-Critical Services in Time-Sensitive SAGIN
abstract
The space-air-ground integrated network(SAGIN), as a core architecture of 6G communications, provides ubiquitous coverage for high-reliability and low-latency scenarios such as industrial Internet and autonomous driving. However, its dynamic environment faces multiple challenges, including resource competition between heterogeneous space-ground services, rapid topology changes, and Doppler shifts. To address the issues of delay jitter and queue congestion caused by mixed transmission of time-triggered (TT) and event-triggered (ET) services in time-sensitive networking (TSN), this paper proposes an analytical framework based on queuing theory with setup periods and multiple working vacations. This framework systematically analyzes the transmission dynamics of mixed services in space-air-ground integrated channels. By constructing a micro-level parameter coupling model, the regulatory effects of the setup period parameter on state transition efficiency, the influence mechanism of the vacation period parameter on resource release cycles, and the dynamic correlation between packet arrival rates and busy period ratios are revealed. The results demonstrate that traditional TSN static resource allocation strategies exhibit multi-dimensional mismatches in space-air-ground scenarios, necessitating dynamic scheduling algorithms to optimize time window orchestration and priority preemption mechanisms. The theoretical model proposed in this paper provides novel insights for adaptive resource scheduling in space-air-ground integrated networks, advancing TSN from fixed-topology systems to space-ground collaborative architectures. This work lays a technical foundation for realizing the 6G vision of "ubiquitous coverage and on-demand service."
Yingge Feng, Weitao Pan
VTC2025-Fall4
2025 Deadline Guaranteed Scheduling Based on Time-Ordered Queues for Time-Sensitive Low-Earth-Orbit Satellite Networks
abstract
Low-Earth-orbit(LEO) satellite communication, as a key technology in the future networking field, which can provide ubiquitous mobile communication services. However, the LEO constellation faces multiple challenges such as low throughput of onboard switching systems, large transmission delay, and the limited computing and storage resources on satellites, making it difficult to meet the requirement of the time-sensitive applications such as the Internet of Things and industrial Internet. To realize LEO-based time-sensitive networking (TSN), this paper proposes a deadline guaranteed scheduling algorithm based on time-ordered queues for time-sensitive LEO satellite networks. The key idea of the proposed algorithm is to simplify the deadline guaranteed scheduling problem of multiple periods into a single period. Thus, the problem of high storage complexity of scheduling table caused by period extension is solved successfully. To do this, the Push-In-First-Out(PIFO) queues, usually used in terrestrial networks, are configured at the output end of the onboard switching fabric. Through a reordering mechanism based on the deadline priority, packets are pushed into the appropriate position for waiting and then output in time order. Simulation results indicate that compared with traditional scheduling algorithms, the proposed algorithm saves 45% - 97% of the time flow table storage resources of the onboard switching system under the same conditions, effectively improving resource utilization. The throughput performance of the onboard switch is enhanced by approximately 15% - 24%, and the average delay is reduced by around 17% - 49%.
Guodong Wei, Qianxi Men, Yingge Feng, Weitao Pan
VTC2025-Fall5
2025 Group-based windows scheduling method for non-deterministic periodic flows in time-sensitive networks
abstract
Time-triggered (TT) flows are usually periodic in time-sensitive networks. However, nondeterministic end systems can generate TT flow frames with significant jitter (i.e., jittery TT flows). Jitter can cause frames to miss the TT windows scheduled for the current period, resulting in excessive access delays, which in turn affect the end-to-end deterministic transmission of the TT flows. In our previous study, we proposed the use of a dynamic multiwindow approach to achieve deterministic access to jittery TT flows; however, its window schedule computation is too slow, and this method is only suitable for small networks with a few TT flows. We therefore propose a group-based, fast scheduling method for accessing and transmitting the windows of jittery TT flows based on multiple windows. A combination of heuristic algorithms and solvers, including the establishment of TT window groups, division of the solution region, and integrated parallel and serial incremental coarse- and fine-grained computations, significantly improves the efficiency of TT window scheduling. For coarse-grained scheduling, by establishing large window clusters and central alignment, the complexity of scheduling is considerably reduced while keeping success rates high. Furthermore, the integer linear programming constraints and objective functions for this method are provided. Compared with the conventional dynamic multiwindow approach, the proposed approach reduces the scheduling time for TT windows by two orders of magnitude for a small star network with a small number of jittery TT flows. Moreover, the reduction in scheduling time becomes more pronounced as the network topology complexity and number of jittery TT flows increase. Finally, the scheduling time performance of the proposed method is verified in commonly used star, tree, and bus networks. Evaluations demonstrate that the access and transmission windows for 500 jittery TT flows can be scheduled in these networks, enabling deterministic access and significantly improving scheduling efficiency.
Yongjun Li 0002, Qin Tian, Kai Zhang 0034, Weitao Pan
Comput. Networks7
2025 Hardware Trojan Detection Methods for Gate-Level Netlists Based on Graph Neural Networks
abstract
Currently, untrusted third-party entities are increasingly involved in various stages of IC design and manufacturing, posing a significant threat to the reliability and security of SoCs due to the presence of hardware Trojans (HTs). In this paper, gate-level HT detection methods based on graph neural networks (GNNs) are established to overcome the defects of existing machine learning, which makes it difficult to characterize circuit connection relationships. We introduce harmonic centrality in the feature engineering of gate-level HT detection, which reflects the positional information of nodes and their adjacent nodes in the graph, thereby enhancing the quality of feature engineering. We use the golden section weight optimization algorithm to configure penalty weights to alleviate the problem of extreme data imbalance. In the SAED database, GraphSAGE-LSTM model obtained a TPR of 88.06% and an average F1 score of 90.95%. In the combined HT netlist of LEDA datasets, GraphSAGE-POOL model obtains a TPR of 88.50% and the best F1 score of 92.17%. In sequential HT netlist, GraphSAGE-LSTM model performs optimally, with a TPR of 98.25% and an average F1 score of 98.59%. Compared to existing detection models, the F1 score is enhanced by 8.86% and 2.48% on combined and sequential HT datasets, respectively.
Peijun Ma, Hongjin Liu, Jiangyi Shi, Weitao Pan, Yue Hao 0001
IEEE Trans. Computers6
2025 GNN-Based Hardware Trojan Detection at Register Transfer Level Leveraging Multiple-Category Features
abstract
The existing hardware Trojan (HT) detection technology usually relies on the golden reference model. With the continuous improvement of circuit integration, the detection accuracy of traditional methods such as side-channel analysis has declined. Methods based on testability and switch probability analysis have shown high detection accuracy. However, these techniques have a limited detection scope and are generally ineffective at identifying HT where the Sandia controllability/observability analysis program (SCOAP) values or switch probabilities are similar to those of normal signals. Against this backdrop, this article proposes a detection method based on graph neural networks (GNNs), which can achieve HT detection at the register transfer level (RTL) without the golden reference model. First, the RTL code is transformed into a data flow graph (DFG), and node feature extraction and node label marking are carried out during the transformation process. To mitigate the impact of insufficient initial features on the GNN model performance, the node feature vector used in this article comprises 37-D node types and 6-D structural features such as the minimum distance from the primary input (PI) and primary output (PO), in-degree, and out-degree. Subsequently, several GNN models are built for node classification tasks. The best model achieves an average of 99.1% recall and 96.7% F1-score on the open-source dataset on the Trust-Hub platform. Compared to the state-of-the-art detection results at RTL, the F1-score in this article has increased by an average of 3.8%.
Peijun Ma, Ge Shang, Hongjin Liu, Jiangyi Shi, Weitao Pan, Yue Hao 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Corrections to "GNN-Based Hardware Trojan Detection at Register Transfer Level Leveraging Multiple-Category Features"
abstract
Presents corrections to the paper, (Corrections to “GNN-Based Hardware Trojan Detection at Register Transfer Level Leveraging Multiple-Category Features”).
Peijun Ma, Ge Shang, Hongjin Liu, Jiangyi Shi, Weitao Pan, Yue Hao 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2024 Design and implementation of a frame preemption model without guard bands for time-sensitive networking
abstract
In time-sensitive networks, the frames of time-triggered (TT) flows need to be transmitted in scheduled slots. To avoid the interference of frames from other flows to the frames of the TT flows, guard bands are generally reserved prior to scheduled slots. The time-sensitive networking (TSN) standard, IEEE 802.1Qbu, specifies a frame preemption model that enables frames of TT flows (preemption frames) to preempt frames of other flows (preempted frames). According to this model, preempted frames are divided into fragments that are transmitted in the gaps of preemption frames. Consequently, the IEEE 802.1Qbu model reduces the guard band from the longest Ethernet frame (typically 1518 bytes) to 123 bytes, which improves the bandwidth utilization and delay performance of preempted frames without causing frame disorder. However, in scenarios in which the preempted frame load is heavy, and the length is short, numerous guard bands smaller than 123 bytes are generated. These guard bands prevent the IEEE 802.1Qbu model from transmitting preempted frame fragments, resulting in a considerable decrease in bandwidth utilization. To solve this problem, we propose a novel frame preemption model without guard bands, based on a padding and splicing mechanism. While ensuring that preemption frames are transmitted according to scheduled slots, this model can transmit preempted frames within arbitrary byte gaps based on the schedule without obtaining the length of the preempted frames in advance. We compared the transmission-delay performance of the proposed model with that of the IEEE 802.1Qbu model using a theoretical analysis. The evaluation results obtained from heavily loaded preemption frame scenarios revealed that the proposed model improved link utilization by 32.8% relative to the IEEE 802.1Qbu model and reduced the transmission delay by more than one order of magnitude. Moreover, when the IEEE 802.1Qbu model fails, the proposed model still transmits 60% of preempted frames.
Zhiliang Qiu, Weitao Pan, Ya Gao 0003
Comput. Networks3
2024 DDC-PIM: Efficient Algorithm/Architecture Co-Design for Doubling Data Capacity of SRAM-Based Processing-in-Memory
abstract
Processing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other nonvolatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states$(Q/\overline {Q})$, thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about$2.84\times $speedup on MobileNetV2 and$2.69\times $on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to$8.41\times $and$2.75\times $improvement in weight density and area efficiency, respectively.
Cenlin Duan, Jianlei Yang 0001, Xiaolin He, Yingjie Qi, Yiou Wang, Ziyan He, Bonan Yan, Xiaotao Jia, Weitao Pan, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.11
2024 An Energy-Efficient Bayesian Neural Network Implementation Using Stochastic Computing Method
abstract
The robustness of Bayesian neural networks (BNNs) to real-world uncertainties and incompleteness has led to their application in some safety-critical fields. However, evaluating uncertainty during BNN inference requires repeated sampling and feed-forward computing, making them challenging to deploy in low-power or embedded devices. This article proposes the use of stochastic computing (SC) to optimize the hardware performance of BNN inference in terms of energy consumption and hardware utilization. The proposed approach adopts bitstream to represent Gaussian random number and applies it in the inference phase. This allows for the omission of complex transformation computations in the central limit theorem-based Gaussian random number generating (CLT-based GRNG) method and the simplification of multipliers as AND operations. Furthermore, an asynchronous parallel pipeline calculation technique is proposed in computing block to enhance operation speed. Compared with conventional binary radix-based BNN, SC-based BNN (StocBNN) realized by FPGA with 128-bit bitstream consumes much less energy consumption and hardware resources with less than 0.1% accuracy decrease when dealing with MNIST/Fashion-MNIST datasets.
Xiaotao Jia, Huiyi Gu, Jianlei Yang 0001, Weitao Pan, Youguang Zhang, Sorin Cotofana, Weisheng Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 A Machine Learning Based Approach to Detect Machine Learning Design Patterns
abstract
As machine learning expands to various domains, the demand for reusable solutions to similar problems increases. Machine learning design patterns are reusable solutions to design problems of machine learning applications. They can significantly enhance programmers' productivity in programming that requires machine learning algorithms. Given the critical role of machine learning design patterns, the automated detection of them becomes equally vital. However, identifying design patterns can be time-consuming and error-prone. We propose an approach to detect their occurrences in Python files. Our approach uses an Abstract Syntax Tree (AST) of Python files to build a corpus of data and train a refined Text-CNN model to automatically identify machine learning design patterns. We empirically validate our approach by conducting an exploratory study to detect four common machine learning design patterns: Embedding, Multilabel, Feature Cross, and Hashed Feature. We manually label 450 Python code files containing these design patterns from repositories of projects in GitHub. Our approach achieves accuracy values ranging from 80 % to 92% for each of the four patterns.
Weitao Pan, Hironori Washizaki, Nobukazu Yoshioka, Yoshiaki Fukazawa, Foutse Khomh, Yann-Gaël Guéhéneuc
APSEC1
2023 Access mechanism for period flows of non-deterministic end systems for time-sensitive networks
abstract
The IEEE 802.1Qbv standard schedules time-triggered (TT) flows (i.e., period flows) in a fixed TT Window each period. However,period flows generated by non-deterministic end systems exhibit significant jitter, whichleads to a mismatch between the generation times and the scheduled TT Windows. In the worst case scenario, this mismatch can cause an additional send delay of approximately one period in the flow. In this study,we model the send delay issue due to the maximum send delay requirement for all frames in a period flow not being satisfied simultaneously, and measure the jitter of period flows in a typical non-deterministic end system. Moreover, we propose an access mechanism for jittered period flows to schedule multiple conflict-free TT Windows in each period of flows at the source end system and control the send delay within the required delay tolerance. This mechanism enables deterministic access to jittered period flows, providing a prerequisite for reliable end-to-end transmission in the network. Moreover, this mechanism adopts a multi-objective integer linear programming (MILP) solver to optimize the TT Windows schedule. Furthermore, we establish the constraints and objective functions for the MILP solver and evaluate the mechanism with actual sampled frames in a real non-deterministic end system. Compared with the conventional fixed single TT Window and worst-case delay analysis mechanisms, the proposed mechanism satisfies the send delay requirements and considerably reduces the buffer usage at the source end system.
Weitao Pan, Zhiliang Qiu, Ya Gao 0003
Comput. Networks3
2022 Eventor: an efficient event-based monocular multi-view stereo accelerator on FPGA platform
abstract
Event cameras are bio-inspired vision sensors that asynchronously represent pixel-level brightness changes as event streams. Event-based monocular multi-view stereo (EMVS) is a technique that exploits the event streams to estimate semi-dense 3D structure with known trajectory. It is a critical task for event-based monocular SLAM. However, the required intensive computation workloads make it challenging for real-time deployment on embedded platforms. In this paper, Eventor is proposed as a fast and efficient EMVS accelerator by realizing the most critical and time-consuming stages including event back-projection and volumetric ray-counting on FPGA. Highly paralleled and fully pipelined processing elements are specially designed via FPGA and integrated with the embedded ARM as a heterogeneous system to improve the throughput and reduce the memory footprint. Meanwhile, the EMVS algorithm is reformulated to a more hardware-friendly manner by rescheduling, approximate computing and hybrid data quantization. Evaluation results on DAVIS dataset show that Eventor achieves up to 24X improvement in energy efficiency compared with Intel i5 CPU platform.
Jianlei Yang 0001, Yingjie Qi, Meng Dong, Yuhao Yang 0008, Runze Liu 0001, Weitao Pan, Bei Yu 0001, Weisheng Zhao 0001
DAC7
2021 HyperParser: A High-Performance Parser Architecture for Next Generation Programmable Switch and SmartNIC
abstract
Programmable switches and SmartNICs motivate the programmable network. ASIC is adopted in programmable switches to achieve high throughput, and FPGA-based SmartNIC is becoming increasingly popular. The programmable parser is a key element in programmable switches and SmartNICs, which can identify the protocol types and extract the relevant fields. The programmable parser for the next generation programmable switches and SmartNICs requires a significant improvement in PPAL (performance, power, area, and latency), which is quite challenging. According to the Ethernet roadmap, 800 Gbps and 1.6 Tbps are expected to be the future switch interface speeds after 2022, which leads to higher throughput of the parser. Meanwhile, the end of Dennard scaling and the slowdown of Moore’s Law result in limited power and area. Besides, the need for low-latency and low-jitter operations at the datacenter scale continues to grow.
Huan Liu 0021, Zhiliang Qiu, Weitao Pan, Jinjian Huang
APNet3
2021 Combined Shared-Memory and Buffered-Crossbar Architecture for High-Bandwidth Onboard Switching Fabric
abstract
No abstract available.
Weitao Pan, Huan Liu 0021
APNet3
2021 Architecture design and performance analysis of a novel memory system for high-bandwidth onboard switching fabric
Weitao Pan, Ya Gao 0003, Huan Liu 0021
Comput. Networks2
2020 Dual-Plane Switch Architecture for Time-Triggered Ethernet
abstract
Time-triggered Ethernet (TTE) technology introduces the concept of time-triggered on the basis of traditional Ethernet, so that it can achieve conflict-free and deterministic service forwarding without sacrificing compatibility. However, storage resources in industrial, aviation, aerospace and other equipment are limited. Therefore, it is important for TTEthernet to develop switching technologies with high storage efficiency and scalability. This paper proposes a dual plane switching (DPS) architecture for TTEthernet, which divides time-triggered services and event-triggered services into two planes for data forwarding. Experimental results show that using the TTE switch of this architecture has the advantages of high clock synchronization accuracy, high throughout, low transmission delay and small jitter of TTE service.
Meng Dong, Zhiliang Qiu, Weitao Pan, Chenglei Kong, Jianlei Yang 0001
ACM Great Lakes Symposium on VLSI3