Jincheng Zhong

dblp:257/2831 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-9827-1981ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Achieving Unified Memory for FPGA-Based String Matching
abstract
String matching serves as a critical module for network security systems. To meet escalating network bandwidth demands, recent studies have transitioned to hardware platforms like FPGA, leveraging the parallel processing ability to accelerate string matching. However, existing hardware solutions face a critical challenge in parallel matching of variable-length string patterns. They require length-specific memory blocks, such as separate hash tables to store patterns of different lengths. This distributed memory architecture causes a memory fragmentation issue when patterns are unevenly distributed, which impacts the scalability of prior works. To address the memory fragmentation issue of the distributed memory architecture, this paper proposes (1) a unified memory architecture that enables unified storage of variable-length patterns, and (2) a collision-free hash scheme that supports parallel matching of variable-length strings in the architecture. We implemented the proposed memory-efficient scheme on an FPGA-based prototype. Extensive evaluations demonstrate that the unified memory architecture achieves 4-5× lower memory usage compared to state-of-the-art alternatives while achieving comparable throughput. Meanwhile, the architecture can store patterns across arbitrary length distributions within a specified length range until its maximum capacity.
Zhuoxuan Sun, Jincheng Zhong, Jiayao Wang 0002, Shuhui Chen
APNet2
2026 AdaptTree: A Practical and Adaptive Packet Classification Scheme on FPGA
Jincheng Zhong, Gaofeng Lv, Shuhui Chen
SECON1
2026 Doubling the speed of large-scale packet classification through compressing decision tree nodes
Jincheng Zhong, Gaofeng Lv, Shuhui Chen
Comput. Networks1
2026 Blocking Is Not Stagnation: A Synchronous FPGA-CPU Architecture for Regular Expression Matching in Real-Time DPI
abstract
Regular expression matching is a crucial step in traffic analysis. Many hardware-based architectures are proposed to improve the matching throughput, such as FPGA. To date, however, the existing FPGA-CPU architectures are difficult to implement in DPI systems due to the following two reasons. First, existing architectures use asynchronous workflows to interact data between FPGA and CPU, making them difficult to be compatible with synchronous DPI systems. Second, asynchronous architectures require batch input, which does not meet the requirements of real-time environments. In this paper, we concentrate on the real-time deployment of a regular expression matching architecture. To improve the deployment throughput, we propose an FPGA-CPU architecture with a parallel layer between the driver and DPI systems. Then, coroutines are introduced and proved to have significant advantages. Meanwhile, some optimization methods are proposed to address idle time, memory allocation, and MMIO control. Our experiments demonstrate that directly deploying an asynchronous architecture on a synchronous DPI would result in a throughput degradation of 3 orders of magnitude. Our approach enhances throughput by 2-3 orders of magnitude. This indicates that we reach a throughput in synchronous mode that is comparable to that in asynchronous mode, and it is over 10 times faster than the software solution, making the direct deployment of asynchronous architectures on mainstream DPI systems feasible. To the best of our knowledge, this is the first attempt to improve hardware-based regular expression matching under synchronous logic, achieving both high throughput and usability.
Shuhui Chen, Ziling Wei, Jincheng Zhong, Puguang Liu
IEEE Trans. Netw.4
2025 Domain Guidance: A Simple Transfer Approach for a Pre-trained Diffusion Model
abstract
Recent advancements in diffusion models have revolutionized generative modeling. However, the impressive and vivid outputs they produce often come at the cost of significant model scaling and increased computational demands. Consequently, building personalized diffusion models based on off-the-shelf models has emerged as an appealing alternative. In this paper, we introduce a novel perspective on conditional generation for transferring a pre-trained model. From this viewpoint, we propose *Domain Guidance*, a straightforward transfer approach that leverages pre-trained knowledge to guide the sampling process toward the target domain. Domain Guidance shares a formulation similar to advanced classifier-free guidance, facilitating better domain alignment and higher-quality generations. We provide both empirical and theoretical analyses of the mechanisms behind Domain Guidance. Our experimental results demonstrate its substantial effectiveness across various transfer benchmarks, achieving over a 19.6\% improvement in FID and a 23.4\% improvement in FD$_\text{DINOv2}$ compared to standard fine-tuning. Notably, existing fine-tuned models can seamlessly integrate Domain Guidance to leverage these benefits, without additional training. Code is available at this repository: https://github.com/thuml/DomainGuidance.
Jincheng Zhong, Xiangcheng Zhang, Jianmin Wang 0001, Mingsheng Long
ICLR1
2025 Memory-Efficient Packet Classification at High-Speed: The pRFC Architecture with Heuristic Partitioning
abstract
Packet classification is essential for modern networked systems, the rapid growth of rule sets and strategies in SDN and NFV environments demands higher performance and better memory-efficient solutions than ever. Existing RFC-based approaches, such as HybridRFC, suffer trade-offs between speed and memory usage. This paper presents pRFC, a partitioningenhanced recursive flow classification architecture that improves classification performance while significantly reducing memory consumption. By introducing a prefix-length-guided partitioning strategy and a lightweight compression mechanism, pRFC mitigates cross-product explosion and reduces bitwise processing overhead. Compared to uniform partitioning, it achieves up to 16.86% lower memory usage and 34.19% faster construction. Evaluations on ClassBench show that pRFC reduces memory usage by up to 80%, accelerates construction by up to 97%, and improves throughput by$4.0 \times$over standard RFC. Against HybridRFC, it achieves 72% lower memory consumption, 10% faster construction, and$2.45 \times$higher software throughput. An FPGA prototype demonstrates that pRFC fits entirely within on-chip memory and supports 100 Gbps line-rate classification via pipelining. These results highlight the effectiveness and practicality of pRFC for large-scale rule classification in resourceconstrained programmable networks.
Yuanfeng Chen, Xiangrui Yang 0002, Xuyan Jiang, Jincheng Zhong, Gaofeng Lv
IWQoS4
2024 Diffusion Tuning: Transferring Diffusion Models via Chain of Forgetting
abstract
Diffusion models have significantly advanced the field of generative modeling. However, training a diffusion model is computationally expensive, creating a pressing need to adapt off-the-shelf diffusion models for downstream generation tasks. Current fine-tuning methods focus on parameter-efficient transfer learning but overlook the fundamental transfer characteristics of diffusion models. In this paper, we investigate the transferability of diffusion models and observe a monotonous chain of forgetting trend of transferability along the reverse process. Based on this observation and novel theoretical insights, we present Diff-Tuning, a frustratingly simple transfer approach that leverages the chain of forgetting tendency. Diff-Tuning encourages the fine-tuned model to retain the pre-trained knowledge at the end of the denoising chain close to the generated data while discarding the other noise side. We conduct comprehensive experiments to evaluate Diff-Tuning, including the transfer of pre-trained Diffusion Transformer models to eight downstream generations and the adaptation of Stable Diffusion to five control conditions with ControlNet. Diff-Tuning achieves a 24.6% improvement over standard fine-tuning and enhances the convergence speed of ControlNet by 24%. Notably, parameter-efficient transfer learning techniques for diffusion models can also benefit from Diff-Tuning. Code is available at this repository: https://github.com/thuml/Diffusion-Tuning.
Jincheng Zhong, Xingzhuo Guo, Jiaxiang Dong, Mingsheng Long
NeurIPS1
2024 A Large-Scale Mobile Traffic Dataset For Mobile Application Identification
abstract
Abstract With Internet access shifting from desktop-driven to mobile-driven, application-level mobile traffic identification has become a research hotspot. Although considerable progress has been made in this research field, two obstacles are hindering its further development. Firstly, there is a lack of sharable labeled mobile traffic datasets. Although it is easy to capture mobile traffic, labeling traffic at the application level is non-trivial. Besides, researchers usually hold a conservative attitude toward publishing their datasets for privacy concerns. Secondly, most of the datasets used by existing studies are inadequate to evaluate the proposed methods, since they usually have the problems of inaccurate labels, small scale and simple collection configurations. To tackle these two obstacles, a mobile traffic collection is carried out in this paper. The collected traffic has the advantages of large-scale data size, accurate application-level labels and diverse collection configurations. Then, the collected traffic is anonymized carefully to make it public. Several mobile traffic identification methods are compared based on our anonymized dataset, which proves the applicability of our dataset.
Shuhui Chen, Fei Wang 0076, Ziling Wei, Jincheng Zhong, Jianbing Liang
Comput. J.5
2023 Bi-tuning: Efficient Transfer from Pre-trained Models
Jincheng Zhong, Ximei Wang, Zhi Kou, Mingsheng Long
ECML/PKDD (5)1
2023 FPGA-CPU Architecture Accelerated Regular Expression Matching With Fast Preprocessing
abstract
Abstract Regular Expression Matching (REM) is the core of Deep Packet Inspection (DPI), which is important for various network security applications. The burgeoning Software Defined Network and Network Function Virtualization technologies make the network evolve more dynamic, which brings serious challenges for DPI engines to achieve high matching performance with fast rule-set update capability. To meet these challenges, this paper proposes a heterogeneous Field Programmable Gate Array (FPGA)-Central Processing Unit (CPU) architecture to accelerate Deterministic Finite Automaton (DFA)-based REM with high preprocessing performance. Firstly, a novel regex decomposition technique is proposed to solve the DFA state explosion problem, which splits each regex into one prefix and several postfixes. Secondly, heterogeneous architecture is presented to collaboratively handle regex matching, in which prefixes are matched in parallel in an FPGA and postfixes are matched in a CPU. To further improve the matching performance, several well-designed DFA compression techniques and regex decomposition optimizations are proposed. Our design has been implemented in a DPI prototype employing a medium-end FPGA. Extensive experiments are conducted to evaluate the performance. Results reveal that our proposed architecture achieves 6.33 Gbps matching throughput on the Snort rule-set (v3.0), which is close to state-of-the-art FPGA NFA-based schemes. However, the rule-set preprocessing time is significantly reduced to <7 minutes, compared with up to several hours of FPGA NFA-based countermeasures.
Jincheng Zhong, Shuhui Chen, Biao Han 0003
Comput. J.1
2023 TupleTree: A High-Performance Packet Classification Algorithm Supporting Fast Rule-Set Updates
abstract
Packet classification plays a crucial role in various network functions such as access control and routing. In recent years, the rapid development of SDN and NFV poses new challenges for packet classification to support fast rule-set updates as introducing strong dynamics for the structure of networks. To this end, this paper proposes a novel scheme, TupleTree, to perform high-speed packet classification while providing fast rule-set update ability. TupleTree is a hybrid scheme combining decision tree and tuple space. In TupleTree, it organizes rules in a decision tree-like structure, but distributes rules in each node into child nodes through hashing rather than cutting or splitting. With the decision tree structure, for each classification, one leaf node containing a few rules can be rapidly indexed. Hence, a high classification performance can be achieved. Meanwhile, with hashing instead of cutting or splitting, it is easy to support fast rule-set updates due to having avoided the rule replication problem. Compared to state-of-the-art schemes that support fast rule-set updates, experimental results show that our proposed scheme achieves a classification performance improvement of 85% to 237% while retaining close update performance for large rule-sets.
Jincheng Zhong, Ziling Wei, Shuhui Chen
IEEE/ACM Trans. Netw.1
2022 FATSS: Filter-Assisted Tuple Space Search for Packet Classification
abstract
Packet Classification is a key part of supporting lots of network functions. Various algorithms have been proposed over the years to meet the increasing performance requirements of packet classification. Tuple space search (TSS) is one of the most popular algorithms and well-suited to scenarios requiring efficient online updates. However, the huge number of tuples in the algorithm leads to numerous memory accesses during packet classification, which limits the classification performance. This paper proposes a novel model named FATSS, which uses Filters to Assist the Tuple Space Search algorithm and reduces the number of tuple accesses. We first create the ImCuckoo Filter by improving the Cuckoo Filter from its structure, capacity and hash calculation. Then, we embed ImCuckoo Filter into TSS in two ways (online and offline) to adapt to diverse scenarios and requirements. By the experiments, it can be found that the ImCuckoo Filter can reduce more than 80% of tuple accesses. Furthermore, the access time of the filter is no more than 60% compared with that of the hash table. The experimental results show that the classification time of FATSS is 17%–19% faster than that of existing widely used algorithms.
Jiayao Wang 0002, Ziling Wei, Jincheng Zhong, Shuhui Chen
IPCCC4
2022 Robust Packet Classification with Field Missing
abstract
Packet classification shows a key role in kinds of network functions, such as access control, routing, and quality of service (QoS). With the rapid growth of the network size, users have to ignore some fields in packet classification due to resource constraints. In addition, some fields may not always be available in some networks. However, traditional packet classification algorithms can hardly handle packet classification if some fields are missing. In this paper, we propose a novel model to build a robust classifier. In the classifier, we utilize the advantage of Recursive Flow Classification (RFC) in handling fields concurrently. Then, we design a new workflow to deal with field missing based on flows. In addition, two complementary bitmap models are designed to accelerate matching packets to flows, and a buffer mechanism is introduced to further improve the classification accuracy. Our experiments show that the proposed classifier can classify packets with an accuracy of 94%-99.5% when the field missing probability is lower than 0.3.
Jiayao Wang 0002, Ziling Wei, Baokang Zhao, Jincheng Zhong
LCN5
2022 RTSS: Robust Tuple Space Search for Packet Classification
abstract
Packet classification shows an essential role in net-work functions. Traditional classification algorithms assume that all field values are available and valid. However, such a premise is being challenged as networks become more complex now. Scenarios with field-missing poses great challenges to packet classifiers. Existing approaches can only list all possible situations in such cases, increasing the workload exponentially. RFC algorithm is proved to be helpful for this issue in our previous work, but its spacial performance is much poor. In this paper, we propose a novel classification scheme using Tuple Space Search (TSS) to deal with missing fields. We redesign the hash calculation method and raise a new data structure to recover field-missing packets. The experiment shows that RTSS reduce the memory consumption and construction time by several orders of magnitude. At the same time, RTSS has better classification performance than previous work, while supporting fast updates.
Jiayao Wang 0002, Ziling Wei, Shuhui Chen, Jincheng Zhong
MSN5
2022 Comprehensive Mobile Traffic Characterization Based on a Large-Scale Mobile Traffic Dataset
Jincheng Zhong, Shuhui Chen, Jianbing Liang
NSS2
2021 CMT: An Efficient Algorithm for Scalable Packet Classification
abstract
Abstract Packet classification plays an essential role in diverse network functions such as quality of service, firewall filtering and load balancer. However, implementing an efficient packet classifier is a challenging problem. The problem even gets worse in the era of software-defined network, in which frequent rule updates are performed, and complex flow tables are used. This paper proposes CMT, a new software algorithm named by its novel data structure—common mask tree—to implement an efficient multi-field packet classifier. The core idea of CMT is to combine the strengths of both decision-tree and tuple-space schemes by employing tree-like structures and hash tables simultaneously. The objective of CMT is to achieve both high classification performance and fast rule updates. In the evaluation section, CMT is compared with decision-tree and tuple-space schemes. Compared to the state-of-the-art decision-tree methods, CMT performs rule updates at two orders of magnitude faster. CMT has a stable performance on different rulesets and achieves a 40% improvement in memory access compared to the state-of-the-art tuple-space method.
Shuhui Chen, Jincheng Zhong, Ziling Wei
Comput. J.2
2021 Efficient multi-category packet classification using TCAM
Jincheng Zhong, Shuhui Chen
Comput. Commun.1