Zhikang Chen

dblp:271/5802 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
18since 2021 · last 2026
0009-0003-2568-6184ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Think Faster Than Words: Efficient LLM Chain-of-Thought Reasoning via Dynamic Shortcut Decoding
abstract
This paper proposes shortcut decoding, an efficient framework for accelerating Chain-of-Thought (CoT) reasoning in Large Language Models (LLMs).Existing methods that prune or employ early stopping to reduce latency often compromise reasoning reliability.Motivated by the observation that LLMs frequently converge to correct solutions internally before completing explicit textual reasoning, we propose a dual-signal adaptive controller that integrates lightweight probes over internal hidden states with step-level entropy.This controller detects convergence of reasoning during generation and adaptively selects between a fastexit path and a stability-verified path to remove redundant steps while preserving answer correctness.Experiments across multiple mathematical reasoning benchmarks demonstrate that shortcut decoding reduces token usage by approximately 35%, maintains accuracy comparable to full CoT decoding, and achieves finalanswer accuracy comparable to the full CoT baseline, outperforming existing early-stopping methods without updating the base model.Our code is available at https://github.com/ kuromi9527/shortcut_decoding.
Yanhao Wang 0001, Zhikang Chen, Lewei He, Jiahui Pan 0003
ACL (1)4
2026 P4XC: A Unified Compiler Framework for Network Dataplane with Heterogeneous Processors
Zhuang Ling, TianYing Tang, Haoyu Song 0001, Zhikang Chen, Bin Liu 0001
IWQoS4
2026 ALPS: ACK-induced Latency-based Packet Spraying for Multipath Transmission
Jinyu Xiao, Haoyu Song 0001, Zhikang Chen, Ying Wan 0001
IWQoS3
2026 Design of a configurable SoC for Alzheimer's disease detection based on multimodal signals
abstract
Alzheimer’s disease (AD) is an irreversible neurodegenerative disorder that remains difficult to cure. However, early screening and timely intervention can significantly slow its progression. Traditional AD detection methods are plagued by high misdiagnosis rates, low hardware integration, and lack of diagnostic diversity. To address these challenges, this paper proposes a configurable System-on-Chip (SoC) design based on a multimodal fusion Artificial Neural Network (ANN) for high-precision diagnosis. The proposed design integrates Electroencephalogram (EEG) and Magnetic Resonance Imaging (MRI) signals. First, a discretized reverse training method was employed to compress the features of the MRI images and reduce the input dimensionality. Second, intra-layer parallel computation and inter-layer pipeline scheduling were implemented to enhance the computational throughput. Finally, a dynamic configuration strategy for Processing Elements (PE) was introduced to optimize the hardware resource utilization. The proposed design achieves a six-fold improvement in throughput and provides multiple diagnostic approaches for AD. In conclusion, this work provides an efficient and scalable hardware solution for the early screening and dynamic monitoring of AD, which is expected to promote the development of portable and intelligent AD diagnostic devices and has good prospects for clinical transformation and application.
Yannan Yuan, Liufang Sheng, Zhikang Chen, Yuejun Zhang, Qikang Li, Junping Chen, Qiaoxia Hu, Wenming He
BMC Bioinform.3
2026 LAPS: Latency Aware Packet Spraying on Unequal-Cost Multi-Path Data Center Networks
Ying Wan 0001, Jinyu Xiao, Haoyu Song 0001, Zhikang Chen, Yunhui Yang, Bin Liu 0001, Tao Huang 0005
IEEE Trans. Netw.4
2025 Learning without Isolation: Pathway Protection for Continual Learning
abstract
Deep networks are prone to catastrophic forgetting during sequential task learning, i.e., losing the knowledge about old tasks upon learning new tasks. To this end, continual learning (CL) has emerged, whose existing methods focus mostly on regulating or protecting the parameters associated with the previous tasks. However, parameter protection is often impractical, since the size of parameters for storing the old-task knowledge increases linearly with the number of tasks, otherwise it is hard to preserve the parameters related to the old-task knowledge. In this work, we bring a dual opinion from neuroscience and physics to CL: in the whole networks, the pathways matter more than the parameters when concerning the knowledge acquired from the old tasks. Following this opinion, we propose a novel CL framework, learning without isolation (LwI), where model fusion is formulated as graph matching and the pathways occupied by the old tasks are protected without being isolated. Thanks to the sparsity of activation channels in a deep network, LwI can adaptively allocate available pathways for a new task, realizing pathway protection and addressing catastrophic forgetting in a parameter-effcient manner. Experiments on popular benchmark datasets demonstrate the superiority of the proposed LwI.
Zhikang Chen, Abudukelimu Wuerkaixi, Sen Cui, Haoxuan Li 0001, Jingfeng Zhang, Bo Han 0003, Gang Niu 0001, Houfang Liu, Yi Yang 0039, Sifan Yang, Changshui Zhang
ICML1
2025 Advancing Personalized Learning with Neural Collapse for Long-Tail Challenge
abstract
Personalized learning, especially data-based methods, has garnered widespread attention in recent years, aiming to meet individual student needs. However, many works rely on the implicit assumption that benchmarks are high-quality and well-annotated, which limits their practical applicability. In real-world scenarios, these benchmarks often exhibit long-tail distributions, significantly impacting model performance. To address this challenge, we propose a novel method called Neural-Collapse-Advanced personalized Learning (NCAL), designed to learn features that conform to the same simplex equiangular tight frame (ETF) structure. NCAL introduces Text-modality Collapse (TC) regularization to optimize the distribution of text embeddings within the large language model (LLM) representation space. Notably, NCAL is model-agnostic, making it compatible with various architectures and approaches, thereby ensuring broad applicability. Extensive experiments demonstrate that NCAL effectively enhances existing works, achieving new state-of-the-art performance. Additionally, NCAL mitigates class imbalance, significantly improving the model’s generalization ability.
Hanglei Hu, Zhikang Chen, Sen Cui, Fei Wu 0001, Kun Kuang 0001, Min Zhang 0068, Bo Jiang 0016
ICML3
2025 Decentralized Dynamic Cooperation of Personalized Models for Federated Continual Learning
abstract
Federated continual learning (FCL) has garnered increasing attention for its ability to support distributed computation in environments with evolving data distributions. However, the emergence of new tasks introduces both temporal and cross-client shifts, making catastrophic forgetting a critical challenge. Most existing works aggregate knowledge from clients into a global model, which may not enhance client performance since irrelevant knowledge could introduce interference, especially in heterogeneous scenarios. Additionally, directly applying decentralized approaches to FCL suffers from ineffective group formation caused by task changes. To address these challenges, we propose a decentralized dynamic cooperation framework for FCL, where clients establish dynamic cooperative learning coalitions to balance the acquisition of new knowledge and the retention of prior learning, thereby obtaining personalized models. To maximize model performance, each client engages in selective cooperation, dynamically allying with others who offer meaningful performance gains. This results in non-overlapping, variable coalitions at each stage of the task. Moreover, we use coalitional affinity game to simulate coalition relationships between clients. By assessing both client gradient coherence and model similarity, we quantify the client benefits derived from cooperation. We also propose a merge-blocking algorithm and a dynamic cooperative evolution algorithm to achieve cooperative and dynamic equilibrium. Comprehensive experiments demonstrate the superiority of our method compared to various baselines. Code is available at: https://github.com/ydn3229/DCFCL.
Danni Yang, Zhikang Chen, Sen Cui, Mengyue Yang, Abudukelimu Wuerkaixi, Haoxuan Li 0001, Jinke Ren, Mingming Gong
NeurIPS2
2025 ClubHeap: A High-Speed and Scalable Priority Queue for Programmable Packet Scheduling
Zhikang Chen, Haoyu Song 0001, Zhiyu Zhang 0012, Yang Xu 0010, Bin Liu 0001
NSDI1
2025 Enhancing Stateful Processing in Programmable Data Planes: Model and Improved Architecture
abstract
Stateful data plane network applications are indispensable but their efficiency is hindered by prevalent network device architectures which utilize a Blocking Scheme to maintain state consistency. The Blocking Scheme results in poor throughput and latency performance due to its frequent halts during packet processing. In response to this issue, we propose an innovative Non-Blocking Scheme and construct a theoretical model based on the G/GI/m queueing model. The new scheme leverages the speculative execution method, avoiding unnecessary blocking by taking advantage of the fact that the state update ratio is usually much smaller than the incoming packet rate. In case of speculation failures, the affected packets are reprocessed to guarantee state consistency. We provide an approximate model for the Blocking Scheme, and show that, even with relaxed approximations, the Blocking Scheme still performs worse than the Non-Blocking Scheme. The superior performance of the Non-Blocking Scheme is further corroborated through rigorous simulations, conducted under both realistic and synthetic traces. Based on our model, we propose an enhanced architecture: SN_RAPID (Sequence Number and unidirectional Reverse path-Augmented PIpeline Dataplane), to support speculative execution. The architecture is simpler than the other architectures supporting stateful network functions. Serving as a design foundation for future iterations, we implement a prototype of SN_RAPID in FPGA which can run at line speed, and also develop a software ASIC emulator. The experiments show the superiority of the improved architecture.
Hanyi Zhou, Zhikang Chen, Haoyu Song 0001, Bin Liu 0001
IEEE Trans. Netw.4
2024 Neural Collapse Inspired Feature Alignment for Out-of-Distribution Generalization
abstract
The spurious correlation between the background features of the image and its label arises due to that the samples labeled with the same class in the training set often co-occurs with a specific background, which will cause the encoder to extract non-semantic features for classification, resulting in poor out-of-distribution generalization performance. Although many studies have been proposed to address this challenge, the semantic and spurious features are still difficult to accurately decouple from the original image and fail to achieve high performance with deep learning models. This paper proposes a novel perspective inspired by neural collapse to solve the spurious correlation problem through the alternate execution of environment partitioning and learning semantic masks. Specifically, we propose to assign an environment to each sample by learning a local model for each environment and using maximum likelihood probability. At the same time, we require that the learned semantic mask neurally collapses to the same simplex equiangular tight frame (ETF) in each environment after being applied to the original input. We conduct extensive experiments on four datasets, and the results demonstrate that our method significantly improves out-of-distribution performance.
Zhikang Chen, Min Zhang 0068, Sen Cui, Haoxuan Li 0001, Gang Niu 0001, Mingming Gong, Changshui Zhang, Kun Zhang 0001
NeurIPS1
2024 Empower Programmable Pipeline for Advanced Stateful Packet Processing
Zhikang Chen, Haoyu Song 0001, Yinchao Zhang, Hanyi Zhou, Ruoyu Sun 0009, Wenkuo Dong, Chuwen Zhang, Yang Xu 0010, Bin Liu 0001
NSDI2
2024 OptimusPrime: Unleash Dataplane Programmability through a Transformable Architecture
abstract
Network dataplane calls for better programmability. Current programmable network processing chips are based on either pipeline or multi-core Run-To-Completion (RTC) architecture with various trade-offs in flexibility, performance, and cost. The existing attempts to amalgamate the strengths of the two are stilted and inflexible. In this paper, we challenge the status quo by introducing a more fluid and organic programmable chip architecture, OptimusPrime, built from identical hardware blocks. Unlike the conventional static hybrid architecture, OptimusPrime allows each block to be transformed into either a pipeline stage processor or a multi-core RTC processor through software-defined configuration, enabling versatile data plane programming tailored to a wide range of applications (e.g., stateful packet processing and in-network computing). We integrate the C and P4 languages for application programming and develop algorithms to map a user program to the optimal distribution of pipeline stages and RTC cores. We demonstrate the viability of OptimusPrime through practical use cases such as in-network aggregation, in-network caching, and network function integration. We developed an FPGA-based prototype and a software-based ASIC simulator to validate the feasibility of OptimusPrime, which can be used by switches and smartNICs to enhance their programmability to a new level with high performance and low cost.
Zhikang Chen, Haoyu Song 0001, Hanyi Zhou, Tong Yun, Wenquan Xu, Tian Pan 0001, Bin Liu 0001
SIGCOMM1
2023 FlowBench: A Flexible Flow Table Benchmark for Comprehensive Algorithm Evaluation
abstract
Flow table is a fundamental and critical component in network data plane. Numerous algorithms and architectures have been devised for efficient flow table construction, lookup, and update. The diversity of flow tables and the difficulty to acquire real data sets make it challenging to give a fair and confident evaluation to a design. In the past, researchers rely on ClassBench and its improvements to synthesize flow tables, which become inadequate for today’s networks. In this paper, we present a new flow table benchmark tool, FlowBench. Based on a novel design methodology, FlowBench can generate large-scale flow tables with arbitrary combination of matching types and fields in a short time, and yet keep accurate characteristics to reveal the real performance of the algorithms under evaluation. The open-source tool facilitates researchers to evaluate both existing and future algorithms with unprecedented flexibility.
Zhikang Chen, Ying Wan 0001, Ting Zhang 0010, Haoyu Song 0001, Bin Liu 0001
INFOCOM1
2023 ClickINC: In-network Computing as a Service in Heterogeneous Programmable Data-center Networks
abstract
In-Network Computing (INC) has found many applications for performance boosts or cost reduction. However, given heterogeneous devices, diverse applications, and multi-path network typologies, it is cumbersome and error-prone for application developers to effectively utilize the available network resources and gain predictable benefits without impeding normal network functions. Previous work is oriented to network operators more than application developers. We develop ClickINC to streamline the INC programming and deployment using a unified and automated workflow. Click-INC provides INC developers a modular programming abstractions, without concerning to the states of the devices and the network topology. We describe the ClickINC framework, model, language, workflow, and corresponding algorithms. Experiments on both an emulator and a prototype system demonstrate its feasibility and benefits.
Wenquan Xu, Haoyu Song 0001, Zhikang Chen, Wenfei Wu, Guyue Liu, Yinchao Zhang, Zerui Tian, Bin Liu 0001
SIGCOMM5
2022 Enabling In-situ Programmability in Network Data Plane: From Architecture to Language
Zhikang Chen, Haoyu Song 0001, Wenquan Xu, Tong Yun, Bin Liu 0001
NSDI2
2021 In-situ Programmable Switching using rP4: Towards Runtime Data Plane Programmability
abstract
The existing chip architecture and programming language are incapable of supporting in-service updates by loading or offloading on-demand protocols and functions at runtime. We examine the fundamental reasons for the inflexibility and design a new In-situ Programmable Switch Architecture (IPSA) as a fix. We further design rP4, a P4 extension, for programming IPSA-based devices. To manifest the in-situ programming feasibility, we develop an rP4 compiler and demonstrate several use cases on both a software switch, ipbm, and an FPGA-based prototype. Our preliminary experiments and analysis show that, compared to PISA, IPSA provides higher flexibility in enabling runtime functional update with limited performance and gate-count penalty. The in-situ programming capability enabled by IPSA and rP4 opens a promising design space for programmable networks.
Haoyu Song 0001, Zhikang Chen, Wenquan Xu, Bin Liu 0001
HotNets4
2021 PIPO: Efficient Programmable Scheduling for Time Sensitive Networking
abstract
Time Sensitive Networking (TSN) is an emerging Ethernet technology for real-time systems. To address different Quality-of-Service (QoS) requirements of applications, IEEE 802.1 TSN Task Group has standardized several packet scheduling and shaping algorithms. The software implementation of these algorithms is hard to meet the performance requirements, while the hardware implementation in Application-Specific Integrated Circuit (ASIC) is inflexible. A hardware-programmable scheduler is necessary to deal with this dilemma. Among the existing primitives, the most expressive one is Push-In-Extract-Out (PIEO), but its complexity makes the implementation very expensive. A relatively lower-cost implementation of PIEO cannot guarantee the scheduling correctness for the most critical Time-Triggered (TT) traffic in TSN. As a remedy, in this paper we propose a new Push-In-Pick-Out (PIPO) primitive under a TSN programmable scheduling framework. Composed of simple priority queues, PIPO can express all existing TSN scheduling and shaping algorithms, and is flexible enough to support future ones. Our PIPO implementation guarantees the TT traffic scheduling correctness. The simulation results corroborate the theoretical analysis that the low-cost PIPO can closely approximate PIEO and sustain a high bandwidth utilization. The prototype on Xilinx FPGA shows that, with 2,048 inputs, the PIPO-based scheduler achieves a throughput of 70 Mpps, which is 1.64x higher than the PIEO-based one, but using only 14.7% Look-Up Tables (LUTs) and 40.5% Block RAMs of the latter.
Chuwen Zhang, Zhikang Chen, Haoyu Song 0001, Ruyi Yao, Yang Xu 0010, Yi Wang 0004, Ji Miao, Bin Liu 0001
ICNP2
2020 GlobalInsight: An LSTM Based Model for Multi-Vehicle Trajectory Prediction
abstract
Intelligent Transport System (ITS) raises the increasing demand on accurate vehicle trajectory prediction for navigation efficiency. The rapidly developing 5G networks provides communications with high transmission bandwidth and super-low latency, paving the way for Mobile Edge Computing (MEC) to calculate more accurate trajectory prediction for vehicles, as the MEC server holds more comprehensive vehicular information. However, the current methods for trajectory prediction are not efficient due to the dynamical environment. To address this issue, we propose GlobalInsight, a Long Short-Term Memory (LSTM) based model, which runs on the MEC to perform accurate trajectory prediction for multiple vehicles no matter how scenario changes. In particular, we use three auxiliary layers to respectively capture the principal component of vehicle features, social interaction of adjacent vehicles, and the cross-vehicle correlation of similar vehicles. We further integrate the above information into LSTM in the main layer to enhance the trajectory learning and prediction. We evaluate our model under the NGSIM dataset, and experimental results exhibit that our model outperforms the state-of-the-art approaches.
Wenquan Xu, Zhikang Chen, Chuwen Zhang, Xuefeng Ji, Yunsheng Wang 0001, Bin Liu 0001
ICC2
2020 FastUp: Compute a Better TCAM Update Scheme in Less Time for SDN Switches
abstract
While widely used for flow tables in SDN switches, TCAM faces challenges for rule updates. Both the computation time and interrupt time need to be short. We propose FastUp, a new TCAM update algorithm, which improves the previous dynamic programming-based algorithms. Evaluations show that FastUp shortens the computation time by 40~100× and the interrupt time by 1.2~2.5×. In addition, we are the first to prove the NP-hardness of the optimal TCAM update problem, and provide a practical method to evaluate an algorithm's degree of optimality. Experiments show that FastUp's optimality reaches 90%.
Ying Wan 0001, Haoyu Song 0001, Hao Che, Yang Xu 0010, Yi Wang 0004, Chuwen Zhang, Zhijun Wang 0001, Tian Pan 0001, Hao Li 0011, Hong Jiang 0001, Chengchen Hu, Zhikang Chen, Bin Liu 0001
ICDCS12