EDBT 2026 Demo / reviewers in the wild / expert
Junye Zhang
dblp:310/0695
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Link-level Lossless Mechanisms in AI NetworksabstractThe ever-increasing demand for network performance in large language models promotes the advent of many Scale-up networking schemes. To consistently deliver superior low latency and high bandwidth, these schemes have widely adopted Credit-based Flow Control (CBFC) and Link Layer Retry (LLR) to ensure link-level lossless transmission. However, there is currently a lack of evaluations for these mechanisms in Scale-up domains. This paper builds an FPGA-based Scale-up network to evaluate these lossless mechanisms. Evaluation results show that, while these mechanisms consume a small amount of link bandwidth, CBFC can greatly reduce receive buffer utilization, and LLR can substantially mitigate network performance degradation caused by packet corruption. Kefei Liu, Ruixue Wang, Tianrun Jiang, Runlong Hu, Weiqiang Cheng, Danyuan Zhou, Junye Zhang, Xinhong Deng, Rundi Zhai, Cong Qi, Yijian Qi |
SIGCOMM | 10 |
| 2026 | Delphinus: Ultra-Fast Link Failure Detection and Recovery for AI Data Center NetworksabstractHigh-performance artificial intelligence (AI) applications impose stringent reliability requirements on AI data center networks (DCNs), yet link failures are almost inevitable and can severely disrupt AI workloads such as large language model (LLM) training and inference. Existing deployed link failure detection and recovery mechanisms suffer from slow execution speed and limited failure coverage, failing to meet the demands of production AI DCNs. To address these issues, we propose Delphinus, an ultra-fast link failure detection and recovery solution built on the data-plane of programmable switches. It achieves ultra-fast failure detection via hardware-based port state monitoring, extends recoverable failure coverage through remote failure notification and relay, and enables fast recovery by path switchover. Delphinus can serve as a key generic function of switches, providing host-transparent link failure handling for Ethernet fabrics. We implement Delphinus on commercial hardware switches, and deploy it in large-scale production AI DCNs for over a year. Extensive evaluations demonstrate that Delphinus can complete link failure detection and recovery within sub-milliseconds, with negligible impact on application performance and imperceptible service interruption. Junye Zhang, Zhigang Ji, Kefei Liu, Rui Zhuang, Ruixue Wang, Weiqiang Cheng, Zixuan Guan |
SIGCOMM | 1 |
| 2026 | Dragonfly-Ultra: A Scalable, Low-Cost Network Architecture for High-Performance AI ClustersabstractLarge-scale AI clusters impose higher requirements on network scalability, cost, and communication efficiency. The traditional Clos topology suffers from superlinear cost growth when scaling to over 100k GPUs, while the more cost-effective Dragonfly+ introduces "down-up" detours, deadlock risks, and complex routing design. This paper presents Dragonfly-Ultra, a scalable, low-cost network architecture for high-performance AI clusters. Dragonfly-Ultra can scale to over 260k GPUs with only 82% cost and 81% power consumption of a 3-layer Clos architecture. Dragonfly-Ultra optimizes inter-group connectivity to eliminate intra-group detours entirely. Beyond the topological benefits, Dragonfly-Ultra incorporates three key mechanisms to further improve network performance and optimize collective communication, including lightweight dual-waterline adaptive routing for fast congestion mitigation, virtual-link-based deadlock avoidance with lower hardware overhead, and uniform affinity-aware rank placement for balanced inter-group traffic across all phases. Simulation results on a 4k-node cluster show that, compared to Clos, Dragonfly-Ultra achieves up to 18.8% and 39.2% lower completion time for AllReduce and AlltoAll, respectively. Compared to Dragonfly+, the reductions are up to 27.9% and 62.1%, outperforming current mainstream topologies. Rui Zhuang, Junye Zhang, Kefei Liu, Weiqiang Cheng, Zixuan Guan, Shengnan Yue, Ruixue Wang, Tong Yang 0003 |
SIGCOMM | 3 |
| 2025 | Adaptive Secure Control for Uncertain Cyber-Physical Systems With Markov Switching Against Both Sensor and Actuator AttacksabstractIn this article, adaptive secure controller synthesis for uncertain cyber-physical systems with Markov switching (CPSMSs), both sensor and actuator stealthy attacks as well as generally unknown transition rates (GUTRs), is under consideration via neural sliding mode control (SMC) technique. In order to resist unknown attack signals from both sensor and actuator channels, a novel neural network (NN)-based SMC design is performed, which could not only guarantee the boundedness of relevant adaptive data but also force the actual state trajectories to arrive at the proposed sliding mode surface (SMS) with limited moments almost surely. Then, a fresh stochastically stable criterion for the resultant plant is provided in spite of hidden cyber attacks, GUTRs, and structural uncertainty, relying on the arrival of the SMS and stochastic stability theory. Finally, an F-404 aircraft engine model with performance comparisons is offered to confirm the feasibleness of the theoretical result. Zhen Liu 0024, Junye Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | A Knowledge-driven Self-healing Dual-loop and Validation for Autonomous NetworksabstractIntelligent technology is driving the communication industry to a higher stage of autonomy. To enable the large-scale deployment of Autonomous Networks (ANs), We design a knowledge-driven self-healing dual-loop architecture throughout the fault lifecycle, Then we propose an optimization model that aims to minimize the loss of services operation. Simulations are conducted in a programmable network to validate the loop. Can Tan, Honglin Fang, Junye Zhang, Dahua Lin, Peng Yu 0001 |
APNet | 3 |
| 2024 | Revisiting the Underlying Causes of RDMA Scalability IssuesabstractRemote direct memory access (RDMA) networks are widely deployed in clouds and data centers for low latency and high throughput. Emerging applications like artificial intelligence training and serverless computing demand RDMA networks with high connection scalability. However, increasing connections can reduce throughput, as many academic research states. Conversely, some industry reports indicate the scalability issues are situation-specific. Despite this, existing work lacks a comprehensive analysis of RDMA scalability issue triggers and causes. In this paper, we revisit triggering conditions and underlying causes of RDMA scalability issues. First, we comprehensively analyze RDMA data flows and potential resource contentions, particularly with direct cache access. Second, we conduct extensive tests across varied measurement settings and RDMA NICs (RNICs). We identify triggering conditions including frequent switching of queue pair connections, rapid request posting, intensive memory access demands and inappropriate program configurations. We also systematically uncover causes of scalability issues, mainly due to RNIC cache misses, RNIC processing unit backpressure, host last-level cache misses, and software inefficiency. Finally, we provide guidelines to mitigate the RDMA scalability issues. Junye Zhang, Peng Yu 0001, Hexiang Song, Di Qu |
ISPA | 1 |
| 2024 | Adaptive and low-cost resource synchronization based on data distribution service in high dynamic networks
Peng Yu 0001, Junye Zhang, Wenjing Li 0001 |
Comput. Networks | 3 |
| 2023 | Joint Routing and GCL Scheduling Algorithm Based on Tabu Search in TSNabstractTime sensitive networking (TSN) has been widely adopted and applied in many fields. The scheduling problem of TSN requires that the gate control list (GCL) is calculated according to the flow information in a given topology network. Conventional flow scheduling schemes are usually based on the given routing scheme, which limits the scheduling performance. Besides, current works mostly focus on the time trigger flows (TT). However, AVB flows exist as aperiodic flows in the industrial Internet. The integrated scheduling of these two types of flows is required to improve the overall schedulability. In this paper, a problem model of joint routing and GCL scheduling is proposed. An algorithm based on Tabu search (Tabu-RG) is proposed to solve the problem with specific design of neighborhood movement policy, neighborhood selection policy, as well as diversified function. Experimental results show that compared with the solver method, the proposed algorithm can save 75% of the time cost on the premise of ensuring the solution performance. Ying Wang 0002, Yufan Cheng, Zhihan Zhuang, Junye Zhang, Peng Yu 0001, Shao-Yong Guo 0001, Xuesong Qiu 0001 |
CNSM | 4 |
| 2023 | EVlncRNA-Dpred: improved prediction of experimentally validated lncRNAs by deep learningabstractLong non-coding RNAs (lncRNAs) played essential roles in nearly every biological process and disease. Many algorithms were developed to distinguish lncRNAs from mRNAs in transcriptomic data and facilitated discoveries of more than 600 000 of lncRNAs. However, only a tiny fraction (<1%) of lncRNA transcripts (~4000) were further validated by low-throughput experiments (EVlncRNAs). Given the cost and labor-intensive nature of experimental validations, it is necessary to develop computational tools to prioritize those potentially functional lncRNAs because many lncRNAs from high-throughput sequencing (HTlncRNAs) could be resulted from transcriptional noises. Here, we employed deep learning algorithms to separate EVlncRNAs from HTlncRNAs and mRNAs. For overcoming the challenge of small datasets, we employed a three-layer deep-learning neural network (DNN) with a K-mer feature as the input and a small convolutional neural network (CNN) with one-hot encoding as the input. Three separate models were trained for human (h), mouse (m) and plant (p), respectively. The final concatenated models (EVlncRNA-Dpred (h), EVlncRNA-Dpred (m) and EVlncRNA-Dpred (p)) provided substantial improvement over a previous model based on support-vector-machines (EVlncRNA-pred). For example, EVlncRNA-Dpred (h) achieved 0.896 for the area under receiver-operating characteristic curve, compared with 0.582 given by sequence-based EVlncRNA-pred model. The models developed here should be useful for screening lncRNA transcripts for experimental validations. EVlncRNA-Dpred is available as a web server at https://www.sdklab-biophysics-dzu.net/EVlncRNA-Dpred/index.html, and the data and source code can be freely available along with the web server. Bailing Zhou, Maolin Ding, Baohua Ji, Pingping Huang, Junye Zhang, Zanxia Cao, Yuedong Yang, Yaoqi Zhou, Jihua Wang |
Briefings Bioinform. | 6 |
| 2023 | Digital Twin Driven Service Self-Healing With Graph Neural Networks in 6G Edge Networksabstract6G edge networks strive to offer ubiquitous intelligent services, requiring a greater emphasis on network stability and reliability. However, current networks present a low automation degree of the operation, administration and maintenance process. Consequently, active service migration away from abnormal network nodes and links, as well as automatic and transparent service recovery from sudden anomalies, become challenging tasks. These conditions underscore the urgency for an innovative service self-healing mechanism for 6G edge networks. Digital twin (DT) technology uses modeling to represent physical entities, thereby facilitating lifecycle management. However, the application of DT technology in networks is still a burgeoning field of study. In this paper, we explore the DT-driven service self-healing mechanism in 6G edge networks. Initially, we design a DT-based architecture for service self-healing. Subsequently, we construct a performance prediction mechanism leveraging graph neural networks (GNNs) to devise an efficient prediction model, which aims to accurately infer network performance and promptly detect abnormal network conditions. To maintain fine-grained service stability amidst potential network anomalies, we propose a DT-driven service redeployment mechanism enhanced by GNNs. Comprehensive experimental results reveal that our proposed mechanism can accurately predict flow-level delays and identify abnormal links and nodes. Furthermore, the DT-driven service redeployment mechanism effectively reduces service delay and enhances network load balance. Peng Yu 0001, Junye Zhang, Honglin Fang, Wenjing Li 0001, Lei Feng 0001, Fanqin Zhou, Pei Xiao 0001, Song Guo 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2023 | Energy-Efficient Coverage and Capacity Enhancement With Intelligent UAV-BSs Deployment in 6G Edge NetworksabstractWith the development of 5G/6G networks, the number of wireless users is growing exponentially, and the application scenarios are increasingly diversified. Using unmanned aerial vehicles as base stations (UAV-BSs) to serve ground users has become a trend for wide area coverage and capacity enhancement for rapid access of service in 6G networks. However, as UAV-BSs have limited energy or battery storage, solutions to optimize energy efficiency while providing high-quality services are necessary. Therefore, this paper mainly concentrates on the energy-efficient deployment of coverage-aimed UAV-BSs (Co-UAV-BSs) and capacity-aimed UAV-BSs (Ca-UAV-BSs) for the coverage and capacity enhancement of ground communication under disaster areas or burst data traffic. First, Co-UAV-BSs are deployed with DQN algorithm to to get the UAV-BSs’ optimal flight paths, which mainly adopted to detect out of service users in such areas. Then the users are completely clustered based on the detection results. After that, Co-UAV-BSs and Ca-UAV-BSs are deployed hierarchically based on the user distribution and sought to optimize the energy efficiency with acceptable user services. Still, DQN algorithm and the A3C algorithm are used for obtaining all the UAV-BSs’ location deployment and users’ best connections. The simulation results show that the dynamic flying path requires less energy than the fixed path for user detecting. For the coverage and capacity enhancement, it reveals the solution we proposed could provide high-quality service for users with high energy efficiency comparing to traditional algorithms. Peng Yu 0001, Yahui Ding, Zifan Li, Jingyue Tian, Junye Zhang, Wenjing Li 0001, Xuesong Qiu 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Fine-Grained Service Offloading in B5G/6G Collaborative Edge Computing Based on Graph Neural NetworksabstractFine-grained service offloading in collaborative edge computing can make full use of the limited resource of edge nodes to achieve efficient parallel computing. It is imperative to select appropriate edge nodes for the subtask offloading in order to ensure the network’s load balance. However, there is a lack of research on computing offloading of end-to-end fine-grained services, and existing node selection algorithms can only be used in small-scale scenarios or networks with a fixed number of nodes. In this paper, we construct an end-to-end fine-grained computing offloading model, with load balancing as the optimization goal. Especially, a deep graph matching method, based on graph neural networks, is used for offloading node selection. It can be applied to dynamic and large-scale scenarios with strong generalization capability and fast execution speed. Compared with baseline algorithms, it greatly reduces the network load imbalance degree while ensuring a high acceptance ratio of services and meeting delay, location and resource constraints. Junye Zhang, Peng Yu 0001, Lei Feng 0001, Wenjing Li 0001, Xueqiang Yan, Jianjun Wu 0002 |
ICC | 1 |
| 2022 | Resource and delay aware fine-grained service offloading in collaborative edge computing
Junye Zhang, Peng Yu 0001, Fanqin Zhou, Lei Feng 0001, Wenjing Li 0001, Xuesong Qiu 0001 |
Comput. Networks | 1 |