Jialong Li 0006

dblp:205/9748-6 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-3416-5551ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FMD-AL: Cold-Start Active Learning based on Foundation Model
abstract
Active Learning (AL) is essential for mitigating annotation costs, yet traditional uncertainty-based strategies encounter a severe reliability crisis in cold-start scenarios due to model miscalibration and self-reinforcing bias. To address this, we propose FMD-AL, a framework that shifts the selection logic from susceptible task-model outputs to the robust zero-shot priors of Foundation Models (FMs). By leveraging a pre-trained FM as an inference proxy, FMD-AL quantifies intrinsic anatomical ambiguity through multi-prompt mask discrepancy. This strategy evaluates predictive inconsistency across diverse prompting settings including no-prompt, center-point, and random-point prompts to generate reliable selection signals independent of initial supervision. Consequently, FMD-AL effectively captures high-value samples with complex structures that are typically overlooked by conventional metrics. Extensive experiments on VerSe, ISIC 2017, and Kvasir-SEG datasets show that FMD-AL significantly outperforms eight state-of-the-art baselines, improving the Dice score by 10.01% over the best baseline and achieving an average gain of 16.33%, establishing a robust paradigm for sample selection in low-resource medical image annotation.
Huilin Ai, Ying Zhou 0017, Zhonghua Peng, Jialong Li 0006
ICMR8
2026 LDCS-Net: Local Dual-Context Attention Network for 3D Semantic Segmentation
abstract
High-precision 3D point cloud semantic segmentation in complex environments is often compromised by blurry boundaries and lost fine-grained details, primarily due to restricted receptive fields and geometric information decay. We present the Local Dual-Context and Multi-Scale Semantic Segmentation Network (LDCS-Net), a novel framework that preserves structural integrity through attention-guided graph convolution and dual-context fusion. By explicitly decoupling geometric and semantic cues, LDCS-Net strengthens local feature encoding and integrates dynamic graph construction with global attention to model long-range dependencies. Extensive experiments on the S3DIS dataset demonstrate that our method improves boundary delineation and small object recognition, achieving 65.9% mIoU and 72.2% mAcc. This represents a 1.2% boost in mIoU over the current state-of-the-art and a 3.2% increase in mAcc over the previous best-performing baseline. The source code is available at https://github.com/666-cute-yang/LDCS-NET.
Ying Zhou 0017, Meina Song, Huilin Ai, Jialong Li 0006, Zi Yang Chen
ICMR5
2026 Matryoshka: Realizing Hyperscale Data Center Network Design for the AI Era
Yan Cai 0018, Jialong Li 0006, Kutalmis Akpinar, Hany Morsy, Sunil Khaunte, Yiting Xia, Ying Zhang 0022
NSDI2
2026 SyncWise: Error-Aware Time Synchronization for Reconfigurable Data Center Networks
Yiming Lei 0002, Jialong Li 0006, Zhengqing Liu, Raj Joshi, Yiting Xia
NSDI2
2026 OpenOptics: Enabling Open Research and Implementation of Optical Data Center Networks
Yiming Lei 0002, Federico De Marchi 0002, Jialong Li 0006, Raj Joshi, Shu-Ting Wang, Balakrishnan Chandrasekaran 0002, Yiting Xia
NSDI3
2026 S3AD: Efficient industrial anomaly detection via selection, space, and scale
abstract
Real-world industrial visual anomaly detection (VAD) demands the simultaneous handling of diverse product categories under strict real-time constraints. However, the self-attention mechanism inherent in standard Vision Transformers (ViTs) incurs quadratic computational complexity during global context modeling, thereby limiting their real-time applicability in VAD. Crucially, this bottleneck is largely driven by the uniform processing of massive background redundancy common in industrial scenarios. To bypass this inefficiency while retaining high precision, we propose S 3 AD , a unified framework anchored in Selection, Space, and Scale. Instead of dense global interactions, S 3 AD employs a Salience-Guided Sparse Attention to focus resources solely on salient areas. Specifically, regarding Selection, we devise a discriminating Top-K pruning strategy guided by a hybrid metric of attention scores and feature magnitudes, which filters out non-informative tokens to minimize cost. Simultaneously, for Space, we incorporate Geometry-Preserving Positional Embeddings to explicitly anchor the selected tokens, ensuring structural integrity is maintained despite sparse interaction. Finally, concerning Scale, we integrate an Exchangeable Feature Pyramid Network (E-FPN) to ensure rigorous feature fusion across resolutions. Extensive experiments demonstrate that S 3 AD achieves state-of-the-art performance and establishes a superior accuracy–efficiency Pareto frontier, consistently surpassing strong baselines on multiple benchmarks. The code is available at: https://github.com/CreatedTRYNA/S3AD .
Zhonghua Peng, Chengyang Dong, Bo Liu 0034, Jialong Li 0006
Pattern Recognit.6
2026 Distributed Flow Control for Efficient DNN Training Scheduling
Chengze Du 0001, Ying Zhou 0017, Bo Liu 0034, Jialong Li 0006
IEEE Trans. Netw. Serv. Manag.7
2026 REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks
abstract
Community GPU(Graphics Processing Unit) platforms are emerging as a cost-effective and democratized alternative to centralized GPU clusters for AI(Artificial Intelligence) workloads, aggregating idle consumer GPUs from globally distributed and heterogeneous environments. However, their extreme hardware/software diversity, volatile availability, and variable network conditions render traditional schedulers ineffective, leading to suboptimal task completion. In this work, we present REACH (Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks), a Transformer-based reinforcement learning framework that redefines task scheduling as a sequence scoring problem to balance performance, reliability, cost, and network efficiency. By modeling both global GPU states and task requirements, REACH learns to adaptively co-locate computation with data, prioritize critical jobs, and mitigate the impact of unreliable resources. Extensive simulation results show that REACH improves task completion rates by up to 17%, more than doubles the success rate for high-priority tasks, and reduces bandwidth penalties by over 80% compared to state-of-the-art baselines. Stress tests further demonstrate its robustness to GPU churn and network congestion, while scalability experiments confirm its effectiveness in large-scale, high-contention scenarios.
Chengze Du 0001, Ying Zhou 0017, Bo Liu 0034, Jialong Li 0006
IEEE Trans. Netw. Serv. Manag.6
2026 Unlocking Diversity of Fast-Switched Optical Data Center Networks With Unified Routing
abstract
Optical data center networks (DCNs) are emerging as a promising solution for cloud infrastructure in the post-Moore’s Law era, particularly with the advent of “fast-switched” optical architectures capable of circuit reconfiguration at microsecond or even nanosecond scales. However, frequent reconfiguration of optical circuits introduces a unique challenge: in-flight packets risk loss during these transitions, hindering the deployment of many mature optical hardware designs due to the lack of suitable routing solutions. In this paper, we presentUnifiedRouting forOptical networks (URO), a general routing framework designed to support fast-switched optical DCNs across various hardware architectures. URO combines theoretical modeling of this novel routing problem with practical implementation on programmable switches, enabling precise, time-based packet transmission. Our prototype on Intel Tofino2 switches achieves a minimum circuit duration of$\mathrm {2~\mu \text {s} }$, ensuring end-to-end, loss-free application performance. Large-scale simulations using production DCN traffic validate URO’s generality across different hardware configurations, demonstrating its effectiveness and efficient system resource utilization.
Jialong Li 0006, Federico De Marchi 0002, Yiming Lei 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia
IEEE Trans. Netw.1
2026 Optimizing Mixture-of-Experts Inference Time via Model Deployment and Communication Scheduling
abstract
As machine learning models scale in size and complexity, their computational requirements become a significant barrier. Mixture-of-Experts (MoE) models alleviate this issue by selectively activating relevant experts. Despite this, MoE models are hindered by high communication overhead from all-to-all operations, low GPU utilization, and complications from heterogeneous GPU environments. This paper presents Comet, which optimizes both model deployment and all-to-all communication scheduling to address these challenges in MoE inference. Comet achieves minimal communication times by strategically ordering token transmissions in all-to-all communications. It improves GPU utilization by colocating experts from different models on the same device, avoiding the limitations of all-to-all communication. We analyze Comet’s optimization strategies theoretically across four common GPU cluster settings: exclusive vs. colocated models on GPUs, and homogeneous vs. heterogeneous GPUs. Comet provides optimal solutions for three cases, and for the remaining NP-hard scenario, it offers a polynomial-time sub-optimal solution with only a 1.09× degradation from the optimal, as shown in the simulation results. Comet is the first approach to minimize MoE inference time via optimal model deployment and communication scheduling across various scenarios. Evaluations demonstrate that Comet significantly accelerates inference, achieving speedups of up to 2.63× in homogeneous clusters and 2.91× in heterogeneous environments. Moreover, Comet enhances GPU utilization by up to 2.38× compared to existing methods.
Jialong Li 0006, Shreyansh Tripathi, Lakshay Rastogi, Yiming Lei 0002, Rui Pan 0003, Yiting Xia
IEEE Trans. Netw.1
2026 RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
abstract
Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to exploit the deterministic topology of Rail architectures, leaving multi-NIC bandwidth underutilized. We present RailS, a distributed load-balancing framework that minimizes all-to-all completion time in MoE training. RailS leverages the Rail topology’s symmetry to prove that uniform sending ensures uniform receiving, transforming global coordination into local scheduling. Each node independently executes a Longest Processing Time First (LPT) spraying scheduler to proactively balance traffic using local information. RailS activates N parallel rails for fine-grained, topology-aware multipath transmission. Across synthetic and real-world MoE workloads, RailS improves bus bandwidth by 20%–78% and reduces completion time by 17%–78%. For Mixtral workloads, it shortens iteration time by 18%–40% and achieves near-optimal load balance, fully exploiting architectural parallelism in distributed training.
Chengze Du 0001, Ying Zhou 0017, Weiqiang Cheng, Jialong Li 0006
IEEE Trans. Netw.8
2025 Generalized SRv6 Header Compression Packet Processors based on Multicore Architectures
Weiqiang Cheng, Xiaodong Duan, Han Li 0008, Jialong Li 0006
APNet5
2025 Unlocking Superior Performance in Reconfigurable Data Center Networks with Credit-Based Transport
abstract
The large-scale, end-to-end implementation of microsecond-switched reconfigurable data center networks (RDCNs), coupled with innovative routing and topology designs that provide continuous routes abstracting away frequent topology changes, demonstrates promise as a viable alternative to Clos networks in the post-Moore's Law era. However, the gap remains in transport performance, with current transport solutions falling short of unlocking their full performance. In this paper, we introduce Flare, a novel credit-based transport protocol that ensures reliable traffic delivery, low latency, and leverages the rapidly reconfiguring circuits of the RDCN to opportunistically route traffic over short paths, maximizing throughput. In simulations, Flare enables RDCNs to outperform Clos networks, achieving up to 1.15× higher throughput even under adversarial traffic. Additionally, it delivers up to 2× and 1.5× higher throughput than NDP and ExpressPass, and up to 10×, 15×, and 3.5× shorter flow completion time (FCT) than ExpressPass, TDTCP, and Bolt. Our testbed implementation further demonstrates the feasibility of Flare's mechanisms with programmable switches and DPDK.
Federico De Marchi 0002, Jialong Li 0006, Ying Zhang 0022, Wei Bai 0001, Yiting Xia
SIGCOMM2
2024 Uniform-Cost Multi-Path Routing for Reconfigurable Data Center Networks
abstract
Reconfigurable data center networks (RDCNs) are arising as a promising data center network (DCN) design in the post-Moore's law era. However, the constantly reconfigured network topology in RDCNs invalidates the assumption of using hop count as the cost metric for routing, e.g., the status quo Equal-Cost Multi-Path routing (ECMP) in traditional DCNs. Unfortunately, existing routing solutions in RDCNs stick to the old assumption and deliver suboptimal performance either high in latency or low in bandwidth efficiency. In this paper, we redefine the cost metric for RDCN routing with uniform cost to unify the effects of topology disruption and hop count on latency and bandwidth efficiency. We propose Uniform-Cost Multi-Path routing (UCMP), an ECMP equivalent for RDCNs, where minimizing uniform cost leads flows of various sizes to the right balance between latency and bandwidth efficiency. Our simulation shows that UCMP achieves 53% to 98% lower flow completion time (FCT) and 1.55× bandwidth efficiency compared to the state-of-the-art RDCN routing strategy, and our testbed implementation demonstrates sustainable switch resource usage of UCMP as RDCNs scale.
Jialong Li 0006, Haotian Gong, Federico De Marchi 0002, Aoyu Gong, Yiming Lei 0002, Wei Bai 0001, Yiting Xia
SIGCOMM1
2022 Hop-On Hop-Off Routing: A Fast Tour across the Optical Data Center Network for Latency-Sensitive Flows
abstract
Optical data center networks show promise to serve as the next-generation cloud infrastructure especially with their cost and power benefits. The need to set up dedicated optical circuits between endpoints before they can exchange data, however, delays latency-sensitive (“mice”) flows. We find the state-of-the-art solution to reducing flow latency produces sub-optimal paths. To address this issue, we leverage programmable switches to realize Hop-On Hop-Off (HOHO) routing, where mice flows are forwarded along the minimal-latency paths. We prove the optimality and robustness of our algorithm and sketch an implementation on programmable switches. In our packet-level simulations, HOHO routing reduces the flow-completion times for mice flows by up to 35% and the average path length by 15% compared to the state-of-the-art solution.
Jialong Li 0006, Yiming Lei 0002, Federico De Marchi 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia
APNet1
2022 Efficient flow scheduling in distributed deep learning training with echelon formation
abstract
This paper discusses why flow scheduling does not apply to distributed deep learning training and presents EchelonFlow, the first network abstraction to bridge the gap. EchelonFlow deviates from the common belief that semantically related flows should finish at the same time. We reached the key observation, after extensive workflow analysis of diverse training paradigms, that distributed training jobs observe strict computation patterns, which may consume data at different times. We devise a generic method to model the drastically different computation patterns across training paradigms, and formulate EchelonFlow to regulate flow finish times accordingly. Case studies of mainstream training paradigms under EchelonFlow demonstrate the expressiveness of the abstraction, and our system sketch suggests the feasibility of an EchelonFlow scheduling system.
Rui Pan 0003, Yiming Lei 0002, Jialong Li 0006, Binhang Yuan, Yiting Xia
HotNets3
2019 Provisioning Short-Term Traffic Fluctuations in Elastic Optical Networks
abstract
Transient traffic spikes are becoming a crucial challenge for network operators from both user-experience and network-maintenance perspectives. Different from long-term traffic growth, the bursty nature of short-term traffic fluctuations makes it difficult to be provisioned effectively. Luckily, next-generation elastic optical networks (EONs) provide an economical way to deal with such short-term traffic fluctuations. In this paper, we go beyond conventional network reconfiguration approaches by proposing the novel lightpath-splitting scheme in EONs. In lightpath splitting, we introduce the concept of SplitPoints to describe how lightpath splitting is performed. Lightpaths traversing multiple nodes in the optical layer can be split into shorter ones by SplitPoints to serve more traffic demands by raising signal modulation levels of lightpaths accordingly. We formulate the problem into a mathematical optimization model and linearize it into an integer linear program (ILP). We solve the optimization model on a small network instance and design scalable heuristic algorithms based on greedy and simulated annealing approaches. Numerical results show the tradeoff between throughput gain and negative impacts like traffic interruptions. Especially, by selecting SplitPoints wisely, operators can achieve almost twice as much throughput as conventional schemes without lightpath splitting.
Zhizhen Zhong, Nan Hua, Massimo Tornatore, Jialong Li 0006, Yanhe Li, Xiaoping Zheng, Biswanath Mukherjee
IEEE/ACM Trans. Netw.4