Jingrong Wang

dblp:209/1568 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 4 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ProLet: Proactive Multi-path Load Balancing for Lossless RDMA
abstract
To achieve high-throughput and low-latency Remote Direct Memory Access (RDMA) communication in data center networks, load balancing is critical for preventing congestion and ensuring that traffic is efficiently distributed across available network paths. However, existing schemes may not effectively detect rerouting opportunities in continuous RDMA packet streams and may degrade in-order delivery, limiting their applicability to RDMA traffic. To address these limitations, we propose ProLet, a load balancing scheme that enables proactive probing and reroutes elephant flows at flowlet granularity in lossless RDMA networks. ProLet dynamically fine-tunes per-destination top-of-rack timeouts and enables effective in-network flowlet identification based on real-time network conditions. Meanwhile, it leverages lightweight mice flows as proactive probes to maintain network-wide congestion awareness. This allows ProLet to reroute elephant flows before congestion accumulates, mitigating the persistent queue buildup inherent in subflow-based schemes. Extensive numerical evaluations demonstrate that ProLet reduces average and tail flow completion time slowdowns by 69% and 79%, respectively, compared to state-of-the-art load balancing schemes.
Jinhao Luo, Jing Jie Tan, Jingrong Wang, Kaiyang Liu
APNet4
2026 Social Navigation as Gameplay: A Hybrid Social Physics Engine for Deduction Games
abstract
When LLMs directly drive NPC dialogue in social deduction games, a single information leak collapses the entire puzzle structure, a consequence far more severe than in casual social simulations. This paper presents a three-layer hybrid architecture that relocates LLM involvement to a narrow intent classification boundary, while delegating social state evolution to the Ensemble social physics engine and natural language output to a RiveScript dialogue module. A turn-taking mechanism integrating adjacency pair rules and probabilistic volition-weighted bidding coordinates multi-NPC discussion without central LLM mediation. Evaluation through dialogue breakdown analysis across 20 sessions (598 turns) reveals that repetition accounts for 60% of breakdowns, traceable to a structural bottleneck where the scripting layer collapses the engine’s fine-grained action variants into shallow response pools. An adapted Bi-Fact evaluation of the NLU layer confirms that the narrow-channel design achieves stable factual accuracy at the cost of systematic semantic compression. Playtesting further surfaces an emergent genre shift: players reorient from evidence-based deduction toward social relationship management, suggesting that social physics engines may constitute a generative substrate for a distinct interactive genre.
Jingrong Wang, Ziyuan Jiang
FDG1
2026 Multi-scale feature extraction of road marking from front-view images through hybrid convolutional and vision transformer for high-definition map
abstract
In the process of constructing high-definition maps, road markings play a crucial role, making the precise extraction of these markings an indispensable task for building high-definition maps. Current road marking extraction methods face certain challenges, including inadequate precision in extracting markings at different scales and difficulty in effectively mitigating background interference in the extraction process. This study proposes a novel road marking extraction method called RoadFormer, aiming to advance the construction of high-definition maps. RoadFormer seamlessly integrates the characteristics of Convolutional Neural Networks and Transformers to comprehensively extract road markings at various scales, avoiding the loss of local features and ensuring the continuity of extraction results. To address sample imbalance issues caused by backgrounds, we also design multi-scale data augmentation and a sample imbalance loss function. Experimental results demonstrate that both multi-scale data augmentation and the sample imbalance loss function contribute to improving road marking extraction to a certain extent. Furthermore, a series of comparative experiments between RoadFormer and six other methods were conducted on the CamVid dataset and a self-collected road marking dataset. The results demonstrate that RoadFormer achieves superior performance, attaining a Road Marking F1-score of 81.29%, a Mean Intersection over Union (MIoU) of 76.19%, and an Overall Accuracy (OA) of 99.35% on the road marking dataset. Additionally, on the CamVid dataset, RoadFormer achieves an F1-score of 86.96%, an MIoU of 81.14%, and an OA of 98.08%. RoadFormer excels in multi-scale feature extraction, offering robust support for high-definition map creation.
Jingrong Wang, Yu Chen 0029, Jiao Zhan, Zhifei Liu
Eng. Appl. Artif. Intell.2
2025 Performance Analysis of Communication Scheduling Schemes for Distributed Deep Learning
abstract
With the growing popularity of large-scale deep neural networks, efficient communication scheduling has become crucial in distributed deep learning systems to reduce overall training time. In multi-job distributed training scenarios, current communication scheduling methods do not effectively utilize the periodic communication patterns of deep learning training (DLT) jobs to reduce the potential link contention. When multiple tenants run concurrent jobs and compete for network resources, training performance can degrade due to increased network contention. In this paper, we focus on exploring the potential of leveraging periodic communication patterns in scheduling DLT jobs. We analyze the performance of static shift-based scheduling strategies based on the least common multiple (LCM) alignment in handling multi-job communication conflicts. Through theoretical analysis and validation via real-world experiments, we expose fundamental limitations of shift-based scheduling strategies, which fail to improve training throughput in about 73 % of multi-job scenarios. Our research work provides guidance for future research on understanding traffic patterns of DLT jobs and lays the groundwork for communication scheduling optimization in multi-tenant clusters.
Jinhao Luo, Jingrong Wang, Adrian Fiech, Kaiyang Liu
LCN3
2025 A Transformer-Block-Wise Collaborative Training Mechanism with Hybrid Parallelism Over Heterogeneous Networks
abstract
With the rise of AI-Generated Content (AIGC) services in wireless networks, efficient and high-quality distributed training of Large Language Models (LLMs) has become essential for enabling the large-scale application of next generation AI technologies. However, the extensive parameters of LLMs impose significant demands on memory, computing power and communication resources in heterogeneous networks. To efficiently utilize the dispersed network resources, this paper presents a First-Pipeline- Then-Federated Learning (FPTFL) approach with a hybrid parallel scheduling strategy to facilitate the training of Transformer-based LLMs. We propose a block-wise splitting mechanism to partition the Transformer's encoder into distinct segments, which are deployed cross individual devices. The encoder parameters and intermediate smashed data are uploaded to the edge server, where the whole model is updated through federated aggregation. Particularly, we develop a fine-grained computation-efficient method based on pipeline parallelism, enabling the segments to cooperatively train the entire encoder. An optimization problem is formulated to determine the LLM segments and the number of micro-batches under network resource constraints, with the goal of minimizing the total latency of LLM training services. Simulation results demonstrate that our approach enables Transformer-based model training on resource-constrained devices, preserves model performance, and reduces waiting time.
Jiewei Chen, Jingrong Wang, Shao-Yong Guo 0001, Jiakai Hao, Xuesong Qiu 0001, Zehui Xiong
WCNC2
2025 Communication-Efficient Network Topology in Decentralized Learning: A Joint Design of Consensus Matrix and Resource Allocation
abstract
In decentralized machine learning over a network of workers, each worker updates its local model as a weighted average of its local model and all models received from its neighbors. Efficient consensus weight matrix design and communication resource allocation can increase the training convergence rate and reduce the wall-clock training time. In this paper, we jointly consider these two factors and propose a novel algorithm termed Communication-Efficient Network Topology (CENT), which reduces the latency in each training iteration by removing unnecessary communication links. CENT enforces communication graph sparsity by iteratively updating, with a fixed step size, a trade-off factor between the convergence factor and a weighted graph sparsity. We further extend CENT to one with an adaptive step size (CENT-A), which adjusts the trade-off factor based on the feedback of the objective function value, without introducing additional computation complexity. We show that both CENT and CENT-A preserve the training convergence rate while avoiding the selection of poor communication links. Numerical studies with real-world machine learning data in both homogeneous and heterogeneous scenarios demonstrate the efficacy of CENT and CENT-A and their performance advantage over state-of-the-art algorithms.
Jingrong Wang, Ben Liang 0001, Zhongwen Zhu, Emmanuel Thepie Fapi, Hardik Dalal
IEEE Trans. Netw.1
2024 Sampling-Based Multi-Job Placement for Heterogeneous Deep Learning Clusters
abstract
Heterogeneous deep learning clusters commonly host a variety of distributed learning jobs. In such scenarios, the training efficiency of learning models is negatively affected by the slowest worker. To accelerate the training process, multiple learning jobs may compete for limited computational resources, posing significant challenges to multi-job placement among heterogeneous workers. This paper presents a heterogeneity-aware scheduler to solve the multi-job placement problem while taking into account job sizing and load balancing, minimizing the average Job Completion Time (JCT) of deep learning jobs. A novel scheme based on proportional training workload assignment, feasible solution categorization, and matching markets is proposed with theoretical guarantees. To further reduce the computational complexity for low latency decision-making and improve scheduling fairness, we propose to construct the sparsification of feasible solution categories through sampling, which has negligible performance loss in JCT. We evaluate the performance of our design with real-world deep neural network benchmarks on heterogeneous computing clusters. Experimental results show that, compared to existing solutions, the proposed sampling-based scheme can achieve 1) results within 2.04% of the optimal JCT with orders-of-magnitude improvements in algorithm running time, and 2) high scheduling fairness among learning jobs.
Kaiyang Liu, Jingrong Wang, Zhiming Huang 0002, Jianping Pan 0001
IEEE Trans. Parallel Distributed Syst.2
2023 Knowledge Graph of Urban Firefighting with Rule-Based Entity Extraction
Nady Slam, Zixiang Zhang, Jingrong Wang
EANN5
2023 Distributed Online Min-Max Load Balancing with Risk-Averse Assistance
abstract
Motivated by a wide range of applications from parallel computing to distributed learning, we study distributed online load balancing among multiple workers. We aim to minimize the pointwise maximum over the workers' local cost functions. We propose a novel algorithm termed Distributed Online Load Balancing with rIsk-averse assistancE (DOLBIE), which jointly considers the worker heterogeneity and system dynamics. The workload is distributed to workers in an online manner, where the underloaded workers learn to provide an appropriate amount of assistance to the most overloaded worker for the next online round without making themselves overwhelmed. In DOLBIE, all workers participate in updating the workload simultaneously, and no computationally intensive gradient or projection calculation is required. DOLBIE can be implemented in both the master-worker and fully-distributed architectures. We analyze the worst-case performance of DOLBIE by deriving an upper bound on its dynamic regret. We further demonstrate the application of DOLBIE to online batch-size tuning in distributed machine learning. Our experimental results show that, in comparison with state-of-the-art alternatives, DOLBIE can substantially speed up the training process and reduce the workers' idle time.
Jingrong Wang, Ben Liang 0001
ICDCS1
2023 Adaptive and Scalable Caching With Erasure Codes in Distributed Cloud-Edge Storage Systems
abstract
Erasure codes have been widely used to enhance data resiliency with low storage overheads. However, in geo-distributed cloud storage systems, erasure codes may incur high service latency as they require end users to access remote storage nodes to retrieve data. An elegant solution to achieving low latency is to deploy caching services at the edge servers close to end users. In this paper, we propose adaptive and scalable caching schemes to achieve low latency in the cloud-edge storage system. Based on the measured data popularity and network latencies in real time, an adaptive content replacement scheme is proposed to update caching decisions upon the arrival of requests. Theoretical analysis shows that the reduced data access latency of the replacement scheme is at least 50% of the maximum reducible latency. With the low computation complexity of our design, nearly no extra overheads will be introduced when handling intensive data flows. For further performance improvements without sacrificing its efficiency, an adaptive content adjustment scheme is presented to replace the subset of cached contents that incur the aforementioned performance loss. Driven by real-world data traces, extensive experiments based on Amazon Simple Storage Service demonstrate the effectiveness and efficiency of our design.
Kaiyang Liu, Jun Peng 0001, Jingrong Wang, Zhiwu Huang, Jianping Pan 0001
IEEE Trans. Cloud Comput.3
2023 Sampling-Based Caching for Low Latency in Distributed Coded Storage Systems
abstract
Caching has been considered as a promising solution to achieve low latency in distributed erasure coded storage systems. The previous research work categorizes all feasible caching decisions into a set of cache partitions, and then obtains the optimal solution by applying the market clearing price on each cache partition. While enjoying the ultimate performance of low data access latency, the optimal scheme suffers from high computation overheads when applied to large-scale storage systems. This paper presents SampleX, which constructs the sparsification of cache partitions through sampling to approximate the optimal caching scheme with substantially reduced computation complexity. Theoretical analysis guarantees the performance of SampleX. Furthermore, SampleX is implemented in a streaming fashion, capturing the characteristics of recent traffic for online cache content replacement. Trace-driven experimental results show that online SampleX is up to 95× faster than the state-of-the-art online scheme while only incurring a performance loss of 0.81%.
Kaiyang Liu, Jingrong Wang, Heng Li 0005, Jun Peng 0001, Jianping Pan 0001
IEEE Trans. Serv. Comput.2
2022 A Learning-Based Data Placement Framework for Low Latency in Data Center Networks
abstract
Low-latency data service is an increasingly critical challenge for data center applications. In modern distributed storage systems, proper data placement helps reduce the data movement delay, which can contribute to the service latency reduction tremendously. Existing data placement solutions have often assumed the prior distribution of data requests or discovered it via trace analysis. However, data placement is a difficult online decision-making problem faced with dynamic network conditions and time-varying user request patterns. The conventional static model-based solutions are less effective to handle the dynamic system. With an overall consideration of data movement and analytical latency, we develop a reinforcement learning-based framework DataBot+, automatically learning the optimal placement policies. DataBot+ adopts neural networks, trained with a variant of$Q$-learning, whose input is the real-time data flow measurements and whose output is a value function estimating the near-future latency. For instantaneous decision making, DataBot+ is decoupled into two asynchronous production and training components, ensuring that the training delay will not introduce extra overheads to handle the data flows. Evaluation results driven by real-world traces demonstrate the effectiveness of our design.
Kaiyang Liu, Jun Peng 0001, Jingrong Wang, Boyang Yu 0001, Zhuofan Liao, Zhiwu Huang, Jianping Pan 0001
IEEE Trans. Cloud Comput.3
2022 Optimal Caching for Low Latency in Distributed Coded Storage Systems
abstract
Erasure codes have been widely considered as a promising solution to enhance data reliability at low storage costs. However, in modern geo-distributed storage systems, erasure codes may incur high data access latency as they require data retrieval from multiple remote storage nodes. This hinders the extensive application of erasure codes to data-intensive applications. This paper proposes novel caching schemes to achieve low latency in distributed coded storage systems. Assuming that future data popularity and network latency information are available, an offline caching scheme is proposed to explore the optimal caching solution for low latency. The proposed scheme categorizes all feasible caching decisions into a set of cache partitions, and then obtains the optimal caching decision through market clearing price for each cache partition. Furthermore, guided by the optimal scheme, an online caching scheme is proposed according to the measured data popularity and network latency information in real time, without the need to completely override the existing caching decisions. Both theoretical analysis and experiment results demonstrate that the online scheme can approximate the offline optimal scheme well with dramatically reduced computation complexity.
Kaiyang Liu, Jun Peng 0001, Jingrong Wang, Jianping Pan 0001
IEEE/ACM Trans. Netw.3
2020 Online UAV-Mounted Edge Server Dispatching for Mobile-to-Mobile Edge Computing
abstract
Mobile edge computing (MEC) has been considered as a promising technology to handle computation-intensive and delay-sensitive tasks in the Internet of Things (IoT) ecosystem, such as smart city and smart tourism. However, due to user mobility, edge servers with fixed deployment are not flexible enough to handle time-varying user tasks in hot-spot areas. In this article, a novel online unmanned aerial vehicle (UAV)-mounted edge server dispatching scheme is proposed to provide flexible mobile-to-MEC services. UAVs are dispatched to the appropriate hover locations by geographically merging tasks into several hot-spot areas. Theoretical analysis guarantees the worst case performance bound. Extensive evaluation driven by real-world mobile requests shows that while maintaining a good latency fairness, the mobile server dispatching scheme can serve more user equipments (UEs) as well as achieve a high resource utilization. Moreover, the hybrid scheme can satisfy even more user demands while dispatching fewer UAVs with a higher server utilization.
Jingrong Wang, Kaiyang Liu, Jianping Pan 0001
IEEE Internet Things J.1
2020 Scalable and Adaptive Data Replica Placement for Geo-Distributed Cloud Storages
abstract
In geo-distributed cloud storage systems, data replication has been widely used to serve the ever more users around the world for high data reliability and availability. How to optimize the data replica placement has become one of the fundamental problems to reduce the inter-node traffic and the system overhead of accessing associated data items. In the big data era, traditional solutions may face the challenges of long running time and large overheads to handle the increasing scale of data items with time-varying user requests. Therefore, novel offline community discovery and online community adjustment schemes are proposed to solve the replica placement problem in a scalable and adaptive way. The offline scheme can find a replica placement solution based on the average read/write rates for a certain period of time. The scalability can be achieved as 1) the computation complexity is linear to the amount of data items and 2) the data-node communities can evolve in parallel for a distributed replica placement. Furthermore, the online scheme is adaptive to handle the bursty data requests, without the need to completely override the existing replica placement. Driven by real-world data traces, extensive performance evaluations demonstrate the effectiveness of our design to handle large-scale datasets.
Kaiyang Liu, Jun Peng 0001, Jingrong Wang, Weirong Liu 0001, Zhiwu Huang, Jianping Pan 0001
IEEE Trans. Parallel Distributed Syst.3
2018 Learning Based Mobility Management Under Uncertainties for Mobile Edge Computing
abstract
Mobile edge computing (MEC) offloads computation-intensive applications and overcomes the long latency by pushing data traffic towards the network edges. With base stations (BSs) densely deployed in a hot-spot area to improve user experience, mobile user equipments (UEs) have multiple choices to offload tasks to edge servers by jointly considering both the channel condition and the computing capacity. However, precise full system information is hard to be synchronized between BSs and UEs for mobility management decision making. In this paper, a Q-Iearning based mobility management scheme is proposed to handle the system information uncertainties. Each UE observes the task delay as an experience and automatically learns the optimal mobility management strategy through trial and error. Simulations show that the proposed scheme manifests the superiority in dealing with the uncertainties. Compared with the traditional received signal strength-based handover scheme, the proposed scheme reduces the task delay by about 30%.
Jingrong Wang, Kaiyang Liu, Minming Ni, Jianping Pan 0001
GLOBECOM1
2018 Learning-based Cooperative Sound Event Detection with Edge Computing
abstract
In this paper, we propose a novel real-time sound event detection framework, which combines multi-label learning and edge computing, to classify and localize abnormal sound events for city surveillance. Multiple devices equipped with acoustic sensors are deployed to collect the audio information. A learning-based approach is introduced to address the difficulties of accurately classifying the temporally overlapping acoustic events in a noisy environment. Then, edge computing is adopted to handle the high processing complexity of the learned analytics. Computation-intensive tasks of classification and localization can be offloaded to the nearby edge server for low-latency sound detection. An ensemble-based cooperative decision-making algorithm is also presented to aggregate the information from distributed devices in order to obtain better classification results. Extensive evaluations show the effectiveness of edge computing which helps reduce the time latency as well as the superiority of cooperative post-processing on the edge server to obtain a high accuracy.
Jingrong Wang, Kaiyang Liu, George Tzanetakis, Jianping Pan 0001
IPCCC1
2018 Learning-based Adaptive Data Placement for Low Latency in Data Center Networks
abstract
Low-latency data access is an important challenge for data center networks. Proper placement of the data items can reduce the data travel time in the distributed storage systems, which contributes significantly to the latency reduction. Most existing data placement approaches have often assumed the prior distribution of data requests or discovered so through trace analysis. However, the traditional static model-based solutions are less effective to handle the system uncertainties in a dynamic environment. We present DataBot, a reinforcement learning-based adaptive framework, to learn the optimal data placement policies faced with the dynamic network conditions and time-varying request patterns. DataBot utilizes a neural network, trained with a variant of Q-learning, whose input is the realtime data flow measurements and whose output is a value function estimating the near-future latency. For rapid decision making, DataBot is divided into two decoupled production and training components, ensuring that the convergence time of the training will not introduce more overheads to serve the read/write requests. Evaluation results demonstrate that the average write and read latency of the whole system can be lowered by about 35% and 40%, respectively.
Kaiyang Liu, Jingrong Wang, Zhuofan Liao, Boyang Yu 0001, Jianping Pan 0001
LCN2
2017 Handover Performance Improvement for Ultra Dense Network of High-Speed Railway
abstract
Nowadays the implementation of mobile Internet access services for high-speed railway (HSR) passengers are facing several challenges. In particular, the extra penetration loss of the train, frequent handovers, poor received signal and handover lag caused by fast time-varying fading channel may result in serious communication interruption. This paper investigates the mobility management of heterogeneous ultra dense network for HSR, in which macro base stations (BS) provide seamless coverage and millimeter-wave (mmWave) super micro BS provide high data rate services. With this network approach, a gray model (GM) prediction based handover algorithm is designed to prevent handover lag. To ensure service continuity for passengers, a combination factor is introduced into the handover decision to select target BS. Extensive simulation results prove the superiority of the proposed handover scheme in reducing handover failure probability and providing passengers with continuous services.
Jingrong Wang, Xuanjin Yang, Shuyue Zhao, Lifei Zhang, Minming Ni
VTC Spring1