Lingkun Meng

dblp:384/1059 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0007-9905-6328ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Identifying batch-integrated domains from spatial transcriptomics via graph autoencoder with contrastive learning based on cross-modality and data augmentation
abstract
Spatially resolved transcriptomics (SRT) allows for the comprehensive profiling of gene expression while preserving spatial context, advancing the study of tissue architecture. However, existing computational approaches still face key limitations, particularly the insufficient exploitation of histology information and the lack of cross-modal meaningful contrastive strategies for biological analyses. To overcome these challenges, we propose GCAST, a graph contrastive autoencoder framework for spatial transcriptomics that seamlessly integrates multimodal SRT data. GCAST adopts a self-supervised strategy to derive biologically meaningful representations directly from histology images when available. GCAST constructs dual graph views based on data augmentation and introduces a novel contrastive learning designed to leverage histology-weighted and gene-weighted features and improve biological interpretability. In addition, GCAST employs a block-diagonal graph construction to automatically align multiple datasets, achieving batch-effect correction without manual intervention. The framework not only captures spatial gene expression patterns to identify tissue domains but also adapts to datasets with or without histological images and supports the integration of multiple datasets for joint analyses. Overall, GCAST provides a unified and biologically informed framework that has the potential to facilitate deeper analyses of spatial transcriptomics.
Yexuan Mao, Lijun Quan, Guozheng Zhang, Yelu Jiang, Liangpeng Nie, Tingfang Wu, Lingkun Meng, Qiang Lyu
Briefings Bioinform.10
2026 Octopus: Accuracy-aware resource scheduling for multi-video streaming inference at the edge
Zhuzhong Qian, Andong Zhu 0001, Hesheng Sun, Lingkun Meng
Comput. Networks5
2026 Edge-cloud co-optimized 3D video analytics with synergistic neural codecs
Hebin Sun, Xiaohang Shi 0001, Xiaokun Wang 0002, Sheng Zhang 0001, Lingkun Meng, Andong Zhu 0001, Zhuzhong Qian
Comput. Networks5
2026 UDMP: Unified Delay-Driven Multipath Protocol for AI Clusters
abstract
Distributed AI model training generates bursty, low-entropy elephant flows that challenge existing single-path transport protocols in multi-stage Clos networks, leading to congestion and inefficiency. Multipath transport emerges as a promising solution, leveraging multiple paths to balance traffic and enhance resilience. However, current multipath RDMA solutions suffer from scalability, congestion control, and load-balancing inefficiencies. This paper introduces Unified Delay-driven Multipath Protocol (UDMP), a novel approach that co-designs congestion control and load balancing using network delay as a unified signal. UDMP employs delay-gradient-based congestion control to precisely resolve unavoidable congestion. Moreover, UDMP leverages delay-assisted load balancing to shift traffic across paths with minimal latency adaptively, maintaining throughput when encountering avoidable congestion. A novel Token Pool design integrates these components, eliminating per-path state overhead while achieving fine-grained traffic distribution. Implementations on DPDK and NS3 demonstrate that UDMP achieves up to 2x higher throughput and reduces flow completion times by up to 30% compared to state-of-the-art methods like MPRDMA and QP-Scaling. These results highlight UDMP’s effectiveness in meeting the stringent performance requirements of modern distributed AI training workloads.
Chengyuan Huang, Zhengqi Cui, Jun Xu 0037, Zhaochen Zhang, Li Wang 0110, Peirui Cao, Zhongming Ji, Jilei Chen, Shengju Zhang, Lingkun Meng, Ahmed M. Abdelmoniem, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Keqiang He, Chen Tian 0001
IEEE Trans. Netw.11
2026 Rail: ReArranging Inter-GPU Links for GPU-Centric Clusters
abstract
In modern GPU-centric clusters, large-scale AI training relies on two distinct communication domains: a high-bandwidth intra-node domain using proprietary interconnects (e.g., NVLink), and a scale-out inter-node network domain (e.g., RDMA). We observe that the widely-used ring algorithm, often create a significant load imbalance across these domains. This leads to the counter-intuitive scenario where the expensive, high-bandwidth intra-node domain becomes a performance bottleneck, while the inter-node network remains underutilized. This inefficiency is further exacerbated by the disparity in bandwidth provisioning: inter-node network bandwidth is generally more cost-effective and accessible, whereas intra-node bandwidth is often proprietary and more costly to scale. To address this fundamental imbalance, we propose RAIL, aimed at resolving the intra-node bottleneck by strategically rearranging inter-GPU communication paths. This rebalancing ensures that traffic loads are appropriately matched with the distinct transmission capabilities of each domain, thereby maximizing overall communication performance. RAIL incorporates a Load Distributing Strategy (LDS) that can accurately partition physical nodes into logical nodes based on the a transmission capabilities of both domains, shifting excess traffic from the overloaded intra-node domain to the underutilized network domain. Additionally, the Intra-Rail Strategy (IRS) leverages topological characteristics to ensure optimal communication paths through the network domain between logical nodes. Our evaluation demonstrates that RAIL effectively mitigates congestion and achieves a 30.7% average increase in collective communication bus bandwidth compared to the widely-used NCCL solution.
Haixin Nan, Jun Xu 0037, Peirui Cao, Zhaochen Zhang, Yizhi Wang 0004, Zhehao Lin, Yuhang Li 0002, Chengyuan Huang, Xiaohu Xu, Zhongming Ji, Shengju Zhang, Lingkun Meng, Rong Gu 0001, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.14
2025 Towards strong continuous consistency in edge-assisted VR-SGs: Delay-differences sensitive online task redistribution
Yunqi Sun, Hesheng Sun, Tuo Cao, Mingtao Ji, Zhuzhong Qian, Lingkun Meng
Comput. Networks6
2025 Adaptive scheduling of online inference pipelines at the edge: A post-hoc request-oriented approach
Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001, Sheng Zhang 0001, Sanglu Lu, Lingkun Meng
J. Syst. Archit.7
2025 Mystique: User-Level Adaptation for Real-Time Video Analytics in Edge Networks via Meta-RL
abstract
Deep neural network (DNN)-based real-time video analytics service, as a core module for numerous crucial applications such as augmented reality (AR), has garnered increasing research attention, where mobile edge computing (MEC) is often leveraged to mitigate its real-time processing burden on resource-constrained user devices. For Quality of Experience (QoE) optimization, latest works employ reinforcement learning (RL)-based methods to adaptively adjust configurations (e.g., resolution and frame rate), yet still presenting significant challenges. Firstly, we observe a substantial diversity in QoE patterns among users. Given that existing methods integrate a fixed QoE pattern in parameter training, it is intuitive to customize a policy network for each user. However, this necessitates significant training investment, failing to support on-the-fly deployment for new users. Secondly, given the dual dynamics from both the network and video content in edge video analytics system, existing methods often fall into the dilemma of fitting newly emerged and diverse system states with offline-trained fixed parameters. While it is promising to employ online learning algorithms, most of them struggle to catch up with the high dynamics. We hence proposeMystique. In real-time edge video analytics domain, it is the first meta-RL-based user-level configuration adaptation framework. Mystique establishes an initial model in offline meta training with model-agnostic meta-learning (MAML), enabling swift online adaptation to new users and system states through limited gradient updates from initial parameters. Comprehensive experiments illustrate that Mystique can improve QoE by 42% on average compared to prior works.
Xiaohang Shi 0001, Sheng Zhang 0001, Meizhao Liu, Lingkun Meng, Liu Wei, Yingcheng Gu, Kai Liu 0043, Andong Zhu 0001, Ning Chen 0010, Zhuzhong Qian
IEEE Trans. Mob. Comput.4
2025 End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video Analytics
abstract
Edge video analytics typically rely on conventional encoding standards to transmit device visual data for server-side inference. Unfortunately, general-purpose compression solutions retain unnecessary visual data that does not contribute to accuracy, resulting in significant latency throughout Video Analytics Pipeline (VAP). While previous approaches have made partial progress, they cannot systematically eliminate VAP redundancy due to uncoordinated subsystem-level optimization. Achieving complete redundancy elimination presents a major challenge, as a lack of spatio-temporal coordination risks offsetting latency gains with computational overhead (associated with redundancy elimination).Crucioovercomes these limitations with an end-to-edge framework that integrates temporally adaptive frame filtering and coordinated video compression. It leverages redesigned asymmetric autoencoders to synchronize inter-frame temporal compression with intra-frame spatial feature extraction. Additionally,Crucioemploys a one-pass decoding mechanism for encoded critical frames and dynamically adjusts batching scales to minimize latency. Empirical results demonstrateCrucio's superiority, outperforming existing solutions (e.g., DDS, Reducto, and STAC) by over a 31% reduction in end-to-end latency at 0.9 accuracy thresholds.
Andong Zhu 0001, Sheng Zhang 0001, Lingkun Meng, Xiaohang Shi 0001, Hesheng Sun, Sanglu Lu, Jie Wu 0001, Yu Liang 0001
IEEE Trans. Mob. Comput.3
2024 Multiple-Task Coded Computing for Distributed Computation Framework: Modeling and Delay Analysis
abstract
Coded computing has received significant attention thanks to its advantage in alleviating the straggler effect in distributed computation framework, which would be one of the key fundamental techniques to enable the distributed and decentralized network architectures towards 5G-advanced and 6G era. Specifically, considering the scenario that multiple tasks randomly arrive at the network, the additional task queuing makes the delay analysis of coded computing more challenging. In this paper, we consider the impacts of task queuing and characterize the end-to-end delay for coded computing systems under the multi-task scenario. To this end, we first model the end-to-end coded computing system. Then, based on the redundant task processing strategies, we consider both purging and non-purging coded computing schemes. Although the expected end-to-end delay for both schemes are intractable, we obtain closed-form expressions for their respective lower and upper bounds, which generalizes the delay results of the single-task scenario. Moreover, we show that the multi-task coded computing has a coding gain of Θ(logn) wherendenotes the number of worker nodes, even with task queues considered. Simulation results verify the accuracy of the derived delay bounds and show the effectiveness of coded computing in the multi-task scenario.
Zhongming Ji, Li Chen 0015, Hongguang Fu, Xinghua Zhao, Jun Xu 0037, Lingkun Meng
IEEE Trans. Commun.6