VLDB 2026 Research / reviewers in the wild / expert
A-Long Jin
dblp:25/10245
· DBLP profile ↗
21ranked-venue papers
2as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpaceMeet: Bringing Conferencing Closer Through In-Orbit Conferencing Services
Songshi Dou, Feihu Jin, Jinxian Wu, A-Long Jin, Kwan Lawrence Yeung |
IEEE Trans. Serv. Comput. | 4 |
| 2026 | MetaRS: A Self-Intelligent Rate-Splitting Approach for Co-Existing Space-Air-Ground Integrated NetworksabstractThe rise of heterogeneous aerial and space platforms within Space-Air-Ground Integrated Networks (SAGINs) introduces significant challenges, as the limited spectrum resources force these platforms to operate within shared frequency bands, resulting in co-existing systems. Effective interference management in such networks requires both the design of communication channels and the dynamic mitigation of interference between them. Prior research has largely focused on interference mitigation with fixed communication links, often overlooking adaptive channel selection, which can result in performance degradation. In this study, we address this limitation by introducing MetaRS, an innovative, self-intelligent rate-splitting solution designed for more flexible interference management in co-existing SAGINs. MetaRS enables adaptive channel and communication scheme selection, by leveraging a Fully-Distributed Rate-Splitting Multiple Access (FD-RSMA)-based framework enhanced with a one-pass diffusion model. Specifically, the FD-RSMA-based framework allows MetaRS to dynamically shift its interference management strategy according to the current network status. The integration of the diffusion model further enhances MetaRS by allowing it to recognize and adapt to real-time channel conditions and user deployment, thereby enabling self-intelligent interference mitigation. Simulation results demonstrate that MetaRS significantly outperforms conventional SDMA, RSMA, and FD-RSMA approaches. This improvement stems from MetaRS’s joint optimization of channel selection and its adaptive, intelligent interference management capabilities, which effectively balance channel utilization and mitigate interference in complex, multi-platform environments. Shengyu Zhang 0003, Feng Wang 0049, Jia Shi 0001, A-Long Jin, Zan Li 0001, Tony Q. S. Quek |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingabstractThe parameter size of modern large language models (LLMs) can be scaled up to the trillion-level via the sparsely-activated Mixture-of-Experts (MoE) technique to avoid excessive increase of the computational costs. To further improve training efficiency, pipelining computation and communication has become a promising solution for distributed MoE training. However, existing work primarily focuses on scheduling tasks within the MoE layer, such as expert computing and all-to-all (A2A) communication, while neglecting other key operations including multi-head attention (MHA) computing, gating, and all-reduce communication. In this paper, we propose FlowMoE, a scalable framework for scheduling multi-type task pipelines. First, FlowMoE constructs a unified pipeline to consistently scheduling MHA computing, gating, expert computing, and A2A communication. Second, FlowMoE introduces a tensor chunk-based priority scheduling mechanism to overlap the all-reduce communication with all computing tasks. We implement FlowMoE as an adaptive and generic framework atop PyTorch. Extensive experiments with 675 typical MoE layers and four real-world MoE models across two GPU clusters demonstrate that our proposed FlowMoE framework outperforms state-of-the-art MoE training frameworks, reducing training time by14%-57%, energy consumption by 10%-39%, and memory usage by 7%-32%. FlowMoE’s code is anonymously available at https://anonymous.4open.science/r/FlowMoE. Yunqi Gao, Bing Hu 0002, Mahdi Boloursaz Mashhadi, A-Long Jin, Yanfeng Zhang 0001, Pei Xiao 0001, Rahim Tafazolli, Mérouane Debbah |
NeurIPS | 4 |
| 2025 | A Hybrid Method for Source Direction Finding With Radio Frequency Interference and Gaussian White NoiseabstractThis paper presents a hybrid data-driven method, termed moving average-Hankel-dynamic mode decomposition (MAHankDMD), for joint direction of arrival (DOA) and frequency estimation in environments affected by both radio frequency interference (RFI) and Gaussian white noise. The proposed approach integrates two key components: (1) a moving average-DMD filter that effectively mitigates Gaussian white noise and separates RFI from the source signal, and (2) a Hankel-DMD method that accurately estimates the DOA of the filtered signal and associates it with the corresponding frequency. The moving average-DMD stage first enhances the signal-to-noise ratio and improves the robustness of the estimation process through noise and inference mitigation, while the subsequent Hankel-DMD stage enables reliable parameter extraction even for overlapping sginals or strong interference conditions. Numerical simulations demonstrate the robustness of MAHankDMD, showing its ability to precisely estimate both DOA and frequency under challenging conditions involving RFI and Gaussian white noise interference. The proposed algorithm thus provides an effective solution for channel parameter estimation in complex noisy environments. Wenchao Xu 0001, Antonios Argyriou, A-Long Jin, Tianquan Tang, Peifeng Ma, Lijun Jiang |
IEEE Internet Things J. | 4 |
| 2025 | A Tensor-Based Data-Driven Approach for Multidimensional Harmonic Retrieval and Its Application for MIMO Channel SoundingabstractIn wireless channel sounding, accurately estimating multiple parameters within a multipath signal, such as azimuth, elevation, Doppler shift, and delay, necessitates addressing the challenges posed by the multidimensional harmonic retrieval (MHR) problem. To overcome these complexities, we propose a framework based on high-order dynamic mode decomposition (HODMD) that designed for robustly estimating frequencies of interest from high-dimensional sinusoidal signals, particularly in additive white Gaussian noise conditions. The HODMD approach, a hybrid algorithm amalgamating high-order singular value decomposition (HOSVD) and dynamic mode decomposition (DMD), operates by initially decomposing observed tensorial data into a core tensor and R mode matrices through HOSVD. Subsequently, DMD is applied to analyze each mode matrix individually, decomposing it into dynamic modes and DMD eigenvalues. The imaginary component of the DMD eigenvalues yields frequencies along the rth dimension. By uniformly applying this analysis to all mode matrices, multiple frequencies of interest are efficiently obtained. Furthermore, the integration of HOSVD, DMD, and moving average techniques in the proposed method is designed to mitigate noise interference during the MHR process. We conduct several numerical experiments and present a real-life example, i.e., the double-direction multiple-input and multiple-output (MIMO) channel sounding, to validate the effectiveness of the proposed HODMD approach. Results demonstrate that HODMD outperforms comparable approaches, particularly in scenarios characterized by high-signal-to-noise ratios. Notably, the proposed method exhibits the capability to estimate the number of tones in undamped cases during the decomposition process. Hence, our work contributes a practical and effective tensor-based solution to the MHR problem, particularly in the context of channel parameter estimation for MIMO systems. Wenchao Xu 0001, A-Long Jin, Min Li 0032, Ping Yuan, Lijun Jiang |
IEEE Internet Things J. | 3 |
| 2025 | Enhanced Multidimensional Harmonic Retrieval in MIMO Wireless Channel SoundingabstractThis article introduces a recursive parallel dynamic mode decomposition (RPDMD) scheme tailored for multidimensional harmonic retrieval (MHR), specifically applied to MIMO wireless channel sounding. The RPDMD algorithm is devised to address the complexities inherent in multidimensional scenarios, leveraging the dynamic mode decomposition (DMD) framework within a recursive parallel structure. Initially, the observed tensorial multidimensional harmonic data is transformed into a 2-D matrix format along the rth dimension. Subsequently, DMD dissects this matrix data into eigenvalues and their associated modes. The real and imaginary components of the DMD eigenvalues yield damping factors and frequencies in the rth dimension, respectively. Furthermore, recursive DMD is employed to scrutinize each mode independently for parameter retrieval across the remaining dimensions, enabling parallel analysis. Ultimately, this high-dimensional correlated decomposition scheme delivers paired damping factors and frequencies for all tones. Notably, the proposed approach can ascertain the number of tones in undamped sinusoidal signals, making it particularly suitable for MHR even without prior knowledge of the source count. Numerical experiments demonstrate the accuracy and robustness of the RPDMD scheme, with comparative analysis indicating that RPDMD outperforms similar methods, achieving optimal results with minimal mean square error in high signal-to-noise ratio scenarios. This work presents an effective data-driven solution for the MHR problem in MIMO wireless channel sounding. Wenchao Xu 0001, A-Long Jin, Tianquan Tang, Min Li 0032, Peifeng Ma, Lijun Jiang |
IEEE Internet Things J. | 3 |
| 2025 | HeaPS: Heterogeneity-aware participant selection for efficient federated learning
Duo Yang 0005, Bing Hu 0002, Yunqi Gao, A-Long Jin, Kwan Lawrence Yeung |
J. Parallel Distributed Comput. | 4 |
| 2025 | Rate-Splitting Multiple Access for Near-Field Communications With Imperfect CSIT and SICabstractExtremely Large-scale Antenna Array (ELAA) is increasingly recognized as a promising solution for enhancing spectral efficiency and spatial resolution in the 6G mobile system. However, realizing these benefits necessitates the development of sophisticated interference management strategies, which typically rely on perfect Channel State Information at the Transmitter (CSIT) and involve computationally intensive operations. In real-world scenarios, perfect CSIT is typically infeasible due to inherent channel estimation errors and hardware impairments, which also lead to imperfect Successive Interference Cancellation (SIC). Additionally, the computational complexity associated with precoding schemes poses a formidable challenge. To address these issues, this study proposes a Deep Learning (DL)-assisted Rate-Splitting Multiple Access (RSMA) scheme for ELAA systems. The primary objective is to maximize the geometric mean of ergodic user-rates under imperfect CSIT and SIC, thereby optimizing both fairness and system throughput. Given the prohibitively high computational complexity of conventional optimization approaches to address this optimization problem, we introduce a DL model, named GruCN, to optimize precoder design. Simulation results demonstrate that the proposed RSMA-enabled ELAA system achieves better performance in terms of fairness and robustness under imperfect CSIT. Moreover, the GruCN model exhibits remarkable efficiency and effectiveness in precoder optimization. Shengyu Zhang 0003, Feng Wang 0049, Yijie Mao, A-Long Jin, Tony Q. S. Quek |
IEEE Trans. Commun. | 4 |
| 2024 | Improving Factual Error Correction by Learning to Inject Factual ErrorsabstractFactual error correction (FEC) aims to revise factual errors in false claims with minimal editing, making them faithful to the provided evidence. This task is crucial for alleviating the hallucination problem encountered by large language models. Given the lack of paired data (i.e., false claims and their corresponding correct claims), existing methods typically adopt the ‘mask-then-correct’ paradigm. This paradigm relies solely on unpaired false claims and correct claims, thus being referred to as distantly supervised methods. These methods require a masker to explicitly identify factual errors within false claims before revising with a corrector. However, the absence of paired data to train the masker makes accurately pinpointing factual errors within claims challenging. To mitigate this, we propose to improve FEC by Learning to Inject Factual Errors (LIFE), a three-step distantly supervised method: ‘mask-corrupt-correct’. Specifically, we first train a corruptor using the ‘mask-then-corrupt’ procedure, allowing it to deliberately introduce factual errors into correct text. The corruptor is then applied to correct claims, generating a substantial amount of paired data. After that, we filter out low-quality data, and use the remaining data to train a corrector. Notably, our corrector does not require a masker, thus circumventing the bottleneck associated with explicit factual error identification. Our experiments on a public dataset verify the effectiveness of LIFE in two key aspects: Firstly, it outperforms the previous best-performing distantly supervised method by a notable margin of 10.59 points in SARI Final (19.3% improvement). Secondly, even compared to ChatGPT prompted with in-context examples, LIFE achieves a superiority of 7.16 points in SARI Final. Xingwei He 0003, Qianru Zhang, A-Long Jin, Siu-Ming Yiu |
AAAI | 3 |
| 2024 | GWPF: Communication-efficient federated learning with Gradient-Wise Parameter Freezing
Duo Yang 0005, Yunqi Gao, Bing Hu 0002, A-Long Jin, Wei Wang 0021 |
Comput. Networks | 4 |
| 2024 | WBSP: Addressing stragglers in distributed machine learning with worker-busy synchronous parallel
Duo Yang 0005, Bing Hu 0002, A-Long Jin, Kwan Lawrence Yeung |
Parallel Comput. | 4 |
| 2024 | US-Byte: An Efficient Communication Framework for Scheduling Unequal-Sized Tensor Blocks in Distributed Deep LearningabstractThe communication bottleneck severely constrains the scalability of distributed deep learning, and efficient communication scheduling accelerates distributed DNN training by overlapping computation and communication tasks. However, existing approaches based on tensor partitioning are not efficient and suffer from two challenges: 1) the fixed number of tensor blocks transferred in parallel can not necessarily minimize the communication overheads; 2) although the scheduling order that preferentially transmits tensor blocks close to the input layer can start forward propagation in the next iteration earlier, the shortest per-iteration time is not obtained. In this paper, we propose an efficient communication framework called US-Byte. It can schedule unequal-sized tensor blocks in a near-optimal order to minimize the training time. We build the mathematical model of US-Byte by two phases: 1) the overlap of gradient communication and backward propagation, and 2) the overlap of gradient communication and forward propagation. We theoretically derive the optimal solution for the second phase and efficiently solve the first phase with a low-complexity algorithm. We implement the US-Byte architecture on PyTorch framework. Extensive experiments on two different 8-node GPU clusters demonstrate that US-Byte can achieve up to 1.26x and 1.56x speedup compared to ByteScheduler and WFBP, respectively. We further exploit simulations of 128 GPUs to verify the potential scaling performance of US-Byte. Simulation results show that US-Byte can achieve up to 1.69x speedup compared to the state-of-the-art communication framework. Yunqi Gao, Bing Hu 0002, Mahdi Boloursaz Mashhadi, A-Long Jin, Pei Xiao 0001, Chunming Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionabstractThe dual-encoder has become the de facto architecture for dense retrieval.Typically, it computes the latent representations of the query and document independently, thus failing to fully capture the interactions between the query and document.To alleviate this, recent research has focused on obtaining query-informed document representations.During training, it expands the document with a real query, but during inference, it replaces the real query with a generated one.This inconsistency between training and inference causes the dense retrieval model to prioritize query information while disregarding the document when computing the document representation.Consequently, it performs even worse than the vanilla dense retrieval model because its performance heavily relies on the relevance between the generated queries and the real query.In this paper, we propose a curriculum sampling strategy that utilizes pseudo queries during training and progressively enhances the relevance between the generated query and the real query.By doing so, the retrieval model learns to extend its attention from the document alone to both the document and query, resulting in high-quality queryinformed document representations.Experimental results on both in-domain and out-ofdomain datasets demonstrate that our approach outperforms previous dense retrieval models. Xingwei He 0003, Yeyun Gong, A-Long Jin, Hang Zhang 0029, Anlei Dong, Jian Jiao 0007, Siu-Ming Yiu, Nan Duan 0001 |
EMNLP | 3 |
| 2023 | OF-WFBP: A near-optimal communication mechanism for tensor fusion in distributed deep learning
Yunqi Gao, Zechao Zhang, Bing Hu 0002, A-Long Jin, Chunming Wu 0001 |
Parallel Comput. | 4 |
| 2022 | Metric-guided Distillation: Distilling Knowledge from the Metric to Ranker and Retriever for Generative Commonsense ReasoningabstractXingwei He, Yeyun Gong, A-Long Jin, Weizhen Qi, Hang Zhang, Jian Jiao, Bartuer Zhou, Biao Cheng, Sm Yiu, Nan Duan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xingwei He 0003, Yeyun Gong, A-Long Jin, Weizhen Qi, Hang Zhang 0029, Jian Jiao 0007, Bartuer Zhou, Biao Cheng, Siu-Ming Yiu, Nan Duan 0001 |
EMNLP | 3 |
| 2022 | PS+: A Simple yet Effective Framework for Fast Training on Parameter ServerabstractIn distributed training, workers collaboratively refine the global model parameters by pushing their updates to the Parameter Server and pulling fresher parameters for the next iteration. This introduces high communication costs for training at scale, and incurs unproductive waiting time for workers. To minimize the waiting time, existing approachesoverlap communication and computationfor deep neural networks. Yet, these techniques not only require the layer-by-layer model structures, but also need significant efforts in runtime profiling and hyperparameter tuning. To make the overlapping optimizationsimpleandgeneric, in this article, we propose a new Parameter Server framework. Our solutiondecouplesthe dependency between push and pull operations, and allows workers toeagerlypull the global parameters. This way, both push and pull operations can be easily overlapped with computations. Besides, the overlapping manner offers a different way to address the straggler problem, where the stale updates greatly retard the training process. In the new framework, with adequate information available to workers, they can explicitly modulate the learning rates for their updates. Thus, the global parameters can be less compromised by stale updates. We implement a prototype system in PyTorch and demonstrate its effectiveness on both CPU/GPU clusters. Experimental results show that our prototype saves up to 54% less time for each iteration and up to 37% fewer iterations for model convergence, achieving up to 2.86× speedup over widely-used synchronization schemes. A-Long Jin, Wenchao Xu 0001, Song Guo 0001, Bing Hu 0002, Kwan Lawrence Yeung |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Friends or foes: Revisiting strategy-proofness in cloud network sharingabstractCloud networks consist of a large number of links, on which tenants have correlated and elastic bandwidth demands in the form of coflows. Ideally, a cloud network sharing policy should provide tenants with isolation guarantees on the minimum coflow progress, while at the same time attaining as high utilization as possible. Prior work shows that to achieve the optimal isolation guarantee, strategy-proofness is needed, in that tenants cannot lie about demands to obtain higher progresses. However, this requirement is derived under a simplified assumption that tenants are only interested in maximizing coflow progresses. We show in this work that a rational tenant should pursue more bandwidth allocation as a secondary objective after progress maximization. In this new model, enforcing strategy-proofness inevitably hurts the isolation guarantee. We propose a new network sharing policy to achieve the optimal isolation guarantee while attaining the highest possible utilization in spite of strategic, untruthful tenants. Trace-driven evaluations show that our policy outperforms existing alternatives with better isolation guarantee, higher utilization, and shorter coflow completion time (CCT). Wei Wang 0030, A-Long Jin |
ICNP | 2 |
| 2016 | Auction Mechanisms Toward Efficient Resource Sharing for Cloudlets in Mobile Cloud ComputingabstractMobile cloud computing offers an appealing paradigm to relieve the pressure of soaring data demands and augment energy efficiency for future green networks. Cloudlets can provide available resources to nearby mobile devices with lower access overhead and energy consumption. To stimulate service provisioning by cloudlets and improve resource utilization, a feasible and efficient incentive mechanism is required to charge mobile users and reward cloudlets. Although auction has been considered as a promising form for incentive, it is challenging to design an auction mechanism that holds certain desirable properties for the cloudlet scenario. Truthfulness and system efficiency are two crucial properties in addition to computational efficiency, individual rationality and budget balance. In this paper, we first propose a feasible and truthful incentive mechanism (TIM), to coordinate the resource auction between mobile devices as service users (buyers) and cloudlets as service providers (sellers). Further, TIM is extended to a more efficient design of auction (EDA). TIM guarantees strong truthfulness for both buyers and sellers, while EDA achieves a fairly high system efficiency but only satisfies strong truthfulness for sellers. We also show the difficulties for the buyers to manipulate the resource auction in EDA and the high expected utility with truthful bidding. A-Long Jin, Wei Song 0001, Ping Wang 0001, Dusit Niyato, Peijian Ju |
IEEE Trans. Serv. Comput. | 1 |
| 2014 | A Bayesian game analysis of cooperative MAC with incentive for wireless networksabstractIn this paper, we analyze a cooperative medium access scheme in a wireless relaying network using Bayesian games, where the participating nodes are peers subject to the half-duplex constraint and they choose to cooperate or not cooperate based on its expected utility. We first set up a one-stage game and derive the ex-post utility. A two-stage game with incomplete information is further formulated to incorporate an incentive mechanism, which charges the cooperation requester and rewards the helper via adapting their channel access probabilities. We prove that the not-cooperating strategy can always achieve a Nash equilibrium (NE) in one-stage and two-stage games as long as the access cost is constrained. More importantly, we derive the sufficient conditions so that cooperating is an NE strategy and supports higher utility than the not-cooperating strategy. Numerical results are presented to validate our analysis and demonstrate that optimal tuning factors can be determined to ensure NE and maximize system utility. Peijian Ju, Wei Song 0001, A-Long Jin |
GLOBECOM | 3 |
| 2014 | Performance analysis and enhancement for a cooperative wireless diversity network with spatially random mobile helpers
Peijian Ju, Wei Song 0001, A-Long Jin, Dizhi Zhou |
J. Netw. Comput. Appl. | 3 |
| 2012 | NLOS error mitigation in mobile location based on modified extended Kalman filterabstractGeolocation and tracking of mobile objects is an important issue in wireless communication networks. Various methods have been devised and implemented to deal with such problems whose performance is particularly limited in non-line-of-sight propagation conditions. In this paper, we take advantage of the extended Kalman filter with some extensions, modifications and improvement of previous work to reduce the NLOS error in the location measurement. One of the key contributions of this paper is to present the methods that discriminate the NLOS measurements from the LOS measurements based on the standard deviation and K-means clustering and reconstruct the LOS measurements out of the NLOS measurements by polynomial fit in order to mitigate the NLOS error. Simulation results confirm the effectiveness and accuracy of our approach in comparison with the conventional EKF algorithm. Moreover, we do not model the distribution of the NLOS error due to its intractability. A-Long Jin, Qingmin Meng |
WCNC | 2 |