VLDB 2026 Research / reviewers in the wild / expert
Liangqi Yuan
dblp:327/4255
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-9994-6773ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REST: Holistic Learning for End-to-End Semantic Segmentation of Whole-Scene Remote Sensing ImageryabstractSemantic segmentation of remote sensing imagery (RSI) is a fundamental task that aims at assigning a category label to each pixel. To pursue precise segmentation with one or more fine-grained categories, semantic segmentation often requires holistic segmentation of whole-scene RSI (WRI), which is normally characterized by a large size. However, conventional deep learning methods struggle to handle holistic segmentation of WRI due to the memory limitations of the graphics processing unit (GPU), thus requiring to adopt suboptimal strategies such as cropping or fusion, which result in performance degradation. Here, we introduce the Robust End-to-end semantic Segmentation architecture for whole-scene remoTe sensing imagery (REST). REST is the first intrinsically endtoend framework for truly holistic segmentation of WRI, supporting a wide range of encoders and decoders in a plugandplay fashion. It enables seamless integration with mainstream semantic segmentation methods, and even more advanced foundation models. Specifically, we propose a novel spatial parallel interaction mechanism (SPIM) within REST to overcome GPU memory constraints and achieve global context awareness. Unlike traditional parallel methods, SPIM enables REST to process a WRI effectively and efficiently by combining parallel computation with a divideandconquer strategy. Both theoretical analysis and experiments demonstrate that REST attains nearlinear throughput scalability as additional GPUs are employed. Extensive experiments demonstrate that REST consistently outperforms existing cropping-based and fusion-based methods across a variety of scenarios, ranging from single-class to multi-class segmentation, from multispectral to hyperspectral imagery, and from satellite to drone platforms. The robustness and versatility of REST are expected to offer a promising solution for the holistic segmentation of WRI, with the potential for further extension to large-size medical imagery segmentation. Wei Chen 0089, Lorenzo Bruzzone, Bo Dang 0002, Yuan Gao 0015, Youming Deng, Jin-Gang Yu, Liangqi Yuan, Yansheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Communication-Efficient Multimodal Federated Learning: Joint Modality and Client SelectionabstractMultimodal federated learning (MFL) aims to enrich model training in FL settings where clients are collecting measurements across multiple modalities. However, key challenges to MFL remain unaddressed, particularly in heterogeneous network settings where: (i) the set of modalities collected by each client is diverse, and (ii) communication limitations prevent clients from uploading all their locally trained modality encoders to the server. In this paper, we propose Multimodal Federated learning with joint Modality and Client selection (MFedMC), a communication-efficient MFL framework that tackles these challenges through a decoupled architecture and selective uploading. Unlike traditional holistic fusion approaches, MFedMC separates modality encoders and fusion modules: modality encoders are aggregated at the server for generalization across diverse client distributions, while fusion modules remain local to each client for personalized adaptation to individual modality configurations and data characteristics. Building on this decoupled design, our joint selection algorithm incorporates two main components: (a) A modality selection methodology for each client, which weighs (i) the impact of the modality, gauged by Shapley value analysis, (ii) the modality encoder size as a gauge of communication overhead, and (iii) the frequency of modality encoder updates, denoted recency, to enhance generalizability. (b) A client selection strategy for the server based on the local loss of modality encoders at each client. Experiments on five real-world datasets demonstrate that MFedMC achieves comparable accuracy to several baselines while reducing communication overhead by over 20×. A demo video and our code are available athttps://liangqiy.com/mfedmc/. Liangqi Yuan, Dong-Jun Han, Su Wang 0007, Devesh Upadhyay, Christopher G. Brinton |
IEEE Trans. Mob. Comput. | 1 |
| 2026 | Device-Cloud Collaborative LLM Inference With Multi-Modal, Multi-Task, and Multi-Turn Conversations
Liangqi Yuan, Dong-Jun Han, Shiqiang Wang 0001, Christopher G. Brinton |
IEEE Trans. Netw. | 1 |
| 2025 | Local-Cloud Inference Offloading for LLMs in Multi-Modal, Multi-Task, Multi-Dialogue SettingsabstractCompared to traditional machine learning models, recent large language models (LLMs) can exhibit multi-task-solving capabilities through multiple dialogues and multi-modal data sources. These unique characteristics of LLMs, together with their large model size, make their deployment more challenging. Specifically, (i) deploying LLMs on local devices faces computational, memory, and energy resource issues, while (ii) deploying them in the cloud cannot guarantee real-time service and incurs communication/usage costs. In this paper, we design TMO, a local-cloud LLM inference system with Three-M Offloading: Multi-modal, Multi-task, and Multi-dialogue. TMO incorporates (i) a lightweight local LLM that can process simple tasks at high speed and (ii) a large-scale cloud LLM that can handle multi-modal data sources. We develop a resource-constrained reinforcement learning (RCRL) strategy for TMO that optimizes the inference location (i.e., local vs. cloud) and multi-modal data sources to use for each task/dialogue, aiming to maximize the long-term reward (response quality, latency, and usage cost) while adhering to resource constraints. We also contribute M4A1, a new dataset we curated that contains reward and cost metrics across multiple modality, task, dialogue, and LLM configurations, enabling evaluation of offloading decisions. We demonstrate the effectiveness of TMO compared to several exploration-decision and LLM-as-Agent baselines, showing significant improvements in latency, cost, and response quality. Liangqi Yuan, Dong-Jun Han, Shiqiang Wang 0001, Christopher G. Brinton |
MobiHoc | 1 |
| 2024 | FedMFS: Federated Multimodal Fusion Learning with Selective Modality CommunicationabstractMultimodal federated learning (FL) aims to enrich model training in FL settings where devices are collecting measurements across multiple modalities (e.g., sensors measuring pressure, motion, and other types of data). However, key challenges to multimodal FL remain unaddressed, particularly in heterogeneous network settings: (i) the set of modalities collected by each device will be diverse, and (ii) communication limitations prevent devices from uploading all their locally trained modality models to the server. In this paper, we propose Federated Multimodal Fusion learning with Selective modality communication (FedMFS), a new multimodal fusion FL methodology that can tackle the above mentioned challenges. The key idea is the introduction of a modality selection criterion for each device, which weighs (i) the impact of the modality, gauged by Shapley value analysis, against (ii) the modality model size as a gauge for communication overhead. This enables FedMFS to flexibly balance performance against communication costs, depending on resource constraints and application requirements. Experiments on the real-world ActionSense dataset demonstrate the ability of FedMFS to achieve comparable accuracy to several baselines while reducing the communication overhead by over 4x. Liangqi Yuan, Dong-Jun Han, Vishnu Pandi Chellapandi, Stanislaw H. Zak, Christopher G. Brinton |
ICC | 1 |
| 2024 | Decentralized Federated Learning: A Survey and PerspectiveabstractFederated learning (FL) has been gaining attention for its ability to share knowledge while maintaining user data, protecting privacy, increasing learning efficiency, and reducing communication overhead. Decentralized FL (DFL) is a decentralized network architecture that eliminates the need for a central server in contrast to centralized FL (CFL). DFL enables direct communication between clients, resulting in significant savings in communication resources. In this paper, a comprehensive survey and profound perspective are provided for DFL. First, a review of the methodology, challenges, and variants of CFL is conducted, laying the background of DFL. Then, a systematic and detailed perspective on DFL is introduced, including iteration order, communication protocols, network topologies, paradigm proposals, and temporal variability. Next, based on the definition of DFL, several extended variants and categorizations are proposed with state-of-the-art (SOTA) technologies. Lastly, in addition to summarizing the current challenges in the DFL, some possible solutions and future research directions are also discussed. Liangqi Yuan, Ziran Wang, Lichao Sun 0001, Philip S. Yu, Christopher G. Brinton |
IEEE Internet Things J. | 1 |
| 2023 | 3-D Indoor Positioning Based on Passive Radio Frequency Signal Strength DistributionabstractIn recent years, indoor positioning systems (IPSs) have received attention from many research fields, such as robotics, navigation, human–computer interaction, etc. However, IPS based on passive radio frequency (PRF) technology is still rare. This article proposes a 3-D IPS based on received signal strength (RSS) distribution and Gaussian process regression (GPR). Traditional RSS-based positioning systems have a transmitter with known frequencies, while in the proposed PRf signal of Opportunity—3D IPS (PRO-3DIPS), the system neither deploys new transmitters nor uses any a priori knowledge of transmitters. Furthermore, PRO-3DIPS integrates multiple Signal of Opportunity (SoOP) sources, shadowing, fading, and also captures scenario signatures. Data collection and analysis of PRF-based RSS distribution in 3-D space enables the capability of 3-D positioning. Three methods are applied and compared to find the frequency band most impacted by the scenario to achieve the best positioning performance as well as used in the estimation of RSS distribution. The RSS distribution is estimated by measuring the PRF spectrum on a fixed grid in the scenario. Using the RSS distribution, the GPR can accurately locate the receiver position. RSS at 90-gridded positions were collected in the experiment scenario, with one hundred samples at each position. The experimental result shows that a root-mean-square error (RMSE) of the proposed PRO-3DIPS is 0.292 m when the sampling distance is 1 m. The result demonstrates that the PRF spectrum is a new modality for the positioning task, which demonstrates better performance than most existing RF-based technologies. Liangqi Yuan, Houlin Chen, Robert L. Ewing, Erik Blasch, Jia Li 0010 |
IEEE Internet Things J. | 1 |
| 2023 | Federated Transfer-Ordered-Personalized Learning for Driver Monitoring ApplicationabstractFederated learning (FL) shines through in the Internet of Things (IoT) with its ability to realize collaborative learning and improve learning efficiency by sharing client model parameters trained on local data. Although FL has been successfully applied to various domains, including driver monitoring applications (DMAs) on the Internet of Vehicles (IoV), its usages still face some open issues, such as data and system heterogeneity, large-scale parallelism communication resources, malicious attacks, and data poisoning. This article proposes a federated transfer–ordered–personalized learning (FedTOP) framework to address the above problems and test on two real-world data sets with and without system heterogeneity. The performance of the three extensions, transfer, ordered, and personalized, is compared by an ablation study and achieves 92.32% and 95.96% accuracy on the test clients of two data sets, respectively. Compared to the baseline, there is a 462% improvement in accuracy and a 37.46% reduction in communication resource consumption. The results demonstrate that the proposed FedTOP can be used as a highly accurate, streamlined, privacy-preserving, cybersecurity-oriented, and personalized framework for DMA. Liangqi Yuan, Lu Su 0001, Ziran Wang |
IEEE Internet Things J. | 1 |