VLDB 2026 Research / reviewers in the wild / expert
Qinya Li
dblp:234/3770
· DBLP profile ↗
18ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0002-4881-8376ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM InferenceabstractLong-context large language models (LLMs) inference is increasingly critical, motivating a number of studies devoted to alleviating the substantial storage and computational costs in such scenarios. Layer-wise skipping methods are promising optimizations but rarely explored in long-context inference. We observe that existing layer-wise skipping strategies have several limitations when applied in long-context inference, including the inability to adapt to model and context variability, disregard for sublayer significance, and inapplicability for the prefilling phase. This paper proposes AdaSkip, an adaptive sublayer skipping method specifically designed for long-context inference. AdaSkip adaptively identifies less important layers by leveraging on-the-fly similarity information, enables sublayer-wise skipping, and accelerates both the prefilling and decoding phases. The effectiveness of AdaSkip is demonstrated through extensive experiments on various long-context benchmarks and models, showcasing its superior inference performance over existing baselines. Zhuomin He, Yizhen Yao, Pengfei Zuo, Qinya Li, Zhenzhe Zheng 0001, Fan Wu 0006 |
AAAI | 5 |
| 2025 | Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data CurationabstractIn an era overwhelmed by vast amounts of data, the effective curation of web-crawl datasets is essential for optimizing model performance. This paper tackles the challenges associated with the unstructured and heterogeneous nature of such datasets. Traditional heuristic curation methods often inadequately capture complex features, resulting in biases and the exclusion of relevant data. We introduce an advanced, learning-driven approach, Ensemble Curation Of DAta ThroUgh Multimodal Operators, called EcoDatum, which employs a novel quality-guided deduplication method to balance feature distribution. EcoDatum strategically integrates various unimodal and multimodal data curation operators within a weak supervision ensemble framework, utilizing automated optimization to effectively score each data point. EcoDatum, which significantly improves the data curation quality and efficiency, outperforms existing state-of-the-art (SOTA) techniques, ranking 1st on the DataComp leaderboard with an average performance score of 0.182 across 38 diverse evaluation datasets. This represents a 28% improvement over the DataComp baseline method, demonstrating its effectiveness in improving dataset curation and model training efficiency. Jinda Xu, Yuhao Song, Kangliang Chen, Qinya Li |
AAAI | 7 |
| 2025 | Adaptive Routing of Text-to-Image Generation Requests between Large Cloud Model and Light-Weight Edge Model
Zewei Xin, Qinya Li, Chaoyue Niu, Fan Wu 0006, Guihai Chen |
ICCV | 2 |
| 2025 | Distributed DNN-based Video Analytics with Adaptive Multi-Device CollaborationabstractThis paper optimizes the computation efficiency for DNN-based video analytics on interconnected edge devices like surveillance cameras and unmanned aerial vehicles (UAVs). Existing solutions depend heavily on computation offloading to edge servers but overlook the potential for cross-device collaboration and leave resources on some edge devices underutilized. We instead propose AdaCollab, an adaptive multi-device collaboration framework for resource-efficient distributed edge video analytics. On the one hand, as the machine perception complexity fluctuates with runtime video content, AdaCollab dynamically calibrates DNN inspection configurations on each device without degrading the model accuracy. On the other hand, AdaCollab aligns mismatched resources and workload among edge devices by selectively offloading partial computations from overloaded devices to underloaded devices, optimizing both computation and bandwidth resource utilization. Through extensive evaluations with large-scale real-world surveillance videos on testbeds of heterogeneous NVIDIA Jetson platforms, AdaCollab outperforms the SOTA baselines by up to 25.8% in frame processing through-put and 19.8% in DNN accuracy, while exhibiting enhanced robustness under restricted network bandwidth. Maozhe Zhao, Qinya Li, Shengzhong Liu, Fan Wu 0006, Guihai Chen |
ICDCS | 3 |
| 2025 | Device-Cloud Collaborative Learning Framework for Efficient Unknown Object DetectionabstractUnknown object detection aims to build detectors capable of identifying out-of-distribution objects, a critical need for applications like autonomous driving and traffic monitoring. However, limited device resources restrict existing methods from achieving accurate detection on the device side. Addressing this gap, this paper introduces a device-cloud collaborative framework named DCCUOD that enhances device model performance through efficient cloud collaboration. Our framework employs an energy-based sampling function on devices to target samples with unknown objects, coupled with a collaborative pseudo-labeling strategy to generate accurate pseudo-labels. Additionally, a two-stage training paradigm enables continuous improvements of device models on both known and unknown objects. Our study is the first to explore device-cloud collaborative learning for UOD tasks. Experimental results show that the device model is three times smaller and seven times faster than cloud models, with minimal performance trade-offs. Kewei Zhao, Xiaowei Hu 0001, Qinya Li |
ACM Multimedia | 3 |
| 2025 | MI-VFL: Feature discrepancy-aware distributed model interpretation for vertical federated learning
Rui Xing 0005, Zhenzhe Zheng 0001, Qinya Li, Fan Wu 0006, Guihai Chen |
Comput. Networks | 3 |
| 2025 | Federated multi-task learning with cross-device heterogeneous task subsets
Zewei Xin, Qinya Li, Chaoyue Niu, Fan Wu 0006, Guihai Chen |
J. Parallel Distributed Comput. | 2 |
| 2025 | Scale-Shift Attention in Polarization Domain for Fine-Grained Classification of Satellite ISAR ImagesabstractTraditional fine-grained classification focuses on visible light domains, such as animals and cars. However, these methods often perform poorly when applied to radar images and images of satellites because of challenges such as distinguishing between noise and objects and the significant scale differences among object components. To address these unique scenarios, we propose the scale-shift attention in polarization domain (SAPD) method for fine-grained classification in satellite ISAR images. Specifically, radar emits different types of waves, each with distinct imaging effects. We utilize multipolarization inputs and introduce a polarization domain query module to integrate complementary features from various radar wave types captured from the same viewpoint. This multipolarization learning helps distinguish noise and leverages complementary features from different inputs. Moreover, to handle the substantial scale differences between centimeter-level payloads and the overall meter-level structure of satellites, we propose a scale-shift attention mechanism based on shift kernels. This mechanism extends attention in the direction specified by the shift kernel by incorporating adjacent pixels, allowing for the diffusion of attention. This is beneficial for capturing features of satellite components with varying scales and shapes. Extensive experiments on a novel satellite ISAR image dataset validate the effectiveness and superiority of the SAPD. Zewei Xin, Qinya Li, Bowen Sheng, Fan Wu 0006, Guihai Chen |
IEEE Trans. Multim. | 2 |
| 2024 | Mobility-Aware Device Sampling for Statistical Heterogeneity in Hierarchical Federated LearningabstractHierarchical Federated Learning (HFL) is a practical implementation of federated learning in mobile edge computing, employing edge servers as intermediaries between mobile devices and the cloud server for device coordination and cloud communication. However, the devices are usually mobile users with unpredictable mobile trajectories and statistical heterogeneity, leading to the edge models optimized along dynamic edge data distribution directions and further resulting in instability and slow convergence of the global model. In this work, we propose a Mobility-Aware deviCe sampling algorithm in HFL, namely MACH, which can dynamically maintain the device sampling strategy at each edge to accelerate the convergence of the global model. First, we analyze the convergence bound of HFL with mobile devices under arbitrary device sampling probabilities. Based on this convergence bound, we formalize the sampling optimization problem for mobility-aware device sampling, aiming to minimize the convergence error under time-averaged cost constraints, while taking the limited device-edge wireless channel capacity into account. Next, we introduce the MACH algorithm, consisting of two underlying components: experience updating and edge sampling. Experience updating utilizes an upper confidence bound method to estimate device statistical information online, and edge sampling customizes a sampling strategy on each edge based on the estimated device statistical information. Finally, extensive experimental results through real-world mobile device trajectories validate that MACH can reduce the time required to achieve a target accuracy by 25.00% - 56.86%. Songli Zhang, Zhenzhe Zheng 0001, Qinya Li, Fan Wu 0006, Guihai Chen |
ICDCS | 3 |
| 2024 | Federated Optimization Under Intermittent Client AvailabilityabstractFederated learning is a new distributed machine learning framework, where numerous heterogeneous clients collaboratively train a model without sharing training data. In this work, we consider a practical and ubiquitous issue when deploying federated learning in mobile environments: intermittent client availability, where the set of eligible clients may change during the training process. Such intermittent client availability would seriously deteriorate the performance of the classical federated averaging algorithm (FedAvg). Thus, we propose a simple distributed nonconvex optimization algorithm, called federated latest averaging (FedLaAvg), which leverages the latest gradients of all clients, even when the clients are not available, to jointly update the global model in each iteration. Our theoretical analysis shows that FedLaAvg achieves guaranteed convergence and a sublinear speedup with respect to the total number of clients. We implement FedLaAvg along with several baselines and evaluate them over the benchmarking MNIST and Sentiment140 data sets. The evaluation results demonstrate that FedLaAvg achieves more stable training than FedAvg in both convex and nonconvex settings and reaches a sublinear speedup. Source code and online supplement are available at the IJOC GitHub site ( http://dx.doi.org/10.1287/ijoc.2022.0057.cd , https://github.com/INFORMSJoC/2022.0057 ). History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Leaning. Funding: This work was supported by the National Key R&D Program of China [Grant 2022ZD0119100], the National Natural Science Foundation of China (NSFC) [Grants 61972252, 61972254, 62072303, 62025204, 62132018, 62202296, and 62202297], the Alibaba Innovation Research (AIR) Program, and the Tencent Rhino Bird Key Research Project. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0057 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0057 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Yikai Yan, Chaoyue Niu, Zhenzhe Zheng 0001, Shaojie Tang 0001, Qinya Li, Fan Wu 0006, Chengfei Lyu, Yang-He Feng, Guihai Chen |
INFORMS J. Comput. | 6 |
| 2024 | VAP: Online Data Valuation and Pricing for Machine Learning Models in Mobile HealthabstractMobile health (mHealth) applications, benefiting from mobile computing, have generated numerous mHealth data. However, they are dispersed across isolated devices, which hinders discovering insights underlying the aggregated data. Considering the online characteristics of mHealth, in this work, we present the first online dataVAluation andPricing mechanism, namely VAP, to incentive users to contribute mHealth data for machine learning (ML) tasks in mHealth systems. Under the Bayesian framework, we propose a new metric based on the concept of entropy to calculate data valuation during model training in an online manner. In proportion to the data valuation, we then determine payments as compensations for users to contribute their data. We formulate this pricing problem as a contextual multi-armed bandit with the goal of profit maximization and propose a new algorithm based on the characteristics of pricing. Furthermore, to tackle the budget constraint, we incorporate a two-stage multi-armed bandit with a knapsack method. We also extend VAP to advanced ML models by computing the entropy on the prediction space. Finally, we have evaluated VAP on two real-world mHealth data sets. Evaluation results show that VAP outperforms the state-of-the-art data valuation and pricing mechanisms in terms of computational complexity and extracted profit. Anran Xu 0003, Zhenzhe Zheng 0001, Qinya Li, Fan Wu 0006, Guihai Chen |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Distributed Model Interpretation for Vertical Federated Learning with Feature DiscrepancyabstractVertical federated learning (VFL) allows multiple clients with misaligned feature spaces to collaboratively accomplish the global model training. Applying VFL to high stakes decision scenarios greatly requires model interpretation for decision reliability and diagnosis. However, the feature discrepancy in VFL raises new issues for model interpretation in distributed setting: one is from the local-global perspective, where the local importance of features is not equal to the global importance; and the other is from the local-local perspective, where information asymmetry among clients causes difficulty in identifying overlapped features. In this work, we propose a new distributed Model Interpretation method for Vertical Federated Learning with feature discrepancy, namely MI-VFL. In particular, to deal with the local-global discrepancy, MI-VFL leverages the law of total probability to adjust the local importance of features and ensures the completeness of the selected features using adversarial game. To handle the local-local discrepancy, MI-VFL builds a federated adversarial learning model to efficiently identify the overlapped features once, rather than performing client-to-client intersections multiple times. We extensively evaluate MI-VFL on six synthetic datasets and five real-world datasets. The evaluation results reveal that MI-VFL can accurately identify the important features, suppress the overlapped features, and thus improve the model performance. Rui Xing 0005, Zhenzhe Zheng 0001, Qinya Li, Fan Wu 0006, Guihai Chen |
IWQoS | 3 |
| 2023 | PADP-FedMeta: A personalized and adaptive differentially private federated meta learning mechanism for AIoT
Fang Dong 0001, Xinghua Ge, Qinya Li, Jinghui Zhang 0001, Dian Shen, Xiao Liu 0004, Gang Li 0009, Fan Wu 0006, Junzhou Luo |
J. Syst. Archit. | 3 |
| 2023 | Capitalize Your Data: Optimal Selling Mechanisms for IoT Data ExchangeabstractMore and more IoT data is being traded online in cloud-based data marketplaces due to the fast-growing market demand. Within the current data selling mechanisms, data consumers have difficulties in making purchasing decisions due to uncertain IoT data quality and inflexible pricing interface. To resolve these issues, potential solutions could be to launch data demonstrations and release free sampling data to reduce the uncertainty about data quality, and to charge based on the volume of data actually used to enable flexible pricing. However, there is still no clear understanding of economic benefits of these mechanisms. In this paper, we design the optimal data selling mechanisms for IoT data exchange, and derive the following two results. First, whether to deploy a data demonstration and how much free sampling data to release depend on the extent of data consumers' inaccuracy perceptions for data quality, which varies over a wide range in IoT applications. We found that the data vendor has no incentive to conduct these strategies if data consumers extremely overestimate data quality. Second, although flexible data pricing mechanisms provide convenience for real-time and streaming IoT data exchange, it brings less economic benefits to the data vendor compared with the fixed pricing scheme, which sells the whole data set with a fixed price. We evaluate the optimal selling mechanisms on a real-world Taxi GPS data set, and evaluation results verify the insights derived from our theoretical analysis. Qinya Li, Zun Li 0002, Zhenzhe Zheng 0001, Fan Wu 0006, Shaojie Tang 0001, Zhao Zhang 0002, Guihai Chen |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | An Efficient, Fair, and Robust Image Pricing Mechanism for Crowdsourced 3D ReconstructionabstractA large-scale and high-quality image collection is a fundamental demand in the 3D reconstruction scenario. Crowdsourcing can help us to collect lots of diversified images. However, it is difficult to attract people to accomplish tasks due to their self-interest. Besides, the quality of collected images is various. Low-quality images may degrade the performance of 3D reconstruction. To avoid low-quality images and motivate participants to provide high-quality images, we take image quality into account when allocating rewards. In this article, we propose a pricing mechanism, called ImgPricing, to determine the rewards of participants in 3D reconstruction. We model the process of image collection as a cooperative game, and regard image quality and the arrival sequence of images as critical factors in the reward allocation. ImgPricing differs from traditional pricing schemes, e.g., Shapley value and Banzhaf power index-based methods, in that it introduces the images’ arrival sequence to be an indispensable element. We lastly implement ImgPricing on the Android platform and extensively evaluate its performance. Our evaluation results demonstrate that ImgPricing outperforms other existing schemes in terms of computational efficiency, fairness, and robustness. In brief, our image quality-based pricing mechanism for crowdsourced 3D reconstruction is feasible and effective. Qinya Li, Fan Wu 0006, Guihai Chen |
IEEE Trans. Serv. Comput. | 1 |
| 2020 | Generative Adversarial Networks-based Privacy-Preserving 3D ReconstructionabstractA large-scale image collection is crucial to the success of 3D reconstruction. Crowdsourcing, as a new pattern, can be utilized to collect high-quality images in an efficient way. However, the sensitive information in images may be exposed during the image transmission process. The general privacy policies perhaps will cause the loss or change of critical information, which may give rise to a decline in the performance of 3D reconstruction. Hence, how to achieve image privacy-preserving while guaranteeing to reconstruct a complete 3D model is important and significant. In this paper, we propose PicPrivacy to address this problem, which consists of three parts. (1) Using a pre-trained deep convolution neural network to segment sensitive information and erase it from images. (2) Using a GAN-based image feature completion algorithm to repair blank regions and minimize the absolute information gap between generated images and raw ones. (3) Taking generated images as the input of 3D reconstruction and using a structure-from-motion algorithm to reconstruct 3D models. Finally, we extensively evaluate the performance of PicPrivacy on realworld datasets. The results demonstrate that PicPrivacy not only achieves individual privacy-preserving but also can guarantee to create complete 3D models. Qinya Li, Zhenzhe Zheng 0001, Fan Wu 0006, Guihai Chen |
IWQoS | 1 |
| 2019 | A Strategy-Proof Model-Based Online Auction for Ad Reservation
Qinya Li, Fan Wu 0006, Guihai Chen |
PRICAI (1) | 1 |
| 2018 | ImgPricing: Everyone Can Earn Proper Rewards by Simply Taking PhotosabstractA high-quality and large-scale image collection is a fundamental demand in the 3D reconstruction. Crowdsourcing can help us collect lots of diversified images. However, it is not easy to attract people to do this task due to their self-interest. Moreover, the collected images are quality-varying. Those low-quality images may disturb the performance of reconstruction. To avoid low-quality images and lead participants to collect high-quality data, we take images quality into account when allocating rewards. The rewards of participants should be proportionable with their contribution. In this paper, we propose a pricing mechanism, called ImgPricing, to determine the reward of participants in 3D reconstruction system. We model the process of image collection as a cooperative game, and regard each participant's contribution and corresponding image quality as critical factors when allocating rewards. ImgPricing differs from traditional pricing schemes, such as Shapley value, as it introduces the image sequence as an indispensable factor. Finally, we implement our design on the Android platform and evaluate its performance. We use some metrics, such as computational efficiency, fairness and anti-interference, to evaluate ImgPricing and compare with other traditional schemes. Our analyses show ImgPricing is superior to others in terms of computational efficiency and fairness. Qinya Li, Fan Wu 0006, Guihai Chen |
IWQoS | 1 |