Tianyu Qi

dblp:07/10151 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-5484-985XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hybrid Granularity Distribution Estimation for Few-Shot Learning: Statistics Transfer From Categories and Instances
abstract
Distribution estimation is a pivotal strategy in few-shot learning (FSL) to mitigate data scarcity by sampling from estimated distributions, utilizing statistical properties (mean and variance) transferred from related base categories. However, category-level estimation alone often fails to generate representative samples due to significant dissimilarities between base and novel categories, leading to suboptimal performance. To address this limitation, we propose Hybrid Granularity Distribution Estimation (HGDE), which integrates both coarse-grained category-level statistics and fine-grained instance-level statistics. By leveraging instance statistics from the nearest base samples, HGDE enhances the characterization of novel categories, capturing subtle features that category-level estimation overlooks. These statistics are fused through linear interpolation to form a robust distribution for novel categories, ensuring both diversity and representativeness in generated samples. Additionally, HGDE employs refined estimation techniques, such as weighted summation for mean calculation and principal component retention for covariance, to further improve accuracy. Empirical evaluations on four FSL benchmarks, including Mini-ImageNet, Tiered-ImageNet, CUB and CIFAR-FS, demonstrate that HGDE offers effective distribution estimation capabilities and leads to notable accuracy gains, with improvements of more than 1.8% in 1-shot tasks on CUB. These results highlight HGDE's ability to balance mean precision and variance diversity, making it a versatile and effective solution for FSL.
Shuo Wang 0008, Tianyu Qi, Yanbin Hao, Beier Zhu, Hanwang Zhang, Meng Wang 0001
IEEE Trans. Image Process.2
2026 Radiant: Efficient Timely Large-Scale Scene Analytics Based on Hierarchical Framework
abstract
With the advancement of computer vision, the recently emerged 3D Gaussian Splatting (3DGS) has increasingly become a popular scene analytics algorithm due to its outstanding performance. Existing cloud-based 3DGS architectures overlook the challenges in real-world environments when handling large-scale scene analysis. This exposes issues such as inefficiency, low security, lack of privacy, and limited scalability. In this paper, we propose Radiant, a hierarchical framework for large scene analytics in a heterogeneous cloud-edge-device system, which jointly considers high efficiency, privacy and security, and scalability. Via extensive empirical study, we find that it is crucial to partition the regions for each edge appropriately and allocate varying camera positions to each device for image collection and training. The core of Radiant is partitioning regions based on heterogeneous environment information and allocating workloads to each device accordingly. Furthermore, we provide a 3DGS model aggregation algorithm that enhances the quality and ensures the continuity of models' boundaries. Finally, we develop a testbed, and experiments demonstrate that Radiant improved reconstruction quality by up to 25.7% and reduced up to 79.6% end-to-end latency.
Haosong Peng, Tianyu Qi, Yufeng Zhan, Ren Jin, Hao Li 0075, Yalun Dai, Yuanqing Xia
IEEE Trans. Serv. Comput.2
2025 ScalaSSC: Scalable Stateful Serverless Computing for Stream Processing Applications
abstract
Serverless platforms are increasingly being used to process continuous streams of data. However, data processing units (expressed as serverless 'functions') in such platforms cannot maintain state internally and, instead, rely on remote storage. Stateful serverless is an emerging paradigm that seeks to reduce the latency associated with remote storage access by introducing state servers on worker nodes, i.e., where the serverless functions run. While existing stateful serverless frameworks are successful in reducing latency compared to their stateless counterparts, there are multiple challenges that still need to be addressed in order to meet the real-time data processing requirements of streaming applications. As a result, this paper proposes ScalaSSC, a novel framework for data processing request scheduling and operator state management tailored to stream processing applications within stateful serverless computing. ScalaSSC co-locates requests with the state they act upon to avoid cross-worker state access. To reduce state access contention among executors, ScalaSSC not only introduces the concept of state parallelism to serverless computing, but also processes requests acting upon the same state in the same batch. During batch execution, when multiple states are required, batch APIs are provided to access all states in a single operation. Experimental results show that ScalaSSC achieves up to 845 times higher throughput than another stateful serverless system while maintaining similar end-to-end latency. It also achieves a throughput 14 times higher than that of Amazon Lambda while maintaining end-to-end latency at the microsecond level, in contrast to Lambda's latency at the second level.
Tianyu Qi, Maria Rodriguez Read, Rajkumar Buyya
CCGrid1
2025 Sylva: Tailoring Personalized Adversarial Defense in Pre-trained Models via Collaborative Fine-tuning
abstract
The growing adoption of large pre-trained models in edge computing has made deploying model inference on mobile clients both practical and popular. These devices are inherently vulnerable to direct adversarial attacks, which pose a substantial threat to the robustness and security of deployed models. Federated adversarial training (FAT) has emerged as an effective solution to enhance model robustness while preserving client privacy. However, FAT frequently produces a generalized global model, which struggles to address the diverse and heterogeneous data distributions across clients, resulting in insufficiently personalized performance, while also encountering substantial communication challenges during the training process. In this paper, we propose Sylva, a personalized collaborative adversarial training framework designed to deliver customized defense models for each client through a two-phase process. In Phase 1, Sylva employs LoRA for local adversarial fine-tuning, enabling clients to personalize model robustness while drastically reducing communication costs by uploading only LoRA parameters during federated aggregation. In Phase 2, a game-based layer selection strategy is introduced to enhance accuracy on benign data, further refining the personalized model. This approach ensures that each client receives a tailored defense model that balances robustness and accuracy effectively. Extensive experiments on benchmark datasets demonstrate that Sylva can achieve up to 50× improvements in communication efficiency compared to state-of-the-art algorithms, while achieving up to 29.5% and 50.4% enhancements in adversarial robustness and benign accuracy, respectively.
Tianyu Qi, Lei Xue 0001, Yufeng Zhan, Xiaobo Ma 0001
CCS1
2025 Multi-Objective Partial Computation Offloading for Edge Intelligence with Heterogeneous Components
abstract
Edge intelligence, the fusion of edge computing and artificial intelligence (AI), drives the advancement of intelligent Internet of Things (IoT). Since AI applications are often data-and computation-intensive, resource-scarce edge devices need to migrate data to resource-rich edge servers through computation offloading to meet requirements such as energy efficiency and low latency. Existing studies often focus on CPU-based edge systems and neglect the impacts of other components, such as memory, on offloading. From a parallel processing perspective, this article establishes a system model and a multi-objective optimization model for edge intelligence systems with heterogeneous components, including diverse processors, memory, network, and applications, to minimize system energy consumption, total execution time, and the workload ratio of edge servers. A multi-objective optimization algorithm integrating archive initialization, hybrid perturbation, clustering, and modified simulated annealing is proposed and validated through experiments using real-world software and hardware. Results demonstrate that the proposed algorithm significantly outperforms comparative algorithms in terms of inverted generational distance, pure diversity, and run-time while revealing the influence of application characteristics on offloading performance.
Baoyu Xu, Yancheng Ruan, Tianyu Qi, Guobing Zou, Xiaoyang Kang 0001, Lihua Zhang 0002
ICPADS3
2025 Robin: An Efficient Hierarchical Federated Learning Framework via a Learning-Based Synchronization Scheme
abstract
Hierarchical federated learning (HFL) extends traditional federated learning by introducing a cloud-edge-device framework to enhance scalability. However, the challenge of determining when devices and edges should aggregate models remains unresolved, making the design of an effective synchronization scheme crucial. Additionally, the heterogeneity in computing and communication capabilities, coupled with non-independent and identically distributed ( non-IID) data distributions, makes synchronization particularly complex. In this paper, we proposeRobin, a learning-based synchronization scheme for HFL systems. By collecting data such as models' parameters, CPU usage, communication time,etc., we design a deep reinforcement learning-based approach to decide the frequencies of cloud aggregation and edge aggregation, respectively. The proposed scheme well considers device heterogeneity, non-IID data and device mobility, to maximize the training model accuracy while minimizing the energy overhead. Meanwhile, we prove the convergence ofRobin's synchronization scheme. And we build an HFL testbed and conduct the experiments with real data obtained from Raspberry Pi and Alibaba Cloud. Extensive experiments under various settings are conducted to confirm the effectiveness ofRobin, which can improve 31.2% in model accuracy while reducing energy consumption by 36.4%.
Tianyu Qi, Yufeng Zhan, Peng Li 0017, Yuanqing Xia
IEEE Trans. Cloud Comput.1
2025 Seeking a Hierarchical Prototype for Multimodal Gesture Recognition
abstract
Gesture recognition has drawn considerable attention from many researchers owing to its wide range of applications. Although significant progress has been made in this field, previous works always focus on how to distinguish between different gesture classes, ignoring the influence of inner-class divergence caused by gesture-irrelevant factors. Meanwhile, for multimodal gesture recognition, feature or score fusion in the final stage is a general choice to combine the information of different modalities. Consequently, the gesture-relevant features in different modalities may be redundant, whereas the complementarity of modalities is not exploited sufficiently. To handle these problems, we propose a hierarchical gesture prototype framework to highlight gesture-relevant features such as poses and motions in this article. This framework consists of a sample-level prototype and a modal-level prototype. The sample-level gesture prototype is established with the structure of a memory bank, which avoids the distraction of gesture-irrelevant factors in each sample, such as the illumination, background, and the performers' appearances. Then the modal-level prototype is obtained via a generative adversarial network (GAN)-based subnetwork, in which the modal-invariant features are extracted and pulled together. Meanwhile, the modal-specific attribute features are used to synthesize the feature of other modalities, and the circulation of modality information helps to leverage their complementarity. Extensive experiments on three widely used gesture datasets demonstrate that our method is effective to highlight gesture-relevant features and can outperform the state-of-the-art methods.
Yunan Li 0001, Tianyu Qi, Zhuoqi Ma, Dou Quan, Qiguang Miao
IEEE Trans. Neural Networks Learn. Syst.2
2024 Tomtit: Hierarchical Federated Fine-Tuning of Giant Models based on Autonomous Synchronization
abstract
With the quick evolution of giant models, the paradigm of pre-training models and then fine-tuning them for downstream tasks has become increasingly popular. The adapter has been recognized as an efficient fine-tuning technique and attracts much research attention. However, adapter-based fine-tuning still faces the challenge of lacking sufficient data. Federated fine-tuning has been recently proposed to fill this gap, but existing solutions suffer from a serious scalability issue, and they are inflexible in handling dynamic edge environments. In this paper, we propose Tomtit, a hierarchical federated fine-tuning system that can significantly accelerate fine-tuning and improve the energy efficiency of devices. Via extensive empirical study, we find that model synchronization schemes (i.e., when edge servers and devices should synchronize their models) play a critical role in federated fine-tuning. The core of Tomtit is a distributed design that allows each edge and device to have a unique synchronization scheme with respect to their heterogeneity in model structure, data distribution and computing capability. Furthermore, we provide a theoretical guarantee about the convergence of Tomtit. Finally, we develop a prototype of Tomtit and evaluate it on a testbed. Experimental results show that it can significantly outperform the state-of-the-art.
Tianyu Qi, Yufeng Zhan, Peng Li 0017, Yuanqing Xia
INFOCOM1
2023 Hwamei: A Learning-Based Synchronization Scheme for Hierarchical Federated Learning
abstract
Federated learning (FL) enables collaborative model training among distributed devices without data sharing, but existing FL suffers from poor scalability because of global model synchronization. To address this issue, hierarchical federated learning (HFL) has been recently proposed to let edge servers aggregate models of devices in proximity, while synchronizing via the cloud periodically. However, a critical open challenge about how to design a good synchronization scheme (when devices and edges should be synchronized) is still unsolved. Devices are heterogeneous in computing and communication capability, and their data could be non-IID. No existing work can well synchronize various roles (e.g., devices and edge) in HFL to guarantee high learning efficiency and accuracy. In this paper, we propose a learning-based synchronization scheme for HFL systems. By collecting data such as edge models, CPU usage, communication time, etc., we design a deep reinforcement learning-based approach to decide the frequencies of cloud aggregation and edge aggregation, respectively. The proposed scheme well considers device heterogeneity, non-IID data and device mobility, to maximize the training model accuracy while minimizing the energy overhead. We build an HFL testbed and conduct experiments using real data obtained from Raspberry Pi and Alibaba Cloud. Extensive experimental results have confirmed the effectiveness of Hwamei.
Tianyu Qi, Yufeng Zhan, Peng Li 0017, Jingcai Guo, Yuanqing Xia
ICDCS1