Tongtong Liu 0006

dblp:185/7074-6 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0003-1651-7784ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 30% Distributed systems · 24% GPUs and heterogeneous computing · 15%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer networks
2 papers
Edge and fog computing · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
1.012026
FedMHO: Heterogeneous One-Shot Federated Learning Towards Resource-Constrained Clients · WWW 2026
Machine learning › Efficient and distributed learning › federated learning › heterogeneous federated learning
model-heterogeneous federated learning
1.012026
FedMHO: Heterogeneous One-Shot Federated Learning Towards Resource-Constrained Clients · WWW 2026
Machine learning › Efficient and distributed learning › federated learning › communication-efficient federated learning
one-shot federated learning
1.012026
FedMHO: Heterogeneous One-Shot Federated Learning Towards Resource-Constrained Clients · WWW 2026
Edge and fog computing › edge inference
collaborative inference
0.912025
SPViT: Accelerate Vision Transformer Inference on Mobile Devices via Adaptive Splitting and Offloading · IEEE Trans. Mob. Comput. 2025
Distributed systems › distributed interactive applications
collaborative computing
0.912025
ApSpGEMM: Accelerating Large-scale SpGEMM with Heterogeneous Collaboration and Adaptive Panel · ACM Trans. Archit. Code Optim. 2025
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing
0.912025
ApSpGEMM: Accelerating Large-scale SpGEMM with Heterogeneous Collaboration and Adaptive Panel · ACM Trans. Archit. Code Optim. 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
SPViT: Accelerate Vision Transformer Inference on Mobile Devices via Adaptive Splitting and Offloading · IEEE Trans. Mob. Comput. 2025
Embedded and real-time systems › on-device inference
mobile inference
0.912025
SPViT: Accelerate Vision Transformer Inference on Mobile Devices via Adaptive Splitting and Offloading · IEEE Trans. Mob. Comput. 2025
High-performance computing › sparse linear algebra
sparse matrix multiplication
0.912025
ApSpGEMM: Accelerating Large-scale SpGEMM with Heterogeneous Collaboration and Adaptive Panel · ACM Trans. Archit. Code Optim. 2025
High-performance computing › sparse linear algebra › sparse matrix multiplication
SpGEMM
0.912025
ApSpGEMM: Accelerating Large-scale SpGEMM with Heterogeneous Collaboration and Adaptive Panel · ACM Trans. Archit. Code Optim. 2025
Distributed systems
distributed coordination and fault tolerance
0.312025
SPViT: Accelerate Vision Transformer Inference on Mobile Devices via Adaptive Splitting and Offloading · IEEE Trans. Mob. Comput. 2025
Distributed systems › distributed machine learning
model partitioning
0.312025
SPViT: Accelerate Vision Transformer Inference on Mobile Devices via Adaptive Splitting and Offloading · IEEE Trans. Mob. Comput. 2025

Methods — techniques the papers use, named apart from their topics

knowledge fusion · 2.0knowledge distillation · 2.0generative model · 2.0offline/online optimization · 1.7auto-regression latency prediction · 1.7adaptive splitting · 1.7reordering and splitting algorithms · 0.9asynchronous data transmission · 0.9adaptive panel allocation · 0.9
YearPublicationVenuePosition
2026 FedMHO: Heterogeneous One-Shot Federated Learning Towards Resource-Constrained Clients
abstract
Federated Learning (FL) is increasingly adopted in edge computing scenarios, where a large number of heterogeneous clients operate under constrained or sufficient resources. The iterative training process of FL incurs considerable computation and communication overhead, which is unfriendly for resource-constrained devices. One-shot FL is a promising approach to addressing communication issues inherent in conventional FL, and model-heterogeneous FL solves the problem of diverse computing resources across clients. However, existing methods face challenges in effectively managing model-heterogeneous one-shot FL, often leading to unsatisfactory global model performance or reliance on auxiliary datasets. To address these challenges, we propose a novel FL framework named FedMHO, which leverages deep classification models on resource-sufficient clients and lightweight generative models on resource-constrained devices. On the server side, FedMHO involves a two-stage process that includes data generation and knowledge fusion. Furthermore, we introduce FedMHO-MD and FedMHO-SD to mitigate the knowledge-forgetting problem during the knowledge fusion stage, and an unsupervised data optimization solution to improve the quality of synthetic samples. Comprehensive experiments demonstrate the effectiveness of our methods, as they outperform state-of-the-art baselines in various experimental setups.
Dezhong Yao 0002, Tongtong Liu 0006, Yuexin Shi, Zhiqiang Xu 0003
WWW2
2025 ApSpGEMM: Accelerating Large-scale SpGEMM with Heterogeneous Collaboration and Adaptive Panel
abstract
The Sparse General Matrix-Matrix multiplication (SpGEMM) is a fundamental component for many applications, such as algebraic multigrid methods (AMG), graphic processing, and deep learning. However, the unbearable latency of computing high-dimensional, large-scale sparse matrix multiplication on GPUs hinders the development of these applications. An effective approach is heterogeneous cores collaborative computing, but this method must address three aspects: (1) irregular non-zero elements lead to load imbalance and irregular memory access, (2) different core computing latency differences reduce computational parallelism, and (3) temporary data transfer between different cores introduces additional latency overhead. In this work, we propose an innovative framework for collaborative large-scale sparse matrix multiplication on CPU-GPU heterogeneous cores, named ApSpGEMM. ApSpGEMM is based on sparsity rules and proposes reordering and splitting algorithms to eliminate the impact of non-zero element distribution features on load and memory access. Then adaptive panels allocation with affinity constraints among cores improves computational parallelism. Finally, carefully arranged asynchronous data transmission and computation balance communication overhead. Compared with state-of-the-art SpGEMM methods, our approach provides excellent absolute performance on matrices with different sparse structures. On heterogeneous cores, the GFlops of large-scale sparse matrix multiplication is improved by 2.25 to 7.21 times.
Dezhong Yao 0002, Sifan Zhao, Tongtong Liu 0006, Hai Jin 0001
ACM Trans. Archit. Code Optim.3
2025 SPViT: Accelerate Vision Transformer Inference on Mobile Devices via Adaptive Splitting and Offloading
abstract
The Vision Transformer (ViT), which benefits from utilizing self-attention mechanisms, has demonstrated superior accuracy compared to CNNs. However, due to the expensive computational costs, deploying and inferring ViTs on resource-constrained mobile devices has become a challenge. To resolve this challenge, we conducted an empirical analysis to identify performance bottlenecks in deploying ViTs on mobile devices and explored viable solutions. In this paper, we propose SPViT, an adaptive split and offloading method that accelerates ViT inference on mobile devices. SPViT executes collaborative inference of ViT across available edge devices. We introduce a fine-grained splitting technique for the vision transformer structure. Furthermore, we propose an algorithm based on the Auto Regression model to predict partition latency and adaptive offload partitions. Finally, we design offline and online optimization methods to minimize the computational and communication overhead on each device. Based on real-world prototype experiments, SPViT effectively reduces inference latency by 2.2x to 3.3x across four state-of-the-art models.
Sifan Zhao, Tongtong Liu 0006, Hai Jin 0001, Dezhong Yao 0002
IEEE Trans. Mob. Comput.2
2024 Rethinking Personalized Federated Learning from Knowledge Perspective
abstract
Personalized federated learning (PFL) is a variant of federation learning, which improves the model performance of each participant through collaborative training while providing a customized local model that meets their unique needs. Balancing global (general) knowledge and local (personalized) knowledge is a core challenge in PFL. Existing PFL methods neglect knowledge forgetting during aggregation and updating, impacting the fusion of both types of knowledge. We empirically confirm the existence of knowledge forgetting, which leads to performance degradation during FL. This observation motivates us to rethink PFL from a knowledge perspective. Therefore, we propose Personalized Federated Learning Framework with Adaptive Model Fusion (pFedAMF). We consider the global model as global knowledge and history state models as local knowledge, and preserve them before forgetting occurs. After local training, we use an adaptive knowledge matrix to fuse knowledge from local, global, and history state models. This fused knowledge is then distilled into an individual model, thus transferring the fused knowledge into global models. Extensive experiments show that pFedAMF consistently outperforms FedAvg, achieving up to 5.22% average accuracy boost, and reducing computation cost by 82% and communication cost by 87%.
Dezhong Yao 0002, Ziquan Zhu, Tongtong Liu 0006, Zhiqiang Xu 0003, Hai Jin 0001
ICPP3
2023 FedRKG: A Privacy-Preserving Federated Recommendation Framework via Knowledge Graph Enhancement
Dezhong Yao 0002, Tongtong Liu 0006, Qi Cao 0002, Hai Jin 0001
GPC (2)2