Cheng-Wei Ching

dblp:284/2434 · DBLP profile ↗
← Back
9ranked-venue papers
9as first author
6since 2021 · last 2026
0000-0001-6621-4907ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 100%
Computer networks
4 papers
Edge and fog computing · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 83% Processor architecture and microarchitecture · 17%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
2.432026
Totoro+: An Adaptive and Scalable Edge Federated Learning System · IEEE Trans. Parallel Distributed Syst. 2026
Totoro: A Scalable Federated Learning Engine for the Edge · EuroSys 2024
Dual-Objective Personalized Federated Service System With Partially-Labeled Data Over Wireless Networks · IEEE Trans. Serv. Comput. 2023
Machine learning › Efficient and distributed learning › federated learning
decentralized federated learning
1.822026
Totoro+: An Adaptive and Scalable Edge Federated Learning System · IEEE Trans. Parallel Distributed Syst. 2026
Totoro: A Scalable Federated Learning Engine for the Edge · EuroSys 2024
Edge and fog computing › distributed learning › federated learning
federated edge learning
1.822026
Totoro+: An Adaptive and Scalable Edge Federated Learning System · IEEE Trans. Parallel Distributed Syst. 2026
Totoro: A Scalable Federated Learning Engine for the Edge · EuroSys 2024
Edge and fog computing › edge analytics
stream processing at edge
0.912025
AgileDART: An Agile and Scalable Edge Stream Processing Engine · IEEE Trans. Mob. Comput. 2025
Distributed systems
peer-to-peer systems
0.832026
Totoro+: An Adaptive and Scalable Edge Federated Learning System · IEEE Trans. Parallel Distributed Syst. 2026
AgileDART: An Agile and Scalable Edge Stream Processing Engine · IEEE Trans. Mob. Comput. 2025
Totoro: A Scalable Federated Learning Engine for the Edge · EuroSys 2024
Edge and fog computing
edge intelligence
0.812024
Totoro: A Scalable Federated Learning Engine for the Edge · EuroSys 2024
Machine learning › Efficient and distributed learning › federated learning
personalized federated learning
0.712023
Dual-Objective Personalized Federated Service System With Partially-Labeled Data Over Wireless Networks · IEEE Trans. Serv. Comput. 2023
Distributed systems › peer-to-peer systems
distributed hash table
0.522026
Totoro+: An Adaptive and Scalable Edge Federated Learning System · IEEE Trans. Parallel Distributed Syst. 2026
Totoro: A Scalable Federated Learning Engine for the Edge · EuroSys 2024
Processor architecture and microarchitecture › dataflow architecture › dataflow machine
dynamic dataflow
0.312025
AgileDART: An Agile and Scalable Edge Stream Processing Engine · IEEE Trans. Mob. Comput. 2025

Methods — techniques the papers use, named apart from their topics

distributed hash table · 7.0publish/subscribe · 3.5game theory · 3.0multi-armed bandit · 2.3exploitation-exploration path planning · 2.3publish-subscribe · 1.8stream operator placement · 1.7bandit-based path planning · 1.7approximation algorithm · 1.3
YearPublicationVenuePosition
2026 Totoro+: An Adaptive and Scalable Edge Federated Learning System
abstract
Federated Learning (FL) is an emerging distributed machine learning (ML) technique that enables in-situ model training and inference on decentralized edge devices. We propose Totoro$^+$, a novel scalable FL system that enables massive FL applications to run simultaneously on edge networks. The key insight is to explore a distributed hash table (DHT)-based peer-to-peer (P2P) model to re-architect the centralized FL system design into a fully decentralized one. In contrast to previous studies where many FL applications shared one centralized parameter server, Totoro$^+$assigns a dedicated parameter server to each application. Any edge node can act as any application's coordinator, aggregator, client selector, worker (participant device), or any combination of the above, thereby radically improving scalability and adaptivity. Totoro$^+$introduces three innovations to realize its design: a locality-aware P2P multi-ring structure, a publish/subscribe-based forest abstraction, and a game-theoretic path planning model with a guarantee of an$\epsilon$-approximate Nash equilibrium. Real-world experiments on 500 Amazon EC2 servers show that Totoro$^+$scales gracefully with the number of FL applications and$N$edge nodes speeds up the total training time by$1.2\times -14.0\times$, achieves$\mathcal {O}(\log N)$hops for model dissemination and gradient aggregation with millions of nodes, and efficiently adapts to the practical edge networks and churns.
Cheng-Wei Ching, Xin Chen 0084, Taehwan Kim 0012, Jian-Jhih Kuo, Dilma Da Silva, Liting Hu
IEEE Trans. Parallel Distributed Syst.1
2025 AgileDART: An Agile and Scalable Edge Stream Processing Engine
abstract
Edge applications generate a large influx of sensor data on massive scales, and these massive data streams must be processed shortly to derive actionable intelligence. However, traditional data processing systems are not well-suited for these edge applications as they often do not scale well with a large number of concurrent stream queries, do not support low-latency processing under limited edge computing resources, and do not adapt to the level of heterogeneity and dynamicity commonly present in edge computing environments. As such, we present AgileDart, an agile and scalable edge stream processing engine that enables fast stream processing of many concurrently running low-latency edge applications' queries at scale in dynamic, heterogeneous edge environments. The novelty of our work lies in a dynamic dataflow abstraction that leverages distributed hash table-based peer-to-peer overlay networks to autonomously place, chain, and scale stream operators to reduce query latencies, adapt to workload variations, and recover from failures and a bandit-based path planning model that re-plans the data shuffling paths to adapt to unreliable and heterogeneous edge networks. We show that AgileDart outperforms Storm and EdgeWise on query latency and significantly improves scalability and adaptability when processing many real-world edge stream applications' queries.
Cheng-Wei Ching, Xin Chen 0084, Chaeeun Kim, Tongze Wang, Dong Chen 0025, Dilma Da Silva, Liting Hu
IEEE Trans. Mob. Comput.1
2024 Totoro: A Scalable Federated Learning Engine for the Edge
abstract
Federated Learning (FL) is an emerging distributed machine learning (ML) technique that enables in-situ model training and inference on decentralized edge devices. We propose Totoro, a novel scalable FL engine, that enables massive FL applications to run simultaneously on edge networks. The key insight is to explore a distributed hash table (DHT)-based peer-to-peer (P2P) model to re-architect the centralized FL system design into a fully decentralized one. In contrast to previous studies where many FL applications shared one centralized parameter server, Totoro assigns a dedicated parameter server to each individual application. Any edge node can act as any application's coordinator, aggregator, client selector, worker (participant device), or any combination of the above, thereby radically improving scalability and adaptivity. Totoro introduces three innovations to realize its design: a locality-aware P2P multi-ring structure, a publish/subscribe-based forest abstraction, and a bandit-based exploitation-exploration path planning model. Real-world experiments on 500 Amazon EC2 servers show that Totoro scales gracefully with the number of FL applications and N edge nodes, speeds up the total training time by 1.2 × -14.0×, achieves O (logN) hops for model dissemination and gradient aggregation with millions of nodes, and efficiently adapts to the practical edge networks and churns.
Cheng-Wei Ching, Xin Chen 0084, Taehwan Kim 0012, Bo Ji 0001, Qingyang Wang 0001, Dilma Da Silva, Liting Hu
EuroSys1
2023 Dual-Objective Personalized Federated Service System With Partially-Labeled Data Over Wireless Networks
abstract
Federated learning (FL) emerges to mitigate the privacy concerns in machine learning-based services and applications, and personalized federated learning (PFL) evolves to alleviate the issue of data heterogeneity. However, FL and PFL usually rest on two assumptions: the users' data is well-labeled, or the personalized goals align with sufficient local data. Unfortunately, the two assumptions may not hold in most cases, where data labeling is costly, or most users have no sufficient local data to satisfy their personalized needs. To this end, we first formulate the problem, DoLP, that studies the issue of insufficient and partially-labeled data on FL-based services. DoLP aims to maximize two service objectives: 1) personalized classification objective and 2) the personalized labeling objective for each user within the constraint of training time over wireless networks. Then, we propose a PFL-based service system DoFed-SPP to solve DoLP. The DoFed-SPP's novelty is two-fold. First, we devise an inference-based first-order approximation metric, similarity ratio, to identify the similarity between users' local data. Second, we design an approximation algorithm to determine the appropriate size and set of users for uploading in each round. Extensive experiments show DoFed-SPP outperforms the state-of-the-art in final accuracy and time-to-accuracy performance on CIFAR10/100 and DBPedia.
Cheng-Wei Ching, Jia-Ming Chang, Jian-Jhih Kuo, Chih-Yu Wang 0001
IEEE Trans. Serv. Comput.1
2021 Efficient Online Decentralized Learning Framework for Social Internet of Things
abstract
Online Decentralized Learning (ODL) is suitable for Internet-of-Things (IoT) devices since only parameter updates are exchanged with neighbors to avoid uploading private data to a central server and the training data is allowed to arrive at the devices sequentially. However, the current ODL frameworks cannot support the emerging Social IoT (SIoT) paradigm favorably since the SIoT devices exchange parameter updates with only trust-worthy neighbors based on specific social relations (e.g., parental object relation and ownership object relation). Conversely, sharing parameter updates with untrustworthy neighbors could speed up the training process but may violate social relations. Differential privacy (DP) is thus used to ensure data security while excessive devices engaging DP may downgrade the training performance. However, most research neglects the effect of neighbor selection for each device based on social networks, physical networks, and DP. Thus, in this paper, we innovate an ODL framework ODLF-PDP to allow only a part of devices to engage DP (i.e., partially DP) to improve training performance. Then, an algorithm BeTTa is proposed to build an adequate communication topology based on the interplay among the social networks, physical networks, and DP. Last, the experiment results manifest that ODLF-PDP saves more than 20% physical training time compared to the current frameworks via the benchmark of MNIST.
Cheng-Wei Ching, Hung-Sheng Huang, Chun-An Yang, Jian-Jhih Kuo, Ren-Hung Hwang
GLOBECOM1
2021 Efficient Communication Topology via Partially Differential Privacy for Decentralized Learning
abstract
Decentralized learning (DL) allows IoT devices to exchange local model updates with only their neighboring devices instead of sending their model updates to a central server for aggregation. However, current DL frameworks cannot support the emerging Social IoT(SIoT) paradigm since SIoT devices exchange model updates with only social neighbors based on specific social relations (e.g., ownership and parental relationships). Conversely, sharing model updates with non-social neighbors can improve training performance but may violate social relations. Differential privacy (DP) is thus engaged with DL to ensure data security, while excessive devices engaging DP may downgrade the training performance. However, most research neglects the effect of neighbor selection for each device based on social networks, physical networks, and DP. Therefore, in this paper, we explore the non-trivial relation among the above factors to present a DL framework, DeepPrivacy, and prove its convergence rate and DP. Then, we formulate a novel optimization problem, CoTOPO, to find an efficient communication topology1for model updates exchange among devices in DL, and propose an algorithm, AutoTag, for CoTOPO. Last, experiment results manifest that DeepPrivacy and AutoTag combined outperform the state of the art in terms of convergence rate and physical training time significantly on CIFAR10 and FMNIST.
Cheng-Wei Ching, Hung-Sheng Huang, Chun-An Yang, Yu-Chun Liu, Jian-Jhih Kuo
ICCCN1
2020 Model Partition Defense against GAN Attacks on Collaborative Learning via Mobile Edge Computing
abstract
With growing concerns about privacy issues of machine learning, collaborative learning (CL) is developed to offer on-device training. However, adversarial behaviors of model inversion (MI) are undermining privacy of training data. Specifically, adversaries act as ordinary participants in CL and reproduce private data of a class in training data by training generative adversarial networks (GAN) on the fly, unknowingly. To this end, we design a novel model partition defense, PAMPAS, over user devices and trustworthy edge server to resist GAN attack, and formulate a new optimization problem, TENSOR, to optimize training time. To address the challenges that come with PAMPAS, we propose an algorithm TESLA that yields the optimal solution. Experiment and simulation results manifest that PAMPAS effectively defend GAN attack and TESLA reduces training time by 50% compared with other solutions.
Cheng-Wei Ching, Tzu-Cheng Lin, Kung-Hao Chang, Chih-Chiung Yao, Jian-Jhih Kuo
GLOBECOM1
2020 Energy-Efficient Link Selection for Decentralized Learning via Smart Devices with Edge Computing
abstract
Data privacy preservation has drawn much attention in emerging machine learning applications. Decentralized learning is thus developed to guarantee data security and get rid of the involvement of parameter server to avoid transmission bottleneck. However, the previous research focuses on data compression and exchange rules of model parameters among smart devices but neglects the interplay between link cardinality and transmission power consumption. To jointly optimize these issues, in this paper, we first formulate a new optimization problem, named GreenDL, prove its hardness, and then propose an approximation algorithm termed CoTRAIN. Experiment and simulation results manifest that CoTRAIN reduces more than 20% power compared with traditional methods without sacrificing the convergence rate.
Cheng-Wei Ching, Chung-Kai Yang, Yu-Chun Liu, Chia-Wei Hsu, Jian-Jhih Kuo, Hung-Sheng Huang, Jen-Feng Lee
GLOBECOM1
2020 Optimal Device Selection for Federated Learning over Mobile Edge Networks
abstract
Data privacy preservation has drawn much attention with emerging machine learning applications. Federated Learning is thus developed to offer decentralized learning on user devices. However, it is difficult to jointly address multiple issues such as device selection, upload scheduling, and payment minimization. To jointly optimize the issues above, we first formulate a new optimization problem, named TRAIN, to minimize the training cost (including incentive payment and upload time) while ensuring the data requirement. We then prove the NP-hardness and propose a 3-approximation algorithm, named DETECT to obtain a near-optimal solution. Simulation results manifest that DETECT reduces the training cost by 50% compared with other traditional methods and achieves high accuracy and short convergence time.
Cheng-Wei Ching, Yu-Chun Liu, Chung-Kai Yang, Jian-Jhih Kuo, Feng-Ting Su
ICDCS1