Jinjing Ma

dblp:63/11145 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 50% Cloud and datacenter computing · 50%
Computer networks
1 paper
Edge and fog computing · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
task scheduling
1.012026
Empowering Dragonfly: A Lightweight and Scalable Distribution System for Large Models With High Concurrency · IEEE Trans. Netw. 2026
Edge and fog computing
mobile edge computing
0.912025
Chasing Common Knowledge: Joint Large Model Selection and Pulling in MEC With Parameter Sharing · IEEE Trans. Parallel Distributed Syst. 2025
Edge and fog computing
model selection
0.912025
Chasing Common Knowledge: Joint Large Model Selection and Pulling in MEC With Parameter Sharing · IEEE Trans. Parallel Distributed Syst. 2025
Machine learning › Efficient and distributed learning
federated and distributed training
0.312025
Chasing Common Knowledge: Joint Large Model Selection and Pulling in MEC With Parameter Sharing · IEEE Trans. Parallel Distributed Syst. 2025
Machine learning › Efficient and distributed learning
parameter sharing
0.312025
Chasing Common Knowledge: Joint Large Model Selection and Pulling in MEC With Parameter Sharing · IEEE Trans. Parallel Distributed Syst. 2025

Methods — techniques the papers use, named apart from their topics

passive bandwidth inference · 2.0graph representation learning · 2.0attention mechanism · 2.0active delay probing · 2.0randomized approximation algorithm · 1.7online learning · 1.7multi-armed bandit · 1.7integer linear programming · 1.7
YearPublicationVenuePosition
2026 Empowering Dragonfly: A Lightweight and Scalable Distribution System for Large Models With High Concurrency
abstract
Artificial Intelligence Generated Content (AIGC) models typically have hundreds of billions of parameters, and developers experience prohibitively long pulling time from a central model registry to their local environments. Peer-to-peer (P2P)-enabled model distribution that pulls models from local peers within a cluster instead of the central model registry is emerging as a promising technique to reduce model pulling time. Nonetheless, such model distribution systems have to handle bursty concurrent pulling tasks. This may occupy the network bandwidth of some peers for a long time, thereby making the peers unable to respond to further pulling tasks. We thus aim to design a lightweight and scalable model distribution system to balance the network resource usage of peers, by proposing learning-driven algorithms to accurately predict network status between peers and implementing the design in real production environments. Specifically, we first propose a lightweight network measurement mechanism that combines active delay probing and passive bandwidth inference with low resource overhead. We also propose a learning-driven task scheduling algorithm based on a structural graph representation with a varied-multi-hop attention mechanism, to predict bursty patterns of concurrent pulling tasks. We then design an asynchronous model training and inference method to enable seamless incremental learning based on the dynamic network status data. We finally implement our system design and the learning-driven algorithm in a Cloud Native Computing Foundation (CNCF) projectDragonflythat has already been publicly released since its version$v2.1.0$. Real experiments in the Ant Group’s production environment show that our system reduces the total completion time by at least 10% and increases the average bandwidth utilization of peers by 20%, compared with mainstream systems and algorithms.
Lizhen Zhou, Zichuan Xu, Wenbo Qi, Jinjing Ma, Haomiao Jiang, Qiufen Xia, Guowei Wu 0001
IEEE Trans. Netw.5
2025 Chasing Common Knowledge: Joint Large Model Selection and Pulling in MEC With Parameter Sharing
abstract
Pretrained Foundation Models (PFMs) are regarded as a promising accelerator for the development of various Artificial Intelligence (AI) applications, and have recently been widely fine-tuned to satisfy users' personalized inference demands. As many users are attracted to PFM-based AI applications, remote data centers are increasingly unable to solely bear the enormous computational demands and meet the delay requirements of inference requests. Mobile edge computing (MEC) offers a viable solution for delivering low-latency inference services by pulling fine-tuned PFMs from the remote data center to cloudlets in the proximity of users. However, a fine-tuned PFM typically comprises billions of model parameters, which are highly resource-intensive, time-consuming, and cost-prohibitive to execute at the edge. To address this, we investigate a novel joint large model selection and pulling problem in MEC networks. The novelty of our study lies in exploring parameter sharing among fine-tuned PFMs based on their common knowledge. Specifically, we first formulate a Non-Linear Integer Programming (NLIP) for the problem to minimize the total delay of implementing all inference requests. We then transform the NLIP into an equivalent Integer Linear Program (ILP) that is much simpler to solve. We further propose a randomized algorithm with a provable approximation ratio for the problem. We also consider the online version of the problem with uncertain request demand, and develop an online learning algorithm with a bounded regret. The crux of the online algorithm is the adoption of the multi-armed bandit technique with restricted context for dynamic admissions of inference requests. We finally conduct extensive experiments based on real datasets. Experimental results demonstrate that our algorithms reduce at least 38% in total delays and average costs, while achieving a 5% improvement in average accuracies.
Lizhen Zhou, Zichuan Xu, Qiufen Xia, Wenhao Ren, Wenbo Qi, Jinjing Ma
IEEE Trans. Parallel Distributed Syst.7
2012 Who Resemble You Better, Your Friends or Co-visited Users
Jinjing Ma, Yan Zhang 0004
APWeb1