Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lyudong Jin

dblp:365/4374 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0009-0003-2454-6577ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
3 papers
Edge and fog computing · 82% Internet of things and sensor networks · 18%
Artificial intelligence
2 papers
Deep learning architectures and training · 82% Language models and text generation · 18%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
mixture of experts
1.222026
Poster: MoE2: Optimizing Collaborative Inference for Edge Large Language Models · MobiCom 2025
MoE2: Optimizing Collaborative Inference for Edge Large Language Models · IEEE Trans. Netw. 2026
Edge and fog computing › edge inference
collaborative inference
1.012026
MoE2: Optimizing Collaborative Inference for Edge Large Language Models · IEEE Trans. Netw. 2026
Edge and fog computing
edge inference
0.912025
Poster: MoE2: Optimizing Collaborative Inference for Edge Large Language Models · MobiCom 2025
Internet of things and sensor networks › age of information
age of information minimization
0.812024
Fractional Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing · AAAI 2024
Edge and fog computing
mobile edge computing
0.812024
Fractional Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing · AAAI 2024
Edge and fog computing › mobile edge computing
task offloading and scheduling
0.812024
Fractional Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing · AAAI 2024
Natural language and speech › Language models and text generation
large language model
0.312025
Poster: MoE2: Optimizing Collaborative Inference for Edge Large Language Models · MobiCom 2025

Methods — techniques the papers use, named apart from their topics

monotonic optimization · 2.0gating · 2.0two-level expert selection · 1.7joint gating · 1.7discrete monotonic optimization · 1.7hybrid action space · 0.8fractional reinforcement learning · 0.8deep reinforcement learning · 0.8
YearPublicationVenuePosition
2026 MoE2: Optimizing Collaborative Inference for Edge Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. Exploiting the heterogeneous capabilities of edge LLMs is crucial for diverse emerging applications, as it enables greater cost-effectiveness and reduced latency. In this work, we introduceMixture-of-Edge-Experts (MoE2), a novel collaborative inference framework for edge LLMs. We formulate a joint gating and expert selection problem to optimize inference performance under energy and latency constraints. Unlike conventional MoE problems, LLM expert selection becomes significantly more challenging due to the combinatorial nature and the heterogeneity of edge LLMs across various attributes. To this end, we propose a two-level expert selection mechanism through which we uncover an optimality-preserving property of gating parameters across expert selections. This property enables the decomposition of the training and selection processes, significantly reducing complexity. Furthermore, we leverage the objective’s monotonicity and design a discrete monotonic optimization algorithm for optimal expert selection. We implement edge servers with NVIDIA Jetson AGX Orins and NVIDIA RTX 4090 GPUs, and perform extensive experiments. Our results validate the performance improvements for various LLM models and show that our MoE2 method can achieve optimal trade-offs among different delay and energy budgets, and outperforms baselines under various system resource constraints. We further demonstrate its strong robustness in dynamic, non-stationary environments and its effectiveness in achieving load balancing.
Lyudong Jin, Shurong Wang, Howard H. Yang, Jian Wu 0001, Meng Zhang 0013
IEEE Trans. Netw.1
2025 Poster: MoE2: Optimizing Collaborative Inference for Edge Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks. Exploiting the heterogeneous capabilities of edge LLMs is crucial for emerging applications, enabling greater cost-effectiveness and reduced latency. In this work, we introduce Mixture-of-Edge-Experts (MoE2), a collaborative inference framework for edge LLMs. We formulate the joint gating and expert selection problem to optimize inference under energy and latency constraints. Unlike conventional MoE problems, expert selection here is more challenging due to the combinatorial nature and heterogeneity of edge LLMs. To address this, we propose a two-level expert selection mechanism and uncover an optimality-preserving property of gating parameters that decouples training and selection, reducing complexity. We further leverage the objective's monotonicity and design a discrete monotonic optimization algorithm. Implemented on Jetson Orin and RTX 4090 platforms, MoE2 achieves optimal trade-offs across delay and energy budgets, outperforming baselines under various resource constraints.
Lyudong Jin, Shurong Wang, Howard H. Yang, Jian Wu 0001, Meng Zhang 0013
MobiCom1
2024 Fractional Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing
abstract
Mobile edge computing (MEC) is a promising paradigm for real-time applications with intensive computational needs (e.g., autonomous driving), as it can reduce the processing delay. In this work, we focus on the timeliness of computational-intensive updates, measured by Age-of-Information (AoI), and study how to jointly optimize the task updating and offloading policies for AoI with fractional form. Specifically, we consider edge load dynamics and formulate a task scheduling problem to minimize the expected time-average AoI. The uncertain edge load dynamics, the nature of the fractional objective, and hybrid continuous-discrete action space (due to the joint optimization) make this problem challenging and existing approaches not directly applicable. To this end, we propose a fractional reinforcement learning (RL) framework and prove its convergence. We further design a model-free fractional deep RL (DRL) algorithm, where each device makes scheduling decisions with the hybrid action space without knowing the system dynamics and decisions of other devices. Experimental results show that our proposed algorithms reduce the average AoI by up to 57.6% compared with several non-fractional benchmarks.
Lyudong Jin, Ming Tang 0006, Meng Zhang 0013, Hao Wang 0016
AAAI1