Jianxiong Liao

dblp:65/7868 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 64% Parallel and multicore computing · 18% GPUs and heterogeneous computing · 18%
Computer networks
1 paper
Edge and fog computing · 100%
Software engineering, system software, and programming languages
2 papers
Services computing and microservices · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
1.922026
Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving · PPoPP 2026
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism · SC 2025
Edge and fog computing
edge-native application
1.012026
MicroEdge: An Online Optimization Framework for Cost-Efficient Microservice Orchestration in Edge Native Applications · IEEE Trans. Mob. Comput. 2026
Edge and fog computing › service orchestration
microservice orchestration
1.012026
MicroEdge: An Online Optimization Framework for Cost-Efficient Microservice Orchestration in Edge Native Applications · IEEE Trans. Mob. Comput. 2026
Cloud and datacenter computing › inference serving
LLM inference scheduling
1.012026
Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving · PPoPP 2026
Cloud and datacenter computing › quality of service
SLO-aware scheduling
1.012026
Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving · PPoPP 2026
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.912025
Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared Clusters · ASPLOS (2) 2025
GPUs and heterogeneous computing › GPU programming
dynamic parallelism
0.912025
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism · SC 2025
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous GPU cluster
0.912025
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism · SC 2025
Cloud and datacenter computing › inference serving
LLM serving
0.912025
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism · SC 2025
Parallel and multicore computing
load balancing
0.912025
Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared Clusters · ASPLOS (2) 2025
Parallel and multicore computing › parallelization strategies
model parallelism
0.912025
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism · SC 2025
Machine learning and data management
data management for machine learning
0.312026
Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving · PPoPP 2026
Machine learning and data management › inference serving
LLM serving
0.312026
Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving · PPoPP 2026
Cloud and datacenter computing
resource management
0.312026
MicroEdge: An Online Optimization Framework for Cost-Efficient Microservice Orchestration in Edge Native Applications · IEEE Trans. Mob. Comput. 2026

Methods — techniques the papers use, named apart from their topics

regularization · 3.0online optimization · 3.0dependent rounding · 3.0approximation algorithm · 3.0layer-level scheduling · 2.0dynamic load shifting · 1.7fine-grained parallelism · 0.9dynamic parallelization · 0.9
YearPublicationVenuePosition
2026 Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving
abstract
Engaging applications with diverse SLO requirements has become indispensable for production-scale LLM serving systems. However, existing systems rely on iteration-level scheduling, which enforces inflexible, unified execution across multi-SLO workloads, significantly constraining the serving efficiency.
Jianxiong Liao, Quanxing Dong, Yunkai Liang, Zhi Zhou 0006, Xu Chen 0004
PPoPP1
2026 Duba: Cost-Efficient Serverless Cloud-Edge Collaborative Machine Learning Serving with Dual-Batching
Jianxiong Liao, Zhi Zhou 0006, Fei Xu 0009
J. Comput. Sci. Technol.1
2026 MicroEdge: An Online Optimization Framework for Cost-Efficient Microservice Orchestration in Edge Native Applications
abstract
The rapid proliferation of edge computing infrastructure has significantly accelerated the adoption of edge-native applications, ranging from autonomous vehicles to augmented reality and real-time analytics. Microservice, renowned for its lightweight, loosely coupled, and modular architecture, has emerged as the de-facto standard for developing edge native applications. However, the resource scarcity and heterogeneity, coupled with request dynamics in edge environments, pose substantial challenges for effective microservice orchestration. To address these challenges, we propose MicroEdge, an online optimization framework designed for cost-efficient microservice orchestration in edge-native environments. MicroEdge employs a multi-level optimization approach by strategically coordinating four key dimensions: microservice placement, layer placement, layer pulling, and user request scheduling. The framework confronts two fundamental challenges in solving this joint optimization problem: (1) the time-coupled nature of long-term holistic cost minimization, and (2) the NP-hardness of the underlying problem. MicroEdge tackles these dual challenges by integrating a regularization method for online algorithm design and a dependent rounding technique for approximation algorithm design. Both rigorous theoretical analysis and extensive simulations driven by realistic Alibaba microservice workload traces validate the efficacy of MicroEdge.
Weihan Zeng, Kongyange Zhao, Jianxiong Liao, Zhi Zhou 0006, Deke Guo, Xu Chen 0004
IEEE Trans. Mob. Comput.3
2025 Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared Clusters
abstract
In a unified resource scheduling architecture, containers within the same microservice often encounter temporal and spatial performance imbalance when deployed in large-scale shared clusters. As a result, the commonly employed load-balancing approach often leads to substantial resource wastage as applications are frequently over-provisioned to meet service level agreements (SLAs).
Shutian Luo, Jianxiong Liao, Chenyu Lin, Huanle Xu, Zhi Zhou 0006, Cheng-Zhong Xu 0001
ASPLOS (2)2
2025 Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
abstract
The significant resource demands in LLM serving prompts production clusters to fully utilize heterogeneous hardware by partitioning LLM models across a mix of high-end and low-end GPUs. However, existing parallelization approaches often struggle to scale efficiently in heterogeneous environments due to their coarse-grained and static parallelization strategies.
Zizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou 0006, Cheng-Zhong Xu 0001
SC2
2024 Prediction of the transient emission characteristics from diesel engine using temporal convolutional networks
Jianxiong Liao, Zhizhou Cai, Hanming Wu, Maoxuan Wang
Eng. Appl. Artif. Intell.1