Shuxi Guo

dblp:323/8558 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 47% GPUs and heterogeneous computing · 32% Parallel and multicore computing · 22%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.922026
RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching · INFOCOM 2026
Efficient Inter-Operator Scheduling for Concurrent Recommendation Model Inference on GPU · IJCAI 2025
Parallel and multicore computing › parallelization strategies
fine-grained parallelism
1.012026
RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching · INFOCOM 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator
recommendation model inference
0.912025
Efficient Inter-Operator Scheduling for Concurrent Recommendation Model Inference on GPU · IJCAI 2025
Recommender systems
DLRM inference
0.312026
RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching · INFOCOM 2026
Parallel and multicore computing
parallel programming models
0.312025
Efficient Inter-Operator Scheduling for Concurrent Recommendation Model Inference on GPU · IJCAI 2025

Methods — techniques the papers use, named apart from their topics

incremental batching · 2.0fine-grained parallelism · 2.0asynchronous tensor management · 0.9
YearPublicationVenuePosition
2026 RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching
Siheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao, Jing Wang 0039
INFOCOM4
2025 Efficient Inter-Operator Scheduling for Concurrent Recommendation Model Inference on GPU
abstract
Deep learning-based recommendation systems are increasingly important in the industry. To meet strict SLA requirements, serving frameworks must efficiently handle concurrent queries. However, current serving systems fail to serve concurrent queries due to the following problems: (1) inefficient operator (op) scheduling due to the query-wise op launching mechanism, and (2) heavy contention caused by the mutable nature of recommendation model inference. This paper presents RecOS, a system designed to optimize concurrent recommendation model inference on GPUs. RecOS efficiently schedules ops from different queries by monitoring GPU workloads and assigning ops to the most suitable streams. This approach reduces contention and enhances inference efficiency by leveraging inter-op parallelism and op characteristics. To maintain correctness across multiple CUDA streams, RecOS introduces a unified asynchronous tensor management mechanism. Evaluations demonstrate that RecOS improves online service performance, reducing latency by up to 68%.
Shuxi Guo, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao, Jingyu Wang 0001
IJCAI1