Ruilong Ma

dblp:348/5123 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Video understanding and tracking · 32% Representation and self-supervised learning · 32% Efficient and distributed learning · 28%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 51% Hardware accelerators and domain-specific architectures · 49%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Edge and fog computing › edge inference
collaborative inference
0.912025
DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge Cluster · ACM Trans. Archit. Code Optim. 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge Cluster · ACM Trans. Archit. Code Optim. 2025
Machine learning › Representation and self-supervised learning › representation learning
spatio-temporal representation learning
0.812024
Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation Learning · CVPR 2024
Computer vision › Video understanding and tracking
video anomaly detection
0.812024
Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation Learning · CVPR 2024
Machine learning › Efficient and distributed learning
distributed inference
0.712023
Poster: PipeLLM: Pipeline LLM Inference on Heterogeneous Devices with Sequence Slicing · SIGCOMM 2023
Parallel and multicore computing › parallelization strategies
model parallelism
0.712023
Poster: PipeLLM: Pipeline LLM Inference on Heterogeneous Devices with Sequence Slicing · SIGCOMM 2023
Parallel and multicore computing › parallelization strategies
model and data parallelism
0.312025
DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge Cluster · ACM Trans. Archit. Code Optim. 2025
Natural language and speech › Language models and text generation › large language model
large language model deployment
0.212023
Poster: PipeLLM: Pipeline LLM Inference on Heterogeneous Devices with Sequence Slicing · SIGCOMM 2023

Methods — techniques the papers use, named apart from their topics

model partition · 1.7linear programming · 1.7adaptive workload partition · 1.7sequence slicing · 1.3pipeline parallelism · 1.3proxy task · 0.8missing frame estimation · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2026 SPIRNet: Polarized image reconstruction with adaptive sparse attention and intelligent scaling
Ruilong Ma, Youlin Gu, Fanhao Meng, Yihua Hu 0001
Knowl. Based Syst.1
2025 DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge Cluster
abstract
Parallelizing CNN inference on heterogeneous edge clusters with data parallelism has gained popularity as a way to meet real-time requirements without sacrificing model accuracy. However, existing algorithms struggle to find optimal parallel granularity for complex CNNS, the structure of which is a directed acyclic graph (DAG) rather than a chain, and the parallel dimension is inflexible. To distribute the workload of modern CNNs on heterogeneous devices is also proven as NP-hard problem. In this article, we introduce DeepZoning , a versatile and cooperative inference framework that combines both model and data parallelism to accelerate CNN inference. DeepZoning employs two algorithms at different levels: (1) a low-level Adaptive Workload Partition algorithm that uses linear programming and takes spatial and channel dimensions into optimization during the search for feature map distribution on heterogeneous devices, and (2) a high-level Model Partition algorithm that finds the optimal model granularity and organizes complex CNNs into sequential zones to balance communication and computation during execution. Our experimental evaluations show that DeepZoning is effective, achieving up to a 3.02× speed improvement on our experimental prototype compared to state-of-the-art algorithms.
Jingyu Wang 0001, Ruilong Ma, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Song Guo 0001
ACM Trans. Archit. Code Optim.2
2024 Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation Learning
abstract
Recent progress in video anomaly detection suggests that the features of appearance and motion play crucial roles in distinguishing abnormal patterns from normal ones. However, we note that the effect of spatial scales of anomalies is ignored. The fact that many abnormal events occur in limited localized regions and severe background noise in-terferes with the learning of anomalous changes. Mean-while, most existing methods are limited by coarse-grained modeling approaches, which are inadequate for learning highly discriminative features to discriminate subtle differences between small-scale anomalies and normal patterns. To this end, this paper address multi-scale video anomaly detection by multi-grained spatiotemporal representation learning. We utilize video continuity to design three proxy tasks to perform feature learning at both coarse-grained and fine-grained levels, i.e., continuity judgment, discontinuity localization, and missing frame estimation. In particular, we formulate missing frame estimation as a contrastive learning task in feature space instead of a reconstruction task in RGB space to learn highly discriminative features. Experiments show that our proposed method outperforms state-of-the-art methods on four datasets, especially in scenes with small-scale anomalies.
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Ruilong Ma, Jianxin Liao
CVPR7
2023 Poster: PipeLLM: Pipeline LLM Inference on Heterogeneous Devices with Sequence Slicing
abstract
Large Language Models (LLMs) has fostered the creation of innovative requirements. Locally deployed LLMs for micro-enterprise mitigates potential issues such as privacy infringements and sluggish response. However, they are hampered by the limitations in computing capability and memory space of possessed devices. We introduce PipeLLM, which allocates the model across devices commensurate with their computing capabilities. It enables the parallel execution of layers with slicing input sequence along the token dimension. PipeLLM demonstrates the potential to accelerate LLM inference with heterogeneity devices, offering a solution for LLM deployment in micro-enterprise hardware environment.
Ruilong Ma, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao
SIGCOMM1
2023 Brief Announcement: Accelerate CNN Inference with Zoning Graph at Dynamic Granularity
abstract
Partitioning a CNN and parallel executing inference with multiple IoT devices have gained popularity as a way to meet real-time requirements without sacrificing model accuracy. However, existing algorithms have struggled to find the optimal model partitioning granularity for complex CNNs. Additionally, executing inference with heterogeneous IoT devices is NP-hard when the structure of the CNN is a directed acyclic graph (DAG) rather than a chain. In this paper, we introduce a versatile and cooperative inference framework that combines both model and data parallelism to accelerate CNN inference. DeepZoning employs two algorithms at different levels: (1) a low-level Adaptive Workload Partition algorithm that uses linear programming and takes spatial and channel dimensions into optimization during the search for feature map distribution on heterogeneous devices, and (2) a high-level Model Partition algorithm that finds the optimal model granularity and organizes complex CNNs into sequential zones to balance communication and computation during execution.
Ruilong Ma, Qi Qi 0001, Jingyu Wang 0001, Zirui Zhuang, Jing Wang 0039
SPAA1