Renyuan Liu

dblp:277/0490 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 74% Representation and self-supervised learning · 18% 3D vision · 5%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware accelerators and domain-specific architectures · 38% Cloud and datacenter computing · 28% Embedded and real-time systems · 18%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference serving
1.012026
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling · EuroSys 2026
Wearable and physiological sensing › motion sensing
IMU sensing
1.012026
Physical Self-Supervised Learning: IMU Sensing without Manual Labels · MobiSys 2026
Cloud and datacenter computing › inference serving
LLM serving
1.012026
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling · EuroSys 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.922024
ScaleFlow: Efficient Deep Vision Pipeline with Closed-Loop Scale-Adaptive Inference · ACM Multimedia 2023
DynaSpa: Exploiting Spatial Sparsity for Efficient Dynamic DNN Inference on Devices · SenSys 2024
Machine learning › Efficient and distributed learning › model compression › neural network compression
activation compression
0.912025
DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training · MobiSys 2025
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-device learning
0.912025
DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training · MobiSys 2025
Machine learning › Efficient and distributed learning
model compression
0.812024
DynaSpa: Exploiting Spatial Sparsity for Efficient Dynamic DNN Inference on Devices · SenSys 2024
Machine learning › Efficient and distributed learning
inference efficiency
0.712023
ScaleFlow: Efficient Deep Vision Pipeline with Closed-Loop Scale-Adaptive Inference · ACM Multimedia 2023
Embedded and real-time systems
on-device inference
0.712023
ScaleFlow: Efficient Deep Vision Pipeline with Closed-Loop Scale-Adaptive Inference · ACM Multimedia 2023
Computer vision › 3D vision › motion capture
inertial motion capture
0.312026
Physical Self-Supervised Learning: IMU Sensing without Manual Labels · MobiSys 2026
Memory systems
cache management
0.312026
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling · EuroSys 2026
Memory systems › cache management
KV cache management
0.312026
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling · EuroSys 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN accelerator
sparse DNN accelerator
0.212024
DynaSpa: Exploiting Spatial Sparsity for Efficient Dynamic DNN Inference on Devices · SenSys 2024
Computer vision › Image recognition and object detection
object detection
0.212023
ScaleFlow: Efficient Deep Vision Pipeline with Closed-Loop Scale-Adaptive Inference · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

uncertainty-aware learning · 2.0preemptive scheduling · 2.0physics decoder · 2.0kinematic tree · 2.0autoencoder · 2.0KV cache offloading · 2.0dynamic activation quantization · 1.7tensor compilation · 1.5spatial sparsity exploitation · 1.5scale-equivariant network · 1.3wavelet theory · 0.7
YearPublicationVenuePosition
2026 TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
abstract
Real-time LLM interactions demand streamed token generations, where text tokens are progressively generated and delivered to users while balancing two objectives: responsiveness (i.e., low time-to-first-token) and steady generation (i.e., required time-between-tokens). Standard LLM serving systems suffer from the inflexibility caused by non-preemptive request scheduling and reactive memory management, leading to poor resource utilization and low request processing parallelism under request bursts. Therefore, we present TokenFlow, a novel LLM serving system with enhanced text streaming performance via preemptive request scheduling and proactive key-value (KV) cache management. TokenFlow dynamically prioritizes requests based on real-time token buffer occupancy and token consumption rate, while actively transferring KV cache between GPU and CPU memory in the background and overlapping I/O with computation to minimize request preemption overhead. Extensive experiments on Llama3-8B and Qwen2.5-32B across multiple GPUs (RTX 4090, A6000, H200) demonstrate that TokenFlow achieves up to 82.5% higher effective throughput (accounting for actual user consumption) while reducing P99 TTFT by up to 80.2%, without degrading overall token throughput.
Chuheng Du, Renyuan Liu, Shuochao Yao, Dingtian Yan, Jiang Liao, Shengzhong Liu, Fan Wu 0006, Guihai Chen
EuroSys3
2026 Physical Self-Supervised Learning: IMU Sensing without Manual Labels
abstract
Deep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and self-supervised methods reduce but do not remove this dependence, still requiring labeled data for domain adaptation and largely ignoring known physical structure. We propose physical self-supervised learning, an autoencoder-style paradigm for label-free IMU sensing. We replace the conventional neural decoder with an auto-adaptive physics decoder—a learnable family of kinematic equations that enforces explicit physical structure while adapting across environments—and adopt a hybrid two-stage IMU encoder with reconstruction in a structured latent space to mitigate sensor noise. Our framework further introduces probabilistic frequency-spatial constraints to disentangle sensor and object motion, a multi-view kinematic tree to exploit sparse physical self-supervised signals, and an uncertainty-aware formulation to handle the inherent ambiguity of IMU inference. Evaluated on inertial tracking and full-body motion capture over public datasets and realistic deployments, physical self-supervised learning reduces errors by up to 5× for tracking and 4× for motion capture in challenging generalization scenarios, consistently outperforming state-of-the-art supervised and self-supervised baselines without any labels.
Yuyang Leng, Renyuan Liu, Shaohan Hu, Peijun Zhao, Chun-Fu Chen 0001, Songqing Chen, Shuochao Yao
MobiSys2
2026 Enhancing graph neural networks through universal self-knowledge distillation
Zheng Zhongzhu, Renyuan Liu, Jiangping Zhu
Neural Networks3
2025 Attention-Driven LPLC2 Neural Ensemble Model for Multi-Target Looming Detection and Localization
abstract
Lobula plate/lobula columnar, type 2 (LPLC2) visual projection neurons in the fly’s visual system possess highly looming-selective properties, making them ideal for developing artificial collision detection systems. The four dendritic branches of individual LPLC2 neurons, each tuned to specific directional motion, enhance the robustness of looming detection by utilizing radial motion opponency. Existing models of LPLC2 neurons either concentrate on individual cells to detect centroid-focused expansion or utilize population-voting strategies to obtain global collision information. However, their potential for addressing multi-target collision scenarios remains largely untapped. In this study, we propose a numerical model for LPLC2 populations, leveraging a bottom-up attention mechanism driven by motion-sensitive neural pathways to generate attention fields (AFs). This integration of AFs with highly nonlinear LPLC2 responses enables precise and continuous detection of multiple looming objects emanating from any region of the visual field. We began by conducting comparative experiments to evaluate the proposed model against two related models, highlighting its unique characteristics. Next, we tested its ability to detect multiple targets in dynamic natural scenarios. Finally, we validated the model using real-world video data collected by aerial robots. Experimental results demonstrate that the proposed model excels in detecting, distinguishing, and tracking multiple looming targets with remarkable speed and accuracy. This advanced ability to detect and localize looming objects, especially in complex and dynamic environments, holds great promise for overcoming collision-detection challenges in mobile intelligent machines.
Renyuan Liu, Qinbing Fu
IJCNN1
2025 DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training
abstract
Recent advancements in on-device training for deep neural networks have underscored the critical need for efficient activation compression to overcome the memory constraints of mobile and edge devices. As activations dominate memory usage during training and are essential for gradient computation, compressing them without compromising accuracy remains a key research challenge. While existing methods for dynamic activation quantization promise theoretical memory savings, their practical deployment is impeded by system-level challenges such as computational overhead and memory fragmentation.
Renyuan Liu, Yuyang Leng, Kaiyan Liu, Shaohan Hu, Chun-Fu Chen 0001, Peijun Zhao, Heechul Yun, Shuochao Yao
MobiSys1
2025 A neuromorphic binocular framework fusing directional and depth motion cues towards precise collision prediction
Chuankai Fang, Haoting Zhou, Renyuan Liu, Qinbing Fu
Neurocomputing3
2025 Disentangled feature graph for Hierarchical Text Classification
Renyuan Liu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou
Inf. Process. Manag.1
2025 Feature disentanglement, selection, and reaggregation method for multi-task learning
Renyuan Liu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou
Knowl. Inf. Syst.1
2024 DynaSpa: Exploiting Spatial Sparsity for Efficient Dynamic DNN Inference on Devices
abstract
Recent advancements in exploring machine learning models' dynamic spatial sparsity have demonstrated great potential for superior efficiency and adaptability without compromising accuracy when compared to conventional static-and-dense DNNs. However, realizing theoretical inference acceleration under practical deployment environments is still faced with significant system challenges. Current vendor libraries and tensor compilers fall short due to their extra data copy operations or insufficient computation schemes, especially for DNN operators with dynamic spatial sparsity.
Renyuan Liu, Yuyang Leng, Shilei Tian, Shaohan Hu, Chun-Fu Chen 0001, Shuochao Yao
SenSys1
2023 ScaleFlow: Efficient Deep Vision Pipeline with Closed-Loop Scale-Adaptive Inference
abstract
Deep visual data processing is underpinning many life-changing applications, such as auto-driving and smart cities. Improving the accuracy while minimizing their inference time under constrained resources has been the primary pursuit for their practical adoptions. Existing research thus has been devoted to either narrowing down the area of interest for the detection or miniaturizing the deep learning model for faster inference time. However, the former may risk missing/delaying small but important object detection, potentially leading to disastrous consequences (e.g., car accidents), while the latter often compromises the accuracy without fully utilizing intrinsic semantic information. To overcome these limitations, in this work, we propose ScaleFlow, a closed-loop scale-adaptive inference that can reduce model inference time by progressively processing vision data with increasing resolution but decreasing spatial size, achieving speedup without compromising accuracy. For this purpose, ScaleFlow refactors existing neural networks to be scale-equivariant on multiresolution data with the assistance of wavelet theory, providing predictable feature patterns on different data resolutions. Comprehensive experiments have been conducted to evaluate ScaleFlow. The results show that ScaleFlow can support anytime inference, consistently provide 1.5× to 2.2× speed up, and save around 25% ~ 45% energy consumption with < 1% accuracy loss on four embedded and edge platforms
Yuyang Leng, Renyuan Liu, Hongpeng Guo, Songqing Chen, Shuochao Yao
ACM Multimedia2
2023 An adversarial training-based mutual information constraint method
Renyuan Liu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou
Appl. Intell.1
2022 An associative memory circuit based on physical memristors
Mei Guo, Yongliang Zhu, Renyuan Liu, Kaixuan Zhao, Gang Dou
Neurocomputing3