Liang Mi

dblp:159/8058 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 StreamDuet: Bandwidth Efficient Multi-Drone Video Analytics with Iterative Streaming
Haihan Zhang, Weijun Wang 0001, Haipeng Dai 0001, Ruiben Zhou, Liang Mi, Guihai Chen
IWQoS5
2026 Efficient Remote KV Cache Reuse with GPU-native Video Codec
Liang Mi, Weijun Wang 0001, Jinghan Chen, Ting Cao 0003, Haipeng Dai 0001, Yunxin Liu 0001
SIGCOMM1
2026 CompViT: Real-Time Compressed Video Action Recognition with Asymmetric Transformer Networks
Tao Wu 0020, Shaowei Cen, Liang Mi, Weijun Wang 0001, Haipeng Dai 0001, Limin Wang 0002
Int. J. Comput. Vis.3
2026 Bi-Level Bandwidth Coordination for Multiple Video Inference at the Edge
abstract
High-definition (HD) cameras for surveillance and road traffic have experienced tremendous growth, demanding intensive computation resources for real-time analytics. Recently, offloading frames from the front-end device to the back-end edge server has shown great promise. In multi-stream competitive environments, efficient bandwidth management and proper scheduling are crucial to ensure both high inference accuracy and high throughput. To achieve this goal, we propose BiSwift, a bi-level framework that scales the concurrent real-time video analytics by a novel adaptive hybrid codec integrated with multi-level pipelines, and a global bandwidth controller for multiple video streams. The lower-level front-back-end collaborative mechanism (called adaptive hybrid codec) locally optimizes the accuracy and accelerates end-to-end video analytics for a single stream. The upper-level scheduler aims to accuracy fairness among multiple streams via the global bandwidth controller. The evaluation of BiSwift shows that BiSwift is able to real-time object detection on 9 streams with an edge device only equipped with an NVIDIA RTX3070 (8G) GPU. BiSwift improves 10%~21% accuracy and presents$1.2\sim 9\times $throughput compared with the state-of-the-art video analytics pipelines.
Haipeng Dai 0001, Jinghan Chen, Liang Mi, Weijun Wang 0001, Yuanchun Li 0003, Tingting Yuan 0001, Yuben Qu, Yunxin Liu 0001, Xiaoming Fu 0001, Guihai Chen
IEEE Trans. Netw.3
2025 Empower Vision Applications with LoRA LMM
abstract
Large Multimodal Models (LMMs) have shown significant progress in various complex vision tasks with the solid linguistic and reasoning capacity inherited from large language models (LMMs). Low-rank adaptation (LoRA) offers a promising method to integrate external knowledge into LMMs, compensating for their limitations on domain-specific tasks. However, the existing LoRA model serving is excessively computationally expensive and causes extremely high latency. In this paper, we present an end-to-end solution that empowers diverse vision tasks and enriches vision applications with LoRA LMMs. Our system, VaLoRA, enables accurate and efficient vision tasks by 1) an accuracy-aware LoRA adapter generation approach that generates LoRA adapters rich in domain-specific knowledge to meet application-specific accuracy requirements, 2) an adaptive-tiling LoRA adapters batching operator that efficiently computes concurrent heterogeneous LoRA adapters, and 3) a flexible LoRA adapter orchestration mechanism that manages application requests and LoRA adapters to achieve the lowest average response latency. We prototype VaLoRA on five popular vision tasks on three LMMs. Experiment results reveal that VaLoRA improves 24-62% of the accuracy compared to the original LMMs and reduces 20-89% of the latency compared to the state-of-the-art LoRA model serving systems.
Liang Mi, Weijun Wang 0001, Wenming Tu, Qingfeng He, Xinyu Fang, Yazhu Dong, Yuanchun Li 0003, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Yunxin Liu 0001
EuroSys1
2025 Demo: EdgeMind-OS: A Plug-and-Play Embodied Intelligence System for Real-Time On-Device Deployment
abstract
Building an always-on, contextual AI assistant that proactively supports humans remains a central goal in Embodied AI—yet cloud-based pipelines struggle to meet due to delay, bandwidth, and privacy constraints. This demo presents EdgeMind-OS, a fully on-device intelligence system designed for embodied agents operating in real-world scenarios. Edge-Mind-OS features a hierarchical architecture combining a real-time StreamBrain, modular skill experts, and a dynamic scene-episode memory. Achieving up to 7.3× faster local processing, it enables low-latency, privacy-preserving, and plug-and-play deployment across tasks such as semantic navigation, spatial memory recall, and multimodal interaction. We demonstrate how EdgeMind-OS empowers a mobile robot with only basic locomotion capabilities to perform realtime, free-form user-robot interaction through autonomous perception, reasoning and action —without reliance on external cloud infrastructure.
Jianyu Wei, Fucheng Jia, Liang Mi, Ruofei Ju, Xianye Wang, Yikai Zheng, Weijun Wang 0001, Shiqi Jiang 0002, Yunxin Liu 0001, Ting Cao 0003
MobiCom4
2025 Region-based Content Enhancement for Efficient Video Analytics at the Edge
Weijun Wang 0001, Liang Mi, Shaowei Cen, Haipeng Dai 0001, Yuanchun Li 0003, Xiaoming Fu 0001, Yunxin Liu 0001
NSDI2
2024 BiSwift: Bandwidth Orchestrator for Multi-Stream Video Analytics on Edge
abstract
High-definition (HD) cameras for surveillance and road traffic have experienced tremendous growth, demanding intensive computation resources for real-time analytics. Recently, offloading frames from the front-end device to the back-end edge server has shown great promise. In multi-stream competitive environments, efficient bandwidth management and proper scheduling are crucial to ensure both high inference accuracy and high throughput. To achieve this goal, we propose BiSwift, a bi-level framework that scales the concurrent real-time video analytics by a novel adaptive hybrid codec integrated with multi-level pipelines, and a global bandwidth controller for multiple video streams. The lower-level front-back-end collaborative mechanism (called adaptive hybrid codec) locally optimizes the accuracy and accelerates end-to-end video analytics for a single stream. The upper-level scheduler aims to accuracy fairness among multiple streams via the global bandwidth controller. The evaluation of BiSwift shows that BiSwift is able to real-time object detection on 9 streams with an edge device only equipped with an NVIDIA RTX3070 (8G) GPU. BiSwift improves 10%∼21% accuracy and presents 1.2∼ 9× throughput compared with the state-of-the-art video analytics pipelines.
Weijun Wang 0001, Tingting Yuan 0001, Liang Mi, Haipeng Dai 0001, Yunxin Liu 0001, Xiaoming Fu 0001
INFOCOM4
2024 Accelerated Neural Enhancement for Video Analytics With Video Quality Adaptation
abstract
The quality of the video stream is the key to neural network-based video analytics. However, low-quality video is inevitably collected by existing surveillance systems because of poor-quality cameras or over-compressed/pruned video streaming protocols, e.g., as a result of upstream bandwidth limit. To address this issue, existing studies use quality enhancers (e.g., neural super-resolution) to improve the quality of videos (e.g., resolution) and eventually ensure inference accuracy. Nevertheless, directly applying quality enhancers does not work in practice because it will introduce unacceptable latency. In this paper, we present AccDecoder, a novel accelerated decoder for real-time and neural-enhanced video analytics, selects a few frames adaptively via Deep Reinforcement Learning (DRL) to enhance the quality and inference then reuse on the unselected ones. Next, we extend AccDecoder to AccDecoder$+$by formulating the resolution-involved Markov decision process (MDP) to achieve resolution adaptation; it aims to trade accuracy and latency corresponding under various video resolutions. Proved by experiments, AccDecoder provides efficient inference capability via filtering important frames using DRL for DNN-based inference and reusing the results for the other frames via extracting the reference relationship among frames and blocks, which contributes 6-21% accuracy improvement and a latency reduction of 20-80% than baselines. Compared with AccDecoder, AccDecoder$+$achieves an additional 2-7% accuracy improvement.
Liang Mi, Tingting Yuan 0001, Weijun Wang 0001, Haipeng Dai 0001, Jiaqi Zheng 0001, Guihai Chen, Xiaoming Fu 0001
IEEE/ACM Trans. Netw.1
2023 AccDecoder: Accelerated Decoding for Neural-enhanced Video Analytics
abstract
The quality of the video stream is key to neural network-based video analytics. However, low-quality video is inevitably collected by existing surveillance systems because of poor quality cameras or over-compressed/pruned video streaming protocols, e.g., as a result of upstream bandwidth limit. To address this issue, existing studies use quality enhancers (e.g., neural super-resolution) to improve the quality of videos (e.g., resolution) and eventually ensure inference accuracy. Nevertheless, directly applying quality enhancers does not work in practice because it will introduce unacceptable latency. In this paper, we present AccDecoder, a novel accelerated decoder for real-time and neural-enhanced video analytics. AccDecoder can select a few frames adaptively via Deep Reinforcement Learning (DRL) to enhance the quality by neural super-resolution and then up-scale the unselected frames that reference them, which leads to 6-21% accuracy improvement. AccDecoder provides efficient inference capability via filtering important frames using DRL for DNN-based inference and reusing the results for the other frames via extracting the reference relationship among frames and blocks, which results in a latency reduction of 20-80% than baselines.
Tingting Yuan 0001, Liang Mi, Weijun Wang 0001, Haipeng Dai 0001, Xiaoming Fu 0001
INFOCOM2
2023 A family of pairwise multi-marginal optimal transports that define a generalized metric
Liang Mi, Azadeh Sheikholeslami, José Bento 0001
Mach. Learn.1
2021 Peak temperature analysis and optimization for pipelined hard real-time systems
Long Cheng 0007, Kai Huang 0001, Liang Mi, Gang Chen 0023, Alois C. Knoll
Inf. Sci.3
2020 Regularized Wasserstein Means for Aligning Distributional Data
abstract
We propose to align distributional data from the perspective of Wasserstein means. We raise the problem of regularizing Wasserstein means and propose several terms tailored to tackle different problems. Our formulation is based on the variational transportation to distribute a sparse discrete measure into the target domain. The resulting sparse representation well captures the desired property of the domain while reducing the mapping cost. We demonstrate the scalability and robustness of our method with examples in domain adaptation, point set registration, and skeleton layout.
Liang Mi, Wen Zhang 0010, Yalin Wang 0001
AAAI1
2018 Variational Wasserstein Clustering
Liang Mi, Wen Zhang 0010, Xianfeng Gu, Yalin Wang 0001
ECCV (15)1
2017 An Optimal Transportation Based Univariate Neuroimaging Index
abstract
The alterations of brain structures and functions have been considered closely correlated to the change of cognitive performance due to neurodegenerative diseases such as Alzheimer's disease. In this paper, we introduce a variational framework to compute the optimal transformation (OT) in 3D space and propose a univariate neuroimaging index based on OT to measure such alterations. We compute the OT from each image to a template and measure the Wasserstein distance between them. By comparing the distances from all the images to the common template, we obtain a concise and informative index for each image. Our framework makes use of the Newton's method, which reduces the computational cost and enables itself to be applicable to large-scale datasets. The proposed work is a generic approach and thus may be applicable to various volumetric brain images, including structural magnetic resonance (sMR) and fluorodeoxyglucose positron emission tomography (FDG-PET) images. In the classification between Alzheimer's disease patients and healthy controls, our method achieves an accuracy of 82:30% on the Alzheimers Disease Neuroimaging Initiative (ADNI) baseline sMRI dataset and outperforms several other indices. On FDG-PET dataset, we boost the accuracy to 88:37% by leveraging pairwise Wasserstein distances. In a longitudinal study, we obtain a 5% significance with p-value = 1:13 ×105 in a t-test on FDG-PET. The results demonstrate a great potential of the proposed index for neuroimage analysis and the precision medicine research.
Liang Mi, Wen Zhang 0010, Junwei Zhang 0010, Yonghui Fan, Dhruman Goradia, Kewei Chen 0001, Eric Reiman, Xianfeng Gu, Yalin Wang 0001
ICCV1
2013 Robust matching of SIFT keypoints via adaptive distance ratio thresholding
abstract
This paper presents a robust method to search for the correct SIFT keypoint matches with adaptive distance ratio threshold. Firstly, the reference image is analyzed by extracting some characteristics of its SIFT keypoints, such as their distance to the object boundary and the number of their neighborhood keypoints. The matching credit of each keypoint is evaluated based on its characteristics. Secondly, an adaptive distance ratio threshold for the keypoint is determined based on its matching credit to identify the correctness of its best match in the source image. The adaptive threshold loosens the matching conditions for keypoints of high matching credits and tightens the conditions for those of low matching credits. Our approach improves the scheme of SIFT keypoint matching by applying adaptive distance ratio threshold rather than global threshold that ignores different matching credits of various keypoints. The experiment results show that our algorithm outperforms the standard SIFT matching method in some complicated cases of object recognition, in which it discards more false matches as well as preserves more correct matches.
Liang Mi, Yu Qiao 0003, Jie Yang 0002, Li Bai 0001
ICMV1