Miaohui Song

dblp:330/3631 · also Miao-Hui Song · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-4280-7406ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
2 papers
Edge and fog computing · 67% Datacenter networks · 18% Internet architecture and protocols · 16%
Artificial intelligence
4 papers
Efficient and distributed learning · 33% Information extraction and text analysis · 24% Knowledge representation and reasoning · 24%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 36% Embedded and real-time systems · 36% Performance modeling and evaluation · 14%
Network and information security
1 paper
Cryptographic protocols and secure computation · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic protocols and secure computation
secure multiparty computation
1.012026
STIP: Three-Party Privacy-Preserving and Lossless Inference for Large Transformers in Production · NDSS 2026
Machine learning › Efficient and distributed learning
model compression
0.922024
Efficient Deep Ensemble Inference via Query Difficulty-dependent Task Scheduling · ICDE 2023
InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric Inference · IEEE Trans. Mob. Comput. 2024
Edge and fog computing
model deployment
0.912025
Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models · IEEE Trans. Mob. Comput. 2025
Edge and fog computing › edge inference
on-device inference
0.912025
Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models · IEEE Trans. Mob. Comput. 2025
Datacenter networks › low-latency networking
tail latency reduction
0.912025
Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models · IEEE Trans. Mob. Comput. 2025
Edge and fog computing › edge inference
mobile inference
0.812024
InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric Inference · IEEE Trans. Mob. Comput. 2024
Internet architecture and protocols
redundancy elimination
0.812024
InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric Inference · IEEE Trans. Mob. Comput. 2024
Edge and fog computing
resource-efficient inference
0.812024
InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric Inference · IEEE Trans. Mob. Comput. 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontology-based annotation
0.712023
CoTel: Ontology-Neural Co-Enhanced Text Labeling · WWW 2023
Natural language and speech › Information extraction and text analysis › data annotation
text annotation
0.712023
CoTel: Ontology-Neural Co-Enhanced Text Labeling · WWW 2023
Embedded and real-time systems › real-time scheduling
deadline-aware scheduling
0.712023
Efficient Deep Ensemble Inference via Query Difficulty-dependent Task Scheduling · ICDE 2023
Electronic design automation › high-level synthesis
scheduling
0.712023
Efficient Deep Ensemble Inference via Query Difficulty-dependent Task Scheduling · ICDE 2023
Parallel and multicore computing
load balancing
0.312025
Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models · IEEE Trans. Mob. Comput. 2025
Performance modeling and evaluation
queueing analysis
0.312025
Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models · IEEE Trans. Mob. Comput. 2025
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.212023
Efficient Deep Ensemble Inference via Query Difficulty-dependent Task Scheduling · ICDE 2023
Machine learning and data management
weak supervision
0.212023
CoTel: Ontology-Neural Co-Enhanced Text Labeling · WWW 2023

Methods — techniques the papers use, named apart from their topics

queuing theory · 1.7load balancing · 1.7heterogeneous model deployment · 1.7hypothesis complexity analysis · 1.5feature embedding · 1.5end-to-end learning · 1.5task scheduling · 1.3query difficulty estimation · 1.3ontology extraction · 1.3neural network · 1.3loss prediction · 1.3
YearPublicationVenuePosition
2026 STIP: Three-Party Privacy-Preserving and Lossless Inference for Large Transformers in Production
Mu Yuan, Lan Zhang 0002, Yihang Cheng 0002, Miaohui Song, Guoliang Xing, Xiang-Yang Li 0001
NDSS4
2025 ProxySampler: Proxy Informativeness Estimation for Efficient Data Selection in Active Learning
abstract
Large-scale data analysis services require efficient periodic model updates to adapt to the possibly changing data distributions. Manually labeling all available samples for task model updates is infeasible for a large sample scale. Active learning technique is proposed to iteratively select subsets of the most informative samples for labeling. From our experience of applying active learning in a real-world video analysis system, we identify a previously overlooked bottleneck of time cost: data selection. Existing active learning methods select data by estimating informativeness (e.g., output confidence) over all unlabeled samples in each iteration. This data selection process can take up to 42% of the time cost of end-to-end model updates in our system (totals include the time for manual labeling, data selection, and model updates.). To address the time cost bottleneck caused by data selection, we propose a new idea: proxy informativeness estimation. We start with modeling the time cost of data selection, from which we identify three key factors: unit estimation cost, the number of samples for estimation, and the number of iteration rounds. The influence of the first two factors increases cumulatively with the number of iteration rounds. Correspondingly, we design a proxy estimator and a sample pooling method, respectively. Our proxy estimator is a lightweight neural network for direct informativeness estimation to replace the role of the high-cost task model, thus reducing the unit cost. And, our sample pooling method leverages historical estimation results to narrow the scope of sample candidates. Based on the above design, we develop ProxySampler, which can be integrated with various active learning approaches as a plug-in. Experimental results show that integrating ProxySampler with state-of-the-art active learning methods can reduce the time cost by 53.6-83.3% (a 2.15-6.01x speedup) when achieving the same accuracy.
Miaohui Song, Lan Zhang 0002, Mu Yuan, Yijun Liu 0003
CIKM1
2025 Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models
abstract
Serving machine learning models on edge, mobile, and embedded devices places stringent requirements on inference latency. From operating a real enterprise service, we observed that even a fully optimized model could lead to severe violations of latency objectives when the load surges. A straightforward and mature approach is to auto-scale multiple models to balance the load. However, unlike cloud clusters, edge or mobile devices usually cannot afford to deploy multiple model replicas. Therefore, in this paper, we explore a new idea: in addition to the original model, we deploy one (or more) heterogeneous model(s) with much smaller resource overhead on the device, and perform load balancing among all models. We overcame the technical challenges posed by performance dynamics and developed InferRouter based on queuing theory. We implement and evaluate InferRouter on three real on-device inference systems, covering mobile sensing, video analytics, and natural language processing applications. Experimental results show that compared with strong baselines, InferRouter can decrease 85.2% P99 latency (5.8x faster) and improve 5.9% accuracy on the mobile workload. For a traffic video analytics task, InferRouter achieves 55.1% higher accuracy with zero deadline misses. InferRouter also shows its advantages in saving resources compared with auto-scaling and offloading approaches.
Mu Yuan, Lan Zhang 0002, Di Duan, Liekang Zeng, Miaohui Song, Zichong Li, Guoliang Xing, Xiang-Yang Li 0001
IEEE Trans. Mob. Comput.5
2024 InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric Inference
abstract
Mobile-centric AI applications have high requirements for the resource-efficiency of model inference. Input filtering is a promising approach to eliminate redundancy so as to reduce the cost of inference. Previous efforts have tailored effective solutions for many applications, but left two essential questions unanswered: (1)theoretical filterability of an inference workloadto guide the application of input filtering techniques, thereby avoiding the trial-and-error cost for resource-constrained mobile applications; (2)robust discriminability of feature embeddingto allow input filtering to be widely effective for diverse inference tasks and input content. To answer them, we first formulate the input filtering problem and theoretically compare the hypothesis complexity of inference models and input filters to understand the optimization potential. Then we propose the first end-to-end learnable input filtering framework that covers most state-of-the-art methods and surpasses them in feature embedding with robust discriminability. We design and implementInFithat supports different input modalities and mobile-centric deployments. Comprehensive evaluations confirm our theoretical results and show thatInFioutperforms strong baselines in applicability, accuracy, and efficiency.InFican achieve 8.5× throughput and save 95% bandwidth, while keeping over 90% accuracy, for a video analytics application on mobile platforms.
Mu Yuan, Lan Zhang 0002, Fengxiang He, Xueting Tong, Miaohui Song, Zhengyuan Xu, Xiang-Yang Li 0001
IEEE Trans. Mob. Comput.5
2023 Efficient Deep Ensemble Inference via Query Difficulty-dependent Task Scheduling
abstract
Deep ensemble learning has been widely adopted to boost accuracy through combing outputs from multiple deep models prepared for the same task. However, the extra computation and memory cost it entails could impose an unacceptably high deadline miss rate in latency-sensitive tasks. Conventional approaches, including ensemble selection, focus on accuracy while ignoring deadline constraints, and thus cannot smartly cope with bursty query traffic and queries with different hardness. This paper explores redundancy in deep ensemble model inference and presents Schemble, a query difficulty-dependent task scheduling framework. Schemble treats ensemble inference progress as multiple base model inference tasks and schedules tasks for queries based on their difficulty and queuing status. We evaluate Schemble on real-world datasets, considering intelligent Q&A system, video analysis and image retrieval as the running applications. Experimental results show that Schemble achieves a 5× lower deadline miss rate and improves the accuracy by 30.8% given deadline constraints.
Zichong Li, Lan Zhang 0002, Mu Yuan, Miaohui Song, Qi Song 0004
ICDE4
2023 CoTel: Ontology-Neural Co-Enhanced Text Labeling
abstract
The success of many web services relies on the large-scale domain-specific high-quality labeled dataset. Insufficient public datasets motivate us to reduce the cost of data labeling while maintaining high accuracy in support of intelligent web applications. The rule-based method and the learning-based method are common techniques for labeling. In this work, we study how to utilize the rule-based and learning-based methods for resource-effective text labeling. We propose CoTel, the first ontology-neural co-enhanced framework for text labeling. We propose critical ontology extraction in the rule-based module and ontology-enhanced loss prediction in the learning-based module. CoTel can integrate explicit labeling rules and implicit labeling models and make them help each other to improve resource efficiency in text labeling tasks. We evaluate CoTel on both public datasets and real applications with three different tasks. Compared with the baseline, CoTel can reduce the time cost by 64.75% (a 2.84× speedup) and the number of labeling by 62.07%.
Miaohui Song, Lan Zhang 0002, Mu Yuan, Zichong Li, Qi Song 0004, Yijun Liu 0003, Guidong Zheng
WWW1