VLDB 2026 Research / reviewers in the wild / expert
Miaohui Song
dblp:330/3631 · also Miao-Hui Song
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-4280-7406ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STIP: Three-Party Privacy-Preserving and Lossless Inference for Large Transformers in Production
Mu Yuan, Lan Zhang 0002, Yihang Cheng 0002, Miaohui Song, Guoliang Xing, Xiang-Yang Li 0001 |
NDSS | 4 |
| 2025 | ProxySampler: Proxy Informativeness Estimation for Efficient Data Selection in Active LearningabstractLarge-scale data analysis services require efficient periodic model updates to adapt to the possibly changing data distributions. Manually labeling all available samples for task model updates is infeasible for a large sample scale. Active learning technique is proposed to iteratively select subsets of the most informative samples for labeling. From our experience of applying active learning in a real-world video analysis system, we identify a previously overlooked bottleneck of time cost: data selection. Existing active learning methods select data by estimating informativeness (e.g., output confidence) over all unlabeled samples in each iteration. This data selection process can take up to 42% of the time cost of end-to-end model updates in our system (totals include the time for manual labeling, data selection, and model updates.). To address the time cost bottleneck caused by data selection, we propose a new idea: proxy informativeness estimation. We start with modeling the time cost of data selection, from which we identify three key factors: unit estimation cost, the number of samples for estimation, and the number of iteration rounds. The influence of the first two factors increases cumulatively with the number of iteration rounds. Correspondingly, we design a proxy estimator and a sample pooling method, respectively. Our proxy estimator is a lightweight neural network for direct informativeness estimation to replace the role of the high-cost task model, thus reducing the unit cost. And, our sample pooling method leverages historical estimation results to narrow the scope of sample candidates. Based on the above design, we develop ProxySampler, which can be integrated with various active learning approaches as a plug-in. Experimental results show that integrating ProxySampler with state-of-the-art active learning methods can reduce the time cost by 53.6-83.3% (a 2.15-6.01x speedup) when achieving the same accuracy. Miaohui Song, Lan Zhang 0002, Mu Yuan, Yijun Liu 0003 |
CIKM | 1 |
| 2025 | Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous ModelsabstractServing machine learning models on edge, mobile, and embedded devices places stringent requirements on inference latency. From operating a real enterprise service, we observed that even a fully optimized model could lead to severe violations of latency objectives when the load surges. A straightforward and mature approach is to auto-scale multiple models to balance the load. However, unlike cloud clusters, edge or mobile devices usually cannot afford to deploy multiple model replicas. Therefore, in this paper, we explore a new idea: in addition to the original model, we deploy one (or more) heterogeneous model(s) with much smaller resource overhead on the device, and perform load balancing among all models. We overcame the technical challenges posed by performance dynamics and developed InferRouter based on queuing theory. We implement and evaluate InferRouter on three real on-device inference systems, covering mobile sensing, video analytics, and natural language processing applications. Experimental results show that compared with strong baselines, InferRouter can decrease 85.2% P99 latency (5.8x faster) and improve 5.9% accuracy on the mobile workload. For a traffic video analytics task, InferRouter achieves 55.1% higher accuracy with zero deadline misses. InferRouter also shows its advantages in saving resources compared with auto-scaling and offloading approaches. Mu Yuan, Lan Zhang 0002, Di Duan, Liekang Zeng, Miaohui Song, Zichong Li, Guoliang Xing, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric InferenceabstractMobile-centric AI applications have high requirements for the resource-efficiency of model inference. Input filtering is a promising approach to eliminate redundancy so as to reduce the cost of inference. Previous efforts have tailored effective solutions for many applications, but left two essential questions unanswered: (1)theoretical filterability of an inference workloadto guide the application of input filtering techniques, thereby avoiding the trial-and-error cost for resource-constrained mobile applications; (2)robust discriminability of feature embeddingto allow input filtering to be widely effective for diverse inference tasks and input content. To answer them, we first formulate the input filtering problem and theoretically compare the hypothesis complexity of inference models and input filters to understand the optimization potential. Then we propose the first end-to-end learnable input filtering framework that covers most state-of-the-art methods and surpasses them in feature embedding with robust discriminability. We design and implementInFithat supports different input modalities and mobile-centric deployments. Comprehensive evaluations confirm our theoretical results and show thatInFioutperforms strong baselines in applicability, accuracy, and efficiency.InFican achieve 8.5× throughput and save 95% bandwidth, while keeping over 90% accuracy, for a video analytics application on mobile platforms. Mu Yuan, Lan Zhang 0002, Fengxiang He, Xueting Tong, Miaohui Song, Zhengyuan Xu, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Efficient Deep Ensemble Inference via Query Difficulty-dependent Task SchedulingabstractDeep ensemble learning has been widely adopted to boost accuracy through combing outputs from multiple deep models prepared for the same task. However, the extra computation and memory cost it entails could impose an unacceptably high deadline miss rate in latency-sensitive tasks. Conventional approaches, including ensemble selection, focus on accuracy while ignoring deadline constraints, and thus cannot smartly cope with bursty query traffic and queries with different hardness. This paper explores redundancy in deep ensemble model inference and presents Schemble, a query difficulty-dependent task scheduling framework. Schemble treats ensemble inference progress as multiple base model inference tasks and schedules tasks for queries based on their difficulty and queuing status. We evaluate Schemble on real-world datasets, considering intelligent Q&A system, video analysis and image retrieval as the running applications. Experimental results show that Schemble achieves a 5× lower deadline miss rate and improves the accuracy by 30.8% given deadline constraints. Zichong Li, Lan Zhang 0002, Mu Yuan, Miaohui Song, Qi Song 0004 |
ICDE | 4 |
| 2023 | CoTel: Ontology-Neural Co-Enhanced Text LabelingabstractThe success of many web services relies on the large-scale domain-specific high-quality labeled dataset. Insufficient public datasets motivate us to reduce the cost of data labeling while maintaining high accuracy in support of intelligent web applications. The rule-based method and the learning-based method are common techniques for labeling. In this work, we study how to utilize the rule-based and learning-based methods for resource-effective text labeling. We propose CoTel, the first ontology-neural co-enhanced framework for text labeling. We propose critical ontology extraction in the rule-based module and ontology-enhanced loss prediction in the learning-based module. CoTel can integrate explicit labeling rules and implicit labeling models and make them help each other to improve resource efficiency in text labeling tasks. We evaluate CoTel on both public datasets and real applications with three different tasks. Compared with the baseline, CoTel can reduce the time cost by 64.75% (a 2.84× speedup) and the number of labeling by 62.07%. Miaohui Song, Lan Zhang 0002, Mu Yuan, Zichong Li, Qi Song 0004, Yijun Liu 0003, Guidong Zheng |
WWW | 1 |