VLDB 2026 Research / reviewers in the wild / expert
Mengyu Yao
dblp:393/4601
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0005-8220-3470ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PredComp: Predicting Compiler Optimization Options with Multi-stage LearningabstractStandard compiler optimization levels, such as -O3 , which provides a fixed optimization strategy for all programs, often fail to deliver the optimal performance. Compiler auto-tuning techniques can deliver substantial speedups, but existing methods present a difficult tradeoff. While dynamic iterative approaches are effective, their requirement for repeated compilation and execution incurs high overhead, which limits their practicality. Conversely, static prediction methods offer a low-overhead alternative. However, they face a vast search space and must comprehensively learn both option-option interactions and option-program feature relationships. To overcome the challenge, we propose PredComp , a novel static framework that leverages the divide and conquer paradigm to predict desired option sets. PredComp decomposes the search space by partitioning options into distinct subspaces based on their relationships, making the prediction problem tractable. It first predicts promising option sub-sets within each subspace, focusing only on intra-subspace option interactions and their preferred program features. Then, it adopts a combination model that aggregates these top-ranked sub-sets, prioritizes inter-subspace option interactions and corresponding features to construct globally desired sets. Experiments on three widely used benchmark suites and one real-world application show that PredComp achieves average speedups of 1.1011× over -O3 with a single prediction. Notably, it achieves performance comparable to dynamic iterative methods while reducing tuning time from hours or days to seconds, thereby making static prediction a practical solution for large-scale and frequently evolving software. Bingyu Gao, Mengyu Yao, Zhihong Xue, Xiangqun Chen, Ding Li 0001, Yao Guo 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Grouptuner: Efficient Group-Aware Compiler Auto-tuningabstractModern compilers typically provide hundreds of options to optimize program performance, but users often cannot fully leverage them due to the huge number of options. While standard optimization combinations (e.g., -O3) provide reasonable defaults, they often fail to deliver near-peak performance across diverse programs and architectures. To address this challenge, compiler auto-tuning techniques have emerged to automate the discovery of improved option combinations. Existing techniques typically focus on identifying critical options and prioritizing them during the search to improve efficiency. However, due to limited tuning iterations, the resulting data is often sparse and noisy, making it highly challenging to accurately identify critical options. As a result, these algorithms are prone to being trapped in local optima. To address this limitation, we propose GroupTuner, a group-aware auto-tuning technique that directly applies localized mutation to coherent option groups based on historically best-performing combinations, thus avoiding explicitly identifying critical options. By forgoing the need to know precisely which options are most important, GroupTuner maximizes the use of existing performance data, ensuring more targeted exploration. Extensive experiments demonstrate that GroupTuner can efficiently discover competitive option combinations, achieving an average performance improvement of 12.39% over -O3 while requiring only 77.21% of the time compared to the random search algorithm, significantly outperforming state-of-the-art methods. Bingyu Gao, Mengyu Yao, Ding Li 0001, Xiangqun Chen, Yao Guo 0001 |
LCTES | 2 |
| 2025 | I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
Yifeng Cai, Mengyu Yao, Xiaoke Zhao, Zhe Liu 0001, Xiangqun Chen, Yao Guo 0001, Ding Li 0001 |
USENIX Security Symposium | 3 |
| 2025 | TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments when Attackers Have Pre-Trained ModelsabstractTrusted Execution Environments (TEEs) are used to safeguard on-device models. However, directly employing TEEs to secure the entire DNN model is challenging due to the limited computational speed. Utilizing GPU can accelerate DNN’s computation speed but widely available commercial GPUs usually lack security protection. To this end, scholars introduce TEE-Shielded DNN Partition (TSDP), a method that protects privacy-sensitive weights within TEEs and offloads insensitive weights to GPUs. Nevertheless, current methods do not consider the presence of a knowledgeable adversary who can access abundant publicly available pre-trained models and datasets. This article investigates the security of the existing methods against such a knowledgeable adversary and reveals their inability to fulfill their security promises. Consequently, we introduce a novel partition before training strategy, which effectively separates privacy-sensitive weights from other components of the model. Our evaluation demonstrates that our approach can offer full model protection with a computational cost reduced by a factor of 10. In addition to traditional CNN models, we also demonstrate the scalability to large language models. Our approach can compress the private functionalities of the large language model to lightweight slices and achieve the same level of protection as the shielding-whole-model baseline. Ding Li 0001, Ziqi Zhang 0017, Mengyu Yao, Yifeng Cai, Yao Guo 0001, Xiangqun Chen |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | Not All Exceptions Are Created Equal: Triaging Error Logs in Real-World EnterprisesabstractError logs like Java exceptions play a crucial role in diagnosing and resolving errors within the industry. Nonetheless, the extensive logging of Java exceptions may result in exception fatigue in large-scale Java systems at an industrial level, where the frequency of Java exceptions being generated surpasses developers’ ability to manage them effectively. Regrettably, there is a lack of research on the seriousness, prevalence, and solutions to this problem. To close this gap, we first make a comprehensive investigation into the exception fatigue problem within a prominent Internet corporation in China, namely Alibaba, confirming its importance in the industry. Consequently, we introduce a novel solution called ABEL , designed to automatically pinpoint the most relevant exceptions associated with software failures. The key challenge lies in the randomness of exceptions, which prevents existing sequence-based techniques from being effective. To address this challenge, ABEL establishes correlations between Java exceptions and the Key Performance Indicator (KPI) of applications, enabling the identification of exceptions leading to irregularities in KPI. Our evaluation of ABEL across four Java applications and five business KPIs within Alibaba illustrates its capability to pinpoint the primary cause of exception logs with an AC@5 (top-5 accuracy) exceeding 90%, effectively mitigating the exception fatigue problem within Alibaba. Furthermore, it can identify the root-cause exceptions in a real software failure within just 4 minutes, outperforming the manual investigation process by over an hour. Mengyu Yao, Shaofei Li, Dingyu Yang, Zheshun Wu, Xiaojun Qu, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Xiangqun Chen |
ACM Trans. Softw. Eng. Methodol. | 2 |