VLDB 2026 Research / reviewers in the wild / expert
Mingyi Zhou
dblp:238/4774
· DBLP profile ↗
15ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-3514-0372ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy Protection Against Personalized Text-to-Image Synthesis via Cross-image Consistency ConstraintsabstractThe rapid advancement of diffusion models and personalization techniques has made it possible to recreate individual portraits from just a few publicly available images. While such capabilities empower various creative applications, they also introduce serious privacy concerns, as adversaries can exploit them to generate highly realistic impersonations. To counter these threats, anti-personalization methods have been proposed, which add adversarial perturbations to published images to disrupt the training of personalization models. However, existing approaches largely overlook the intrinsic multi-image nature of personalization and instead adopt a naive strategy of applying perturbations independently, as commonly done in single-image settings. This neglects the opportunity to leverage inter-image relationships for stronger privacy protection. Therefore, we advocate for a group-level perspective on privacy protection against personalization. Specifically, we introduce Cross-image Anti-Personalization (CAP), a novel framework that enhances resistance to personalization by enforcing style consistency across perturbed images. Furthermore, we develop a dynamic ratio adjustment strategy that adaptively balances the impact of the consistency loss throughout the attack iterations. Extensive experiments on the classical CelebA-HQ and VGGFace2 benchmarks show that CAP outperforms eight existing methods. Guanyu Wang 0005, Kailong Wang 0001, Yihao Huang 0001, Mingyi Zhou, Geguang Pu, Li Li 0029 |
ICMR | 4 |
| 2026 | Effective Fine-tuning for Low-resource Languages: A Case Study of Cangjie
Zhaofeng Liu, Mingyi Zhou, Zihan Huang, Wei Ma 0014, Li Li 0029 |
Empir. Softw. Eng. | 3 |
| 2025 | UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action GenerationabstractMobile task automation is an emerging technology that leverages AI to automatically execute routine tasks by users' commands on mobile devices like Android, thus enhancing efficiency and productivity.While large language models (LLMs) excel at general mobile tasks through training on massive datasets, they struggle with app-specific workflows.To solve this problem, we designed UI Map, a structured representation of target app's UI information.We further propose a UI Map-guided LLM-based approach UICOMPASS to automate mobile tasks.Specifically, UICOMPASS first leverages static analysis and LLMs to automatically build UI Map from either source codes of apps or byte codes (i.e., APK packages).During task execution, UICOMPASS mines the task-relevant information from UI Map to feed into the LLMs, generates a planned path, and adaptively adjusts the path based on the actual app state and action history.Experimental results demonstrate that UICOMPASS achieves a 14.52% higher task executing success rate than SOTA approaches.Even when only APK is available, UICOMPASS maintains superior performance, demonstrating its applicability to closed-source apps. Yuanzhang Lin, He Rui, Qingao Dong, Mingyi Zhou, Xiang Gao 0012, Hailong Sun 0001 |
EMNLP | 5 |
| 2025 | An Empirical Study on UI Overlap in OpenHarmony ApplicationsabstractUI overlap is a phenomenon where one UI component visually covers another. While this overlap is necessary to construct rich visual hierarchies, it is also a root cause of usability issues and performance bottlenecks. However, a systematic, data-driven understanding of its prevalence and patterns has been lacking. To bridge this gap, we conduct the first large-scale empirical study on UI overlap in the OpenHarmony ecosystem. We analyze 100 popular apps, classifying 33,262,624 overlap instances through a novel three-tiered taxonomy. Our findings reveal that high-cost occlusion is a critical and previously hard-to-detect performance defect where resource-intensive components are rendered while visually obscured. We propose HCO-Eye, an innovative tool that leverages multimodal vision-language models (VLMs) to automatically detect such issues, successfully identifying 34 high-cost occlusion cases in commercial apps. Our study not only provides the first comprehensive understanding of UI overlap in OpenHarmony but also offers a practical tool to automatically diagnose complex performance-related UI bugs. Our tools are publicly available. Farong Liu, Mingyi Zhou, Li Li 0029 |
ASE | 2 |
| 2025 | Context-Sensitive Pointer Analysis for ArkTSabstractCurrent call graph generation methods for ArkTS, a new programming language for OpenHarmony, exhibit precision limitations when supporting advanced static analysis tasks such as data flow analysis and vulnerability pattern detection, while the workflow of traditional JavaScript(JS)/TypeScript(TS) analysis tools fails to interpret ArkUI component tree semantics. The core technical bottleneck originates from the closure mechanisms inherent in TypeScript’s dynamic language features and the interaction patterns involving OpenHarmony’s framework APIs. Existing static analysis tools for ArkTS struggle to achieve effective tracking and precise deduction of object reference relationships, leading to topological fractures in call graph reachability and diminished analysis coverage. This technical limitation fundamentally constrains the implementation of advanced program analysis techniques.Therefore, in this paper, we propose a tool named ArkAnalyzer Pointer Analysis Kit (APAK), the first context-sensitive pointer analysis framework specifically designed for ArkTS. APAK addresses these challenges through a unique ArkTS heap object model and a highly extensible plugin architecture, ensuring future adaptability to the evolving OpenHarmony ecosystem. In the evaluation, we construct a dataset from 1,663 real-world applications in the OpenHarmony ecosystem to evaluate APAK, demonstrating APAK’s superior performance over CHA/RTA approaches in critical metrics including valid edge coverage (e.g., a 7.1% reduction compared to CHA and a 34.2% increase over RTA). The improvement in edge coverage systematically reduces false positive rates from 20% to 2%, enabling future exploration of establishing more complex program analysis tools based on our framework. Our proposed APAK has been merged into the official static analysis framework ArkAnalyzer for OpenHarmony. Yizhuo Yang 0005, Mingyi Zhou, Li Li 0029 |
ASE | 3 |
| 2025 | LLM for Mobile: An Initial RoadmapabstractWhen mobile meets LLMs, mobile app users deserve to have more intelligent usage experiences. For this to happen, we argue that there is a strong need to apply LLMs for the mobile ecosystem. We therefore provide a research roadmap for guiding our fellow researchers to achieve that as a whole. In this roadmap, we sum up six directions that we believe are urgently required for research to enable native intelligence in mobile devices. In each direction, we further summarize the current research progress and the gaps that still need to be filled by our fellow researchers. Daihang Chen, Yonghui Liu 0001, Mingyi Zhou, Yanjie Zhao 0001, Haoyu Wang 0001, Shuai Wang 0011, Xiao Chen 0002, Tegawendé F. Bissyandé, Jacques Klein, Li Li 0029 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | Concealing Sensitive Samples against Gradient Leakage in Federated LearningabstractFederated Learning (FL) is a distributed learning paradigm that enhances users' privacy by eliminating the need for clients to share raw, private data with the server. Despite the success, recent studies expose the vulnerability of FL to model inversion attacks, where adversaries reconstruct users’ private data via eavesdropping on the shared gradient information. We hypothesize that a key factor in the success of such attacks is the low entanglement among gradients per data within the batch during stochastic optimization. This creates a vulnerability that an adversary can exploit to reconstruct the sensitive data. Building upon this insight, we present a simple, yet effective defense strategy that obfuscates the gradients of the sensitive data with concealed samples. To achieve this, we propose synthesizing concealed samples to mimic the sensitive data at the gradient level while ensuring their visual dissimilarity from the actual sensitive data. Compared to the previous art, our empirical evaluations suggest that the proposed technique provides the strongest protection while simultaneously maintaining the FL performance. Code is located at https://github.com/JingWu321/DCS-2. Jing Wu 0021, Munawar Hayat, Mingyi Zhou, Mehrtash Harandi |
AAAI | 3 |
| 2024 | Investigating White-Box Attacks for On-Device ModelsabstractNumerous mobile apps have leveraged deep learning capabilities. However, on-device models are vulnerable to attacks as they can be easily extracted from their corresponding mobile apps. Although the structure and parameters information of these models can be accessed, existing on-device attacking approaches only generate black-box attacks (i.e., indirect white-box attacks), which are less effective and efficient than white-box strategies. This is because mobile deep learning (DL) frameworks like TensorFlow Lite (TFLite) do not support gradient computing (referred to as non-debuggable models), which is necessary for white-box attacking algorithms. Thus, we argue that existing findings may underestimate the harm-fulness of on-device attacks. To validate this, we systematically analyze the difficulties of transforming the on-device model to its debuggable version and propose a Reverse Engineering framework for On-device Models (REOM), which automatically reverses the compiled on-device TFLite model to its debuggable version, enabling attackers to launch white-box attacks. Our empirical results show that our approach is effective in achieving automated transformation (i.e., 92.6%) among 244 TFLite models. Compared with previous attacks using surrogate models, REOM enables attackers to achieve higher attack success rates (10.23%→89.03%) with a hundred times smaller attack perturbations (1.0→0.01). Our findings emphasize the need for developers to carefully consider their model deployment strategies, and use white-box methods to evaluate the vulnerability of on-device models. Our artifacts 1 are available. Mingyi Zhou, Xiang Gao 0012, Jing Wu 0021, Kui Liu 0001, Hailong Sun 0001, Li Li 0029 |
ICSE | 1 |
| 2024 | Model-less Is the Best Model: Generating Pure Code Implementations to Replace On-Device DL ModelsabstractRecent studies show that on-device deployed deep learning (DL) models, such as those of Tensor Flow Lite (TFLite), can be easily extracted from real-world applications and devices by attackers to generate many kinds of adversarial and other attacks. Although securing deployed on-device DL models has gained increasing attention, no existing methods can fully prevent these attacks. Traditional software protection techniques have been widely explored. If on-device models can be implemented using pure code, such as C++, it will open the possibility of reusing existing robust software protection techniques. However, due to the complexity of DL models, there is no automatic method that can translate DL models to pure code. To fill this gap, we propose a novel method, CustomDLCoder, to automatically extract on-device DL model information and synthesize a customized executable program for a wide range of DL models. CustomDLCoder first parses the DL model, extracts its backend computing codes, configures the extracted codes, and then generates a customized program to implement and deploy the DL model without explicit model representation. The synthesized program hides model information for DL deployment environments since it does not need to retain explicit model representation, preventing many attacks on the DL model. In addition, it improves ML performance because the customized code removes model parsing and preprocessing steps and only retains the data computing process. Our experimental results show that CustomDLCoder improves model security by disabling on-device model sniffing. Compared with the original on-device platform (i.e., TFLite), our method can accelerate model inference by 21.0% and 24.3% on x86-64 and ARM64 platforms, respectively. Most importantly, it can significantly reduce memory consumption by 68.8% and 36.0% on x86-64 and ARM64 platforms, respectively. Mingyi Zhou, Xiang Gao 0012, John C. Grundy, Chunyang Chen 0001, Xiao Chen 0002, Li Li 0029 |
ISSTA | 1 |
| 2024 | DynaMO: Protecting Mobile DL Models through Coupling Obfuscated DL OperatorsabstractDeploying deep learning (DL) models on mobile applications (Apps) has become ever-more popular. However, existing studies show attackers can easily reverse-engineer mobile DL models in Apps to steal intellectual property or generate effective attacks. A recent approach, Model Obfuscation, has been proposed to defend against such reverse engineering by obfuscating DL model representations, such as weights and computational graphs, without affecting model performance. These existing model obfuscation methods use static methods to obfuscate the model representation, or they use half-dynamic methods but require users to restore the model information through additional input arguments. However, these static methods or half-dynamic methods cannot provide enough protection for on-device DL models. Attackers can use dynamic analysis to mine the sensitive information in the inference codes as the correct model information and intermediate results must be recovered at runtime for static and half-dynamic obfuscation methods. We assess the vulnerability of the existing obfuscation strategies using an instrumentation method and tool, DLModelExplorer, that dynamically extracts correct sensitive model information (i.e., weights, computational graph) at runtime. Experiments show it achieves very high attack performance (e.g., 98.76% of weights extraction rate and 99.89% of obfuscating operator classification rate). To defend against such attacks based on dynamic instrumentation, we propose DynaMO, a Dynamic Model Obfuscation strategy similar to Homomorphic Encryption. The obfuscation and recovery process can be done through simple linear transformation for the weights of randomly coupled eligible operators, which is a fully dynamic obfuscation strategy. Experiments show that our proposed strategy can dramatically improve model security compared with the existing obfuscation strategies, with only negligible overheads for on-device models. Our prototype tool is publicly available at https://github.com/zhoumingyi/DynaMO. Mingyi Zhou, Xiang Gao 0012, Xiao Chen 0002, Chunyang Chen 0001, John C. Grundy, Li Li 0029 |
ASE | 1 |
| 2023 | ModelObfuscator: Obfuscating Model Information to Protect Deployed ML-Based SystemsabstractMore and more edge devices and mobile apps are leveraging deep learning (DL) capabilities. Deploying such models on devices – referred to as on-device models – rather than as remote cloud-hosted services, has gained popularity because it avoids transmitting user’s data off of the device and achieves high response time. However, on-device models can be easily attacked, as they can be accessed by unpacking corresponding apps and the model is fully exposed to attackers. Recent studies show that attackers can easily generate white-box-like attacks for an on-device model or even inverse its training data. To protect on-device models from white-box attacks, we propose a novel technique called model obfuscation. Specifically, model obfuscation hides and obfuscates the key information – structure, parameters and attributes – of models by renaming, parameter encapsulation, neural structure obfuscation, shortcut injection, and extra layer injection. We have developed a prototype tool ModelObfuscator to automatically obfuscate on-device TFLite models. Our experiments show that this proposed approach can dramatically improve model security by significantly increasing the difficulty of parsing models’ inner information, without increasing the latency of DL models. Our proposed on-device model obfuscation has the potential to be a fundamental technique for on-device model deployment. Our prototype tool is publicly available at https://github.com/zhoumingyi/ModelObfuscator. Mingyi Zhou, Xiang Gao 0012, Jing Wu 0021, John C. Grundy, Xiao Chen 0002, Chunyang Chen 0001, Li Li 0029 |
ISSTA | 1 |
| 2021 | Deep Learning-Based Regional Sub-models Integration for Parkinson's Disease Diagnosis Using Diffusion Tensor Imaging
Hengling Zhao, Chih-Chien Tsai, Ce Zhu, Mingyi Zhou, Jiun-Jie Wang, Yipeng Liu 0001 |
ICIG (2) | 4 |
| 2020 | DaST: Data-Free Substitute Training for Adversarial AttacksabstractMachine learning models are vulnerable to adversarial examples. For the black-box setting, current substitute attacks need pre-trained models to generate adversarial examples. However, pre-trained models are hard to obtain in real-world tasks. In this paper, we propose a data-free substitute training method (DaST) to obtain substitute models for adversarial black-box attacks without the requirement of any real data. To achieve this, DaST utilizes specially designed generative adversarial networks (GANs) to train the substitute models. In particular, we design a multi-branch architecture and label-control loss for the generative model to deal with the uneven distribution of synthetic samples. The substitute model is then trained by the synthetic samples generated by the generative model, which are labeled by the attacked model subsequently. The experiments demonstrate the substitute models produced by DaST can achieve competitive performance compared with the baseline models which are trained by the same train set with attacked models. Additionally, to evaluate the practicability of the proposed method on the real-world task, we attack an online machine learning model on the Microsoft Azure platform. The remote model misclassifies 98.35% of the adversarial examples crafted by our method. To the best of our knowledge, we are the first to train a substitute model for adversarial attacks without any real data. Mingyi Zhou, Jing Wu 0021, Yipeng Liu 0001, Shuaicheng Liu, Ce Zhu |
CVPR | 1 |
| 2019 | Early diagnosis of Parkinson's disease from multiple voice recordings by simultaneous sample and feature selection
Ce Zhu, Mingyi Zhou, Yipeng Liu 0001 |
Expert Syst. Appl. | 3 |
| 2019 | Tensor rank learning in CP decomposition via convolutional neural network
Mingyi Zhou, Yipeng Liu 0001, Zhen Long, Longxi Chen, Ce Zhu |
Signal Process. Image Commun. | 1 |