EDBT 2026 Demo / reviewers in the wild / expert
Yuantian Miao
dblp:196/4261
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-7333-0305ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 5 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ST-Attention-XAI: Intrinsic Spatio-Temporal Explainability for IoT Intrusion Detection via Attention Analysis
Nimesha Dilini, Nan Sun 0002, Yuantian Miao, Nour Moustafa |
ACISP (2) | 3 |
| 2025 | Large Language Models for Cybersecurity Education: A Survey of Current Practices and Future Directions
Nan Sun 0002, Yuantian Miao, Xiaoxing Mo, Jun Zhang 0010 |
PAKDD (6) | 2 |
| 2025 | BadFU: Backdoor Federated Learning through Adversarial Machine UnlearningabstractFederated learning (FL) has been widely adopted as a decentralized training paradigm that enables multiple clients to collaboratively learn a shared model without exposing their local data. As concerns over data privacy and regulatory compliance grow, machine unlearning, which aims to remove the influence of specific data from trained models, has become increasingly important in the federated setting to meet legal, ethical, or user-driven demands. However, integrating unlearning into FL introduces new challenges and raises largely unexplored security risks. In particular, adversaries may exploit the unlearning process to compromise the integrity of the global model. In this paper, we present the first backdoor attack in the context of federated unlearning, demonstrating that an adversary can inject backdoors into the global model through seemingly legitimate unlearning requests. Specifically, we propose BadFU, an attack strategy where a malicious client uses both backdoor and camouflage samples to train the global model normally during the federated training process. Once the client requests unlearning of the camouflage samples, the global model transitions into a backdoored state. Extensive experiments under various FL frameworks and unlearning strategies validate the effectiveness of BadFU, revealing a critical vulnerability in current federated unlearning practices and underscoring the urgent need for more secure and robust federated unlearning mechanisms. Bingguang Lu, Hongsheng Hu, Yuantian Miao, Shaleeza Sohail, Chaoxiang He, Shuo Wang 0012, Xiao Chen 0002 |
RAID | 3 |
| 2025 | The impact of unsupervised feature selection techniques on the performance and interpretation of defect prediction models
Zhiqiang Li 0003, Wenzhi Zhu, Hongyu Zhang 0002, Yuantian Miao, Jie Ren 0007 |
Autom. Softw. Eng. | 4 |
| 2025 | A survey of coverage-guided greybox fuzzing with deep neural modelsabstractCoverage-guided greybox fuzzing (CGF) has emerged as a powerful technique for software vulnerability detection, yet traditional techniques often struggle with the increasing complexity of modern software systems and the vastness of input spaces. Deep neural networks (DNNs) have begun to fundamentally transform CGF by addressing these limitations through automated feature extraction, adaptive input generation, and intelligent path prioritization. However, despite these advancements, critical gaps persist in understanding the state-of-the-art landscape. Existing studies often lack rigorous benchmarks to evaluate scalability and generalizability, fail to address the interpretability of neural-guided decisions, and overlook the integration of emerging paradigms such as large language models (LLMs) and neurosymbolic reasoning. This survey systematically bridges these gaps by providing a comprehensive taxonomy of DNN-driven CGF techniques, analyzing their strengths and limitations across key fuzzing stages—seed generation, selection, and mutation. We find that although DNNs have significantly improved fuzzing efficiency, challenges such as semantically invalid seeds, high computational overhead, and limited cross-domain adaptability remain unresolved. Most importantly, we identify two transformative directions with the potential to redefine CGF: (1) LLM-powered fuzzing , which combines generative AI with domain-specific fine-tuning to produce context-aware inputs; and (2) neurosymbolic integration , which merges the precision of symbolic execution with the scalability of neural networks to tackle path explosion. By synthesizing these insights, this survey not only clarifies the state-of-the-art but also outlines a roadmap for developing robust, explainable, and widely applicable intelligent fuzzers. The future of CGF lies in hybrid models that integrate data-driven learning with formal methods, paving the way for autonomous vulnerability discovery in an era of increasingly complex software systems. Junyang Qiu, Yupeng Jiang 0002, Yuantian Miao, Wei Luo 0001, Lei Pan 0002, James Xi Zheng |
Inf. Softw. Technol. | 3 |
| 2024 | Optimizing the Utilization of Large Language Models via Schedule Optimization: An Exploratory StudyabstractBackground: Large Language Models (LLMs) have gained significant attention in machine-learning-as-a-service (MLaaS) offerings. In-context learning (ICL) is a technique that guides LLMs towards accurate query processing by providing additional information. However, longer prompts lead to higher costs of LLM service, creating a performance-cost trade-off. Aims: We aim to investigate the potential of combining schedule optimization with ICL to optimize LLM utilization. Method: We conduct an exploratory study. First, we consider the performance-cost trade-off in LLM utilization as a multi-objective optimization problem, aiming to select the most suitable prompt template for each LLM job to maximize accuracy (the percentage of correctly processed jobs) and minimize invocation cost. Next, we investigate three methods for prompt performance prediction to address the challenge of evaluating the accuracy objective in the fitness function, as the result can only be determined after submitting the job to the LLM. Finally, we apply widely used search-based techniques and evaluate their effectiveness. Results: The results indicate that the machine learning-based technique is an effective approach for prompt performance prediction and fitness function calculation. Schedule optimization can achieve higher accuracy or lower cost by selecting a suitable prompt template for each job, compared to simply submitting all jobs using a single prompt template, e.g., saving costs from 21.33% to 86.92% in our experiments on LLM-based log parsing. However, the performance of the evaluated search-based techniques varies across different instances and metrics, with no single technique consistently outperforming the others. Conclusions: This study demonstrates the potential of combining schedule optimization with ICL to improve the utilization of LLMs. However, there is still ample room for improving the searched-based techniques and prompt performance prediction techniques for more cost-effective LLM utilization. Yueyue Liu 0002, Hongyu Zhang 0002, Zhiqiang Li 0003, Yuantian Miao |
ESEM | 4 |
| 2024 | CPLS: Optimizing the Assignment of LLM QueriesabstractLarge Language Models (LLMs) like ChatGPT have gained significant attention because of their impressive capabilities, leading to a dramatic increase in their integration into intelligent software engineering. However, their usage as a service with varying performance and price options presents a challenging trade-off between desired performance and the associated cost. To address this challenge, we propose CPLS, a framework that utilizes transfer learning and local search techniques for assigning intelligent software engineering jobs to LLM-based services. CPLS aims to minimize the total cost of LLM invocations while maximizing the overall accuracy. The framework first leverages knowledge from historical data across different projects to predict the probability of an LLM processing a query correctly. Then, CPLS incorporates problem-specific rules into a local search algorithm to effectively generate Pareto optimal solutions based on the predicted accuracy and cost. To evaluate the proposed approach, we conduct extensive experiments on LLM-based log parsing, a typical software maintenance task. Our experimental results demonstrate that CPLS outperforms the baseline methods, providing solutions with the highest accuracy in 14 out of 16 instances. Compared to the baselines, CPLS achieves an accuracy improvement ranging from 1.24% to 485.54%, or reduces costs by 15.21% to 89.09% while maintaining the highest accuracy achieved by the baselines. Yueyue Liu 0002, Hongyu Zhang 0002, Zhiqiang Li 0003, Yuantian Miao |
ICSME | 4 |
| 2024 | OptLLM: Optimal Assignment of Queries to Large Language ModelsabstractLarge Language Models (LLMs) have garnered considerable attention owing to their remarkable capabilities, leading to an increasing number of companies offering LLMs as services. Different LLMs achieve different performance at different costs. A challenge for users lies in choosing the LLMs that best fit their needs, balancing cost and performance. In this paper, we propose a framework for addressing the cost-effective query allocation problem for LLMs. Given a set of input queries and candidate LLMs, our framework, named OptLLM, provides users with a range of optimal solutions to choose from, aligning with their budget constraints and performance preferences, including options for maximizing accuracy and minimizing cost. OptLLM predicts the performance of candidate LLMs on each query using a multi-label classification model with uncertainty estimation and then iteratively generates a set of non-dominated solutions by destructing and reconstructing the current solution. To evaluate the effectiveness of OptLLM, we conduct extensive experiments on various types of tasks, including text classification, question answering, sentiment analysis, reasoning, and log parsing. Our experimental results demonstrate that OptLLM substantially reduces costs by 2.40% to 49.18% while achieving the same accuracy as the best LLM. Compared to other multi-objective optimization algorithms, OptLLM improves accuracy by 2.94% to 69.05% at the same cost or saves costs by 8.79% and 95.87% while maintaining the highest attainable accuracy. Yueyue Liu 0002, Hongyu Zhang 0002, Yuantian Miao, Van-Hoang Le, Zhiqiang Li 0003 |
ICWS | 3 |
| 2022 | No-Label User-Level Membership Inference for ASR Model Auditing
Yuantian Miao, Chao Chen 0015, Lei Pan 0002, Shigang Liu, Seyit Ahmet Çamtepe, Jun Zhang 0010, Yang Xiang 0001 |
ESORICS (2) | 1 |
| 2021 | The Audio Auditor: User-Level Membership Inference in Internet of Things Voice ServicesabstractAbstract With the rapid development of deep learning techniques, the popularity of voice services implemented on various Internet of Things (IoT) devices is ever increasing. In this paper, we examine user-level membership inference in the problem space of voice services, by designing an audio auditor to verify whether a specific user had unwillingly contributed audio used to train an automatic speech recognition (ASR) model under strict black-box access. With user representation of the input audio data and their corresponding translated text, our trained auditor is effective in user-level audit. We also observe that the auditor trained on specific data can be generalized well regardless of the ASR model architecture. We validate the auditor on ASR models trained with LSTM, RNNs, and GRU algorithms on two state-of-the-art pipelines, the hybrid ASR system and the end-to-end ASR system. Finally, we conduct a real-world trial of our auditor on iPhone Siri, achieving an overall accuracy exceeding 80%. We hope the methodology developed in this paper and findings can inform privacy advocates to overhaul IoT privacy. Yuantian Miao, Minhui Xue 0001, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Benjamin Zi Hao Zhao, Mohamed Ali Kâafar, Yang Xiang 0001 |
Proc. Priv. Enhancing Technol. | 1 |
| 2018 | Comprehensive analysis of network traffic dataabstractSummary With the large volume of network traffic flow, it is necessary to preprocess raw data before classification to gain the accurate results speedily. Feature selection is an essential approach in preprocessing phase. The principal component analysis (PCA) is recognized as an effective and efficient method. In this paper, we classify network traffic flows by using the PCA technique together with 6 machine learning algorithms—Naive Bayes, decision tree, 1‐nearest neighbor, random forest, support vector machine, andH2O. We analyzed the impact of PCA on the classification results by applying each algorithm with and without PCA onto the data set. Experiments were set out by varying the size of input data sets, and the performances were measured from 2 aspects, including average overall accuracy and F‐measure. The computational time was also considered in analyzing the performance. Our results showed that random forest and 1‐nearest neighbor were the top 2 algorithms among all the 6 regarding the 2 metrics mentioned above. Then we continued the study of PCA impact on per class level with these 2 algorithms as examples. And the positive correlation between overall impact and the number of class with significant impact was revealed. Lastly, the visualization was used in exploring the reasons of the impacts caused by PCA. Two factors are considered in PCA's impact on per class level: benefit for classes grouped by PCA and mislabeled error interfered by nearby groups. Yuantian Miao, Zichan Ruan, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Big network traffic data visualization
Zichan Ruan, Yuantian Miao, Lei Pan 0002, Yang Xiang 0001, Jun Zhang 0010 |
Multim. Tools Appl. | 2 |