EDBT 2026 Demo / reviewers in the wild / expert
Yihan Liao
dblp:368/6456
· DBLP profile ↗
19ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0002-8002-9190ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified Benchmark for Out-of-Distribution Detection for Autonomous Driving SystemsabstractAutonomous Driving Systems (ADS) can fail when they encounter inputs that differ from their training data, known as out-of-distribution (OOD) conditions. Such OOD inputs lead ADS to make incorrect driving decisions, resulting in serious safety risks. Reliable OOD detection is therefore essential for enhancing system robustness and preventing hazardous behavior. However, existing literature in the autonomous driving field examines a relatively narrow scope of OOD detectors (e.g., reconstruction-based only) under limited OOD conditions. Jacky W. Keung, Yihan Liao |
AST | 5 |
| 2026 | Where Do Expectations Diverge? an Empirical Analysis of Software Engineering Graduate Preparedness in Industry
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Yihan Liao, Yishu Li |
COMPSAC | 4 |
| 2026 | Designing Psychologically Safe AI Tutors for Students: An Emotion-Aware Post-Hoc Intervention for LLM-Assisted Learning
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Yihan Liao, Zhenyu Mao, Yishu Li |
COMPSAC | 4 |
| 2026 | Artifact-Constrained Agentic Testing for Black-Box System Testing under Actuarial and Regulatory ConstraintsabstractSystem-level testing of insurance software systems is challenging due to long-lived legacy architectures, distributed actuarial logic, regulatory-driven conditional behavior, and exception-heavy workflows, where correctness is defined by compliance with evolving domain constraints rather than deterministic outputs. Existing approaches either focus on artifact-centric validation, rely heavily on human-driven execution and diagnosis, or generate test artifacts in a one-shot manner without sustained adaptation during execution, limiting robustness and reproducibility at the system level. This paper proposes \emph{Artifact-Constrained Agentic Testing} (ACAT), a multi-agent framework that structures black-box system testing as a closed-loop process of planning, execution, diagnosis, and repair, in which LLM-based agents operate under explicitly bounded capabilities and interact with the system under test exclusively through tool-mediated execution. We evaluate ACAT through a controlled industrial study on an insurance broker management system using 50 real-world use cases. The results show that artifact-constrained agentic testing significantly improves test executability and execution stability compared to human-driven testing, while exposing a complementary subset of system-level failures. These findings suggest that agent-assisted testing can enhance system-level automation in complex, regulated software systems without replacing human expertise. Hi Kuen Yu, Jacky W. Keung, Man On Wong, Yicheng Sun, Rachel Samantha Chandra, Yihan Liao |
COMPSAC | 6 |
| 2026 | Industrial log analysis revisited: A task-oriented evaluation of parsing and anomaly detection under real-world constraints
Yicheng Sun, Jacky W. Keung, Yihan Liao, Zhenyu Mao, Hi Kuen Yu |
Inf. Softw. Technol. | 4 |
| 2026 | FedDC: Efficient protection scheme based on chaotic system in federated learningabstractFederated Learning (FL) enables collaborative model training across decentralized clients while keeping raw data local. However, existing privacy-preserving mechanisms, such as secure aggregation and differential privacy (DP), often introduce significant computational overhead or degrade model utility. To address this challenge, we propose FedDC , a lightweight FL framework that combines DP with chaos-based parameter scrambling. Unlike existing approaches that uniformly protect entire model updates, FedDC introduces selective and layer-aware protection for sensitive neural network parameters, enabling flexible privacy protection with minimal overhead. Extensive experiments on four image and text datasets show that FedDC effectively mitigates privacy leakage under both black-box and white-box attacks while maintaining competitive model accuracy and negligible computational overhead. Yihan Liao, Jacky W. Keung, Yurou Dai |
J. Inf. Secur. Appl. | 1 |
| 2025 | PerProb: Indirectly Evaluating Memorization in Large Language ModelsabstractThe rapid advancement of Large Language Models (LLMs) has been driven by extensive datasets that may contain sensitive information, raising serious privacy concerns. One notable threat is the Membership Inference Attack (MIA), where adversaries infer whether a specific sample was used in model training. However, the true impact of MIA on LLMs remains unclear due to inconsistent findings and the lack of standardized evaluation methods, further complicated by the undisclosed nature of many LLM training sets. To address these limitations, we propose PerProb, a unified, label-free framework for indirectly assessing LLM memorization vulnerabilities. PerProb evaluates changes in perplexity and average log probability between data generated by victim and adversary models, enabling an indirect estimation of training-induced memory. Compared with prior MIA methods that rely on member/non-member labels or internal access, PerProb is independent of model and task, and applicable in both black-box and white-box settings. Through a systematic classification of MIA into four attack patterns, we evaluate PerProb’s effectiveness across five datasets, revealing varying memory behaviors and privacy risks among LLMs. Additionally, we assess mitigation strategies, including knowledge distillation, early stopping, and differential privacy, demonstrating their effectiveness in reducing data leakage. Our findings offer a practical and generalizable framework for evaluating and improving LLM privacy. Yihan Liao, Jacky W. Keung, Yicheng Sun |
APSEC | 1 |
| 2025 | Exposing and Defending Membership Leakage in Vulnerability Prediction ModelsabstractNeural models for vulnerability prediction (VP) have achieved impressive performance by learning from large-scale code repositories. However, their susceptibility to Membership Inference Attacks (MIAs), where adversaries aim to infer whether a particular code sample was used during training, poses serious privacy concerns. While MIA has been widely investigated in NLP and vision domains, its effects on security-critical code analysis tasks remain underexplored. In this work, we conduct the first comprehensive analysis of MIA on VP models, evaluating the attack success across various architectures (LSTM, BiGRU, and CodeBERT) and feature combinations, including embeddings, logits, loss, and confidence. Our threat model aligns with black-box and gray-box settings where prediction outputs are observable, allowing adversaries to infer membership by analyzing output discrepancies between training and non-training samples. The empirical findings reveal that logits and loss are the most informative and vulnerable outputs for membership leakage. Motivated by these observations, we propose a Noise-based Membership Inference Defense (NMID), which is a lightweight defense module that applies output masking and Gaussian noise injection to disrupt adversarial inference. Extensive experiments demonstrate that NMID significantly reduces MIA effectiveness, lowering the attack AUC from nearly 1.0 to below 0.65, while preserving the predictive utility of VP models. Our study highlights critical privacy risks in code analysis and offers actionable defense strategies for securing AI-powered software systems. Yihan Liao, Jacky W. Keung, Yicheng Sun |
APSEC | 1 |
| 2025 | Understanding Industrial Log Analysis: A Multi-Dataset Evaluation of Parsing and Anomaly DetectionabstractLog analysis plays a critical role in monitoring and maintaining the safety of industrial software systems. However, most existing research relies heavily on benchmark datasets derived from legacy or open-source systems, which fail to capture the structural diversity and operational complexity of real-world industrial logs. In this study, we present a comprehensive empirical evaluation of log parsing and anomaly detection models across four diverse datasets, including three collected from largescale industrial software deployed in manufacturing, process control, and energy monitoring environments. Our analysis reveals that state-of-the-art models—particularly rule-based parsers and supervised detectors—experience substantial performance degradation when applied to industrial settings. To address this gap, we introduce a unified evaluation framework using representative training subsets, and we highlight the effectiveness of semisupervised and LLM-based approaches in handling heterogeneous, low-resource log environments. The findings offer practical insights into the limitations of current log analysis techniques and suggest design principles for building more robust, domain-adaptive solutions for industrial software risk mitigation. Yicheng Sun, Jacky W. Keung, Yihan Liao, Hi Kuen Yu |
APSEC | 3 |
| 2025 | Towards Lightweight LLM Software Solutions for InsurTech: A Framework for Scalable Question Answering SystemsabstractThe integration of Large Language Models (LLMs) into software systems is transforming regulated sectors like insurance, where precision, compliance, and efficiency are essential. While proprietary LLMs like GPT-4 offer state-of-the-art performance, their closed-source nature and high computational demands constrain adoption in privacy-sensitive and cost-restricted InsurTech environments. In response, this paper investigates how lightweight, open-source LLMs can be effectively deployed for domain-specific question answering in insurance, emphasizing software engineering considerations such as modularity, inference stability, and prompt orchestration. We propose a software-engineered evaluation framework tailored to insurance-related tasks, featuring modular prompt management, automated rubricbased evaluation, and backend support for reproducibility and compliance tracking. A curated benchmark dataset derived from the Hong Kong Insurance Intermediaries Qualifying Examination (IIQE) is constructed to reflect real-world regulatory and operational challenges. Ten open-source models are systematically evaluated across four question types using both standard and Chain-of-Thought (CoT) prompting strategies. Our findings show that compact models such as DeepSeek-R1-1.5B achieve strong accuracy with minimal resource consumption, making them suitable for practical deployment. CoT prompting further enhances reasoning performance, particularly for models with 3B parameters or more. With proper prompt design and modular deployment, lightweight LLMs can support secure, efficient, and interpretable InsurTech applications, enabling trustworthy AI-driven software systems in regulated domains. Hi Kuen Yu, Jacky W. Keung, Yicheng Sun, Yihan Liao, Richard Suen |
APSEC | 4 |
| 2025 | Beyond Log Parsers: A Scalable AI-Driven Framework for Efficient Log Anomaly Detection in Software EngineeringabstractLog anomaly detection is critical for ensuring software system reliability and security, yet challenges persist in log parser dependency, small-scale dataset applicability, and hyperparameter tuning efficiency. Existing methods over-rely on predefined log templates, leading to information loss and high computational overhead. Additionally, anomaly detection models often struggle with limited log data, and hyperparameter tuning remains computationally expensive in dynamic environments. In this paper, we empirically evaluate seven state-of-the-art anomaly detection models across varied software systems, assessing the necessity of log parsers and model performance on small-scale datasets. Furthermore, we propose SMAC-, an enhanced real-time hyperparameter optimization framework, integrating stochastic gradient descent (SGD) and adaptive learning to improve model adaptability and efficiency. Our experiments on six benchmark datasets demonstrate that SMAC-achieves an overall average F1-score improvement of 4.27%, a 27.55% reduction in hyperparameter tuning time compared to other models, and a 1.35% increase in F1-score when adapting to newly emerging logs, compared to its counterpart without SGD integration. These findings underscore the practical advantages of AI-driven log analysis, providing valuable insights into scalable, software-engineered anomaly detection. Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Shuo Liu 0020, Yihan Liao |
COMPSAC | 5 |
| 2025 | StuLAC: An Adaptive LLM-Driven Framework for Scalable Student Feedback Analysis in Software-Driven Educational SystemsabstractWith the growing scalability challenges in higher education, automated student feedback analysis has become crucial for course evaluation and pedagogical improvements. However, traditional methods struggle to handle mixed sentiments, adapt to evolving feedback trends, and maintain computational efficiency. To address these challenges, we propose StuLAC, a Software Engineering-driven framework that integrates Large Language Models (LLMs) with Adaptive Template-Based Caching (ATC). StuLAC employs hierarchical matching for fine-grained classification and dynamically updates feedback templates through context-aware cache refinement. Empirical results on 80,000 student feedback entries demonstrate that StuLAC-generated summaries improve overall quality by 10.5% compared to manually generated reports, while also achieving faster processing times. Additionally, StuLAC attains an 86.4% accuracy and an 86.24% F1-score in sentiment detection. StuLAC’s Feedback Summary Generation provides actionable insights that enhance data-driven decision-making in educational settings. These findings establish StuLAC as a scalable and adaptive solution for improving AI-driven educational feedback systems. Yicheng Sun, Hi Kuen Yu, Jacky W. Keung, Yuchen Cao 0006, Yihan Liao |
COMPSAC | 5 |
| 2025 | Advancing autonomous driving system testing: Demands, challenges, and future directions
Yihan Liao, Jacky W. Keung, Yan Xiao 0002, Yurou Dai |
Inf. Softw. Technol. | 1 |
| 2025 | SemiSMAC: A semi-supervised framework for log anomaly detection with automated hyperparameter tuningabstractContext: Logs generated during software operations are critical for system reliability and anomaly detection. However, their diversity, the scarcity of labeled data, and hyperparameter tuning challenges hinder traditional detection methods. Objective: This paper presents SemiSMAC, a novel semi-supervised framework that leverages the Large Language Model for log parsing and grouping, combined with Sequential Model-based Algorithm Configuration (SMAC) for hyperparameter optimization to enhance anomaly detection. Method: In this work, we leverage ChatGPT for log parsing and introduce a novel log grouping approach. This grouping process requires only a small number of labeled samples, which ChatGPT uses to generate pseudo-labels for the remaining data, thereby expanding the training set. Furthermore, SemiSMAC utilizes a Sequential Model-based Algorithm Configuration (SMAC) to automatically optimize the hyperparameters of the embedded models. This integration leads to consistent performance improvements, particularly in resource-constrained environments. Results: SemiSMAC-LSTM, which uses LSTM as the backbone of the SemiSMAC framework, demonstrates superior performance in experiments on four widely used datasets. It outperforms six benchmark models, including three supervised learning models. In low-resource scenarios, SemiSMAC-LSTM exhibits exceptional robustness, showcasing its effectiveness in handling challenging detection tasks. Conclusion: SemiSMAC demonstrates its potential to revolutionize anomaly detection in both large-scale and low-resource datasets. Its ability to deliver outstanding performance makes it a valuable tool for scalable and automated anomaly detection in real-world applications, paving the way for more reliable and scalable software engineering practices Yicheng Sun, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020, Yihan Liao |
Inf. Softw. Technol. | 5 |
| 2024 | Enhancing the Transferability of Adversarial Attacks for End-to-End Autonomous Driving SystemsabstractAdversarial attacks play an important role in testing and enhancing the reliability of deep learning (DL) systems. Most existing attacks for DL-based autonomous driving systems (ADSs) demonstrate strong performance under the white-box setting but struggle with black-box transferability, while blackbox attacks are more practical in real-world scenarios as they operate without full model access. Numerous transferabilityenhancement techniques have been proposed in other fields (e.g., image classification), however, they remain unexplored for endtoend (E2E) ADSs. Our study fills the gap by conducting the first comprehensive empirical analysis of nine transferability-enhancement methods on E2E ADSs, covering two types: three input transformation enhancements and six attack objective enhancements. We evaluate their effectiveness on two datasets with four steering models. Our findings reveal that, out of nine enhancements, Resizing+ Translation delivers the best black-box transferability, producing up to 9.39° increase in MAE. Pred+Attn serves as the best objective enhancement, producing a maximum of 5.55° (white-box) and 6.21° (black-box) increase in MAE. Through attention heatmap visualizations, we discover that different models focus on similar regions when predicting, thereby enhancing the transferability of attention-based attacks. In conclusion, our study provides valuable results and insights into the transferability-enhancement techniques for E2E ADSs, which also serve as a robust benchmark for further advancements in the autonomous driving field. Jacky W. Keung, Yihan Liao, Yishu Li, Yicheng Sun |
APSEC | 4 |
| 2024 | LLM-Based Class Diagram Derivation from User Stories with Chain-of-Thought PromptingsabstractIn agile requirements engineering, user stories are the primary means of capturing project requirements. However, deriving conceptual models, such as class diagrams, from user stories requires significant manual effort. This paper explores the potential of leveraging Large Language Models (LLMs) and a tailored Chain-of- Thought (CoT) prompting technique to automate this task. We conducted a comprehensive preliminary study to investigate different prompting techniques applied to the task. The study involved comparing LLM-based approaches with guided and unguided human extraction to evaluate the effectiveness of the proposed LLM-based techniques. Our findings demonstrate that LLM-based approaches, particularly when combined with well-crafted few-shot prompts, outperform guided human extraction in identifying classes. However, we also identified areas of suboptimal performance through qualitative analysis. The proposed CoT prompting technique offers a promising pathway to automate the derivation of class diagrams in agile projects, reducing the reliance on manual effort. Our study contributes valuable insights and directions for future research in this field. Yishu Li, Jacky W. Keung, Chun Yong Chong, Yihan Liao |
COMPSAC | 6 |
| 2024 | Delving into Parameter-Efficient Fine-Tuning in Code Change Learning: An Empirical StudyabstractCompared to Full-Model Fine-Tuning (FMFT), Parameter Efficient Fine-Tuning (PEFT) has demonstrated superior performance and lower computational overhead in several code understanding tasks, such as code summarization and code search. This advantage can be attributed to PEFT's ability to alleviate the catastrophic forgetting issue of Pre-trained Language Models (PLMs) by updating only a small number of parameters. As a result, PEFT effectively harnesses the pre-trained general-purpose knowledge for downstream tasks. However, existing studies primarily involve static code comprehension, aligning with the pre-training paradigm of recent PLMs and facilitating knowledge transfer, but they do not account for dynamic code changes. Thus, it remains unclear whether PEFT outperforms FMFT in task-specific adaptation for code-change-related tasks. To address this question, we examine two prevalent PEFT methods, namely Adapter Tuning (AT) and Low-Rank Adaptation (LoRA), and compare their performance with FMFT on five popular PLMs. Specifically, we evaluate their performance on two widely-studied code-change-related tasks: Just-In-Time Defect Prediction (JIT-DP) and Commit Message Generation (CMG). The results demonstrate that both AT and LoRA achieve state-of-the-art (SOTA) results in JIT-DP and exhibit comparable performances in CMG when compared to FMFT and other SOTA approaches. Furthermore, AT and LoRA exhibit superiority in cross-lingual and low-resource scenarios. We also conduct three probing tasks to explain the efficacy of PEFT techniques on JIT-DP and CMG tasks from both static and dynamic perspectives. The study indicates that PEFT, particularly through the use of AT and LoRA, offers promising advantages in code-change-related tasks, surpassing FMFT in certain aspects. This research contributes to a deeper understanding of the capabilities of PEFT in leveraging pre-trained PLMs for dynamic code changes. The replication package is available at https://github.com/ishuoliu/PEFT4CC. Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Fang Liu 0032, Qilin Zhou, Yihan Liao |
SANER | 6 |
| 2024 | TerGEC: A graph enhanced contrastive approach for program termination analysis
Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Yihan Liao, Yishu Li |
Sci. Comput. Program. | 4 |
| 2024 | UniAda: Universal Adaptive Multiobjective Adversarial Attack for End-to-End Autonomous Driving SystemsabstractAdversarial attacks play a pivotal role in testing and improving the reliability of deep learning (DL) systems. Existing literature has demonstrated that subtle perturbations to the input can elicit erroneous outcomes, thereby substantially compromising the security of DL systems. This has emerged as a critical concern in the development of DL-based safety–critical systems like autonomous driving systems (ADSs). The focus of existing adversarial attack methods on end-to-end (E2E) ADSs has predominantly centered on misbehaviors of steering angle, which overlooks speed-related controls or imperceptible perturbations. To address these challenges, we introduce UniAda–a multiobjective white-box attack technique with a core function that revolves around crafting an image-agnostic adversarial perturbation capable of simultaneously influencing both steering and speed controls. UniAda capitalizes on an intricately designed multiobjective optimization function with the adaptive weighting scheme (AWS), enabling the concurrent optimization of diverse objectives. Validated with both simulated and real-world driving data, UniAda outperforms five benchmarks across two metrics, inducing steering and speed deviations from 3.54$^{\circ }$to 29$^{\circ }$and 11 to 22 km/h on average. This systematic approach establishes UniAda as a proven technique for adversarial attacks on modern DL-based E2E ADSs. Jacky W. Keung, Yan Xiao 0002, Yihan Liao, Yishu Li |
IEEE Trans. Reliab. | 4 |