EDBT 2026 Demo / reviewers in the wild / expert
Hi Kuen Yu
dblp:384/5839
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0009-8451-188XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where Do Expectations Diverge? an Empirical Analysis of Software Engineering Graduate Preparedness in Industry
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Yihan Liao, Yishu Li |
COMPSAC | 3 |
| 2026 | Designing Psychologically Safe AI Tutors for Students: An Emotion-Aware Post-Hoc Intervention for LLM-Assisted Learning
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Yihan Liao, Zhenyu Mao, Yishu Li |
COMPSAC | 3 |
| 2026 | Artifact-Constrained Agentic Testing for Black-Box System Testing under Actuarial and Regulatory ConstraintsabstractSystem-level testing of insurance software systems is challenging due to long-lived legacy architectures, distributed actuarial logic, regulatory-driven conditional behavior, and exception-heavy workflows, where correctness is defined by compliance with evolving domain constraints rather than deterministic outputs. Existing approaches either focus on artifact-centric validation, rely heavily on human-driven execution and diagnosis, or generate test artifacts in a one-shot manner without sustained adaptation during execution, limiting robustness and reproducibility at the system level. This paper proposes \emph{Artifact-Constrained Agentic Testing} (ACAT), a multi-agent framework that structures black-box system testing as a closed-loop process of planning, execution, diagnosis, and repair, in which LLM-based agents operate under explicitly bounded capabilities and interact with the system under test exclusively through tool-mediated execution. We evaluate ACAT through a controlled industrial study on an insurance broker management system using 50 real-world use cases. The results show that artifact-constrained agentic testing significantly improves test executability and execution stability compared to human-driven testing, while exposing a complementary subset of system-level failures. These findings suggest that agent-assisted testing can enhance system-level automation in complex, regulated software systems without replacing human expertise. Hi Kuen Yu, Jacky W. Keung, Man On Wong, Yicheng Sun, Rachel Samantha Chandra, Yihan Liao |
COMPSAC | 1 |
| 2026 | Improving anomaly detection in software logs through hybrid language modeling and reduced reliance on parser
Yicheng Sun, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020, Hi Kuen Yu |
Autom. Softw. Eng. | 5 |
| 2026 | Industrial log analysis revisited: A task-oriented evaluation of parsing and anomaly detection under real-world constraints
Yicheng Sun, Jacky W. Keung, Yihan Liao, Zhenyu Mao, Hi Kuen Yu |
Inf. Softw. Technol. | 6 |
| 2026 | LogMeta: A few-shot model-agnostic meta-learning framework for robust and adaptive log anomaly detectionabstractContext: Log anomaly detection is critical for maintaining the security, stability, and operational efficiency of modern software systems, especially as they generate vast and diverse log data. However, existing deep learning models struggle with the challenges of heterogeneous log formats across systems and the scarcity of labeled anomaly logs, limiting their real-world deployment and generalization capabilities. Objective: To address these challenges, we propose LogMeta, a novel semi-supervised framework designed for adaptive and efficient log anomaly detection in diverse and low-resource environments. Method: LogMeta integrates Model-Agnostic Meta-Learning (MAML) with a hybrid language model to address key challenges. MAML enables LogMeta to rapidly adapt to unseen log systems using few-shot samples, while the hybrid model combines RoBERTa for extracting semantic representations with Bi-LSTM and attention mechanisms to capture sequential dependencies and critical features within log sequences. This design reduces reliance on large-scale labeled datasets and enhances adaptability in heterogeneous environments. Results: Experimental evaluations on multiple benchmark datasets demonstrate that LogMeta consistently outperforms state-of-the-art supervised and unsupervised methods, achieving up to a 28.3% improvement in F1-scores under low-resource scenarios compared to other models. Furthermore, LogMeta exhibits exceptional domain transfer capabilities, maintaining robust performance across diverse log datasets with minimal fine-tuning. In terms of efficiency, LogMeta achieves competitive training and inference times, making it suitable for real-time anomaly detection in large-scale systems. Conclusion: LogMeta provides a scalable and practical solution for real-world log anomaly detection, overcoming challenges related to data heterogeneity and label scarcity. Its strong generalization capabilities, minimal supervision requirements, and adaptability to new log systems make it a promising tool for enhancing software system reliability and security. © 2026 The Author(s). Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Wenqiang Luo |
J. Syst. Softw. | 3 |
| 2025 | Understanding Industrial Log Analysis: A Multi-Dataset Evaluation of Parsing and Anomaly DetectionabstractLog analysis plays a critical role in monitoring and maintaining the safety of industrial software systems. However, most existing research relies heavily on benchmark datasets derived from legacy or open-source systems, which fail to capture the structural diversity and operational complexity of real-world industrial logs. In this study, we present a comprehensive empirical evaluation of log parsing and anomaly detection models across four diverse datasets, including three collected from largescale industrial software deployed in manufacturing, process control, and energy monitoring environments. Our analysis reveals that state-of-the-art models—particularly rule-based parsers and supervised detectors—experience substantial performance degradation when applied to industrial settings. To address this gap, we introduce a unified evaluation framework using representative training subsets, and we highlight the effectiveness of semisupervised and LLM-based approaches in handling heterogeneous, low-resource log environments. The findings offer practical insights into the limitations of current log analysis techniques and suggest design principles for building more robust, domain-adaptive solutions for industrial software risk mitigation. Yicheng Sun, Jacky W. Keung, Yihan Liao, Hi Kuen Yu |
APSEC | 4 |
| 2025 | Towards Lightweight LLM Software Solutions for InsurTech: A Framework for Scalable Question Answering SystemsabstractThe integration of Large Language Models (LLMs) into software systems is transforming regulated sectors like insurance, where precision, compliance, and efficiency are essential. While proprietary LLMs like GPT-4 offer state-of-the-art performance, their closed-source nature and high computational demands constrain adoption in privacy-sensitive and cost-restricted InsurTech environments. In response, this paper investigates how lightweight, open-source LLMs can be effectively deployed for domain-specific question answering in insurance, emphasizing software engineering considerations such as modularity, inference stability, and prompt orchestration. We propose a software-engineered evaluation framework tailored to insurance-related tasks, featuring modular prompt management, automated rubricbased evaluation, and backend support for reproducibility and compliance tracking. A curated benchmark dataset derived from the Hong Kong Insurance Intermediaries Qualifying Examination (IIQE) is constructed to reflect real-world regulatory and operational challenges. Ten open-source models are systematically evaluated across four question types using both standard and Chain-of-Thought (CoT) prompting strategies. Our findings show that compact models such as DeepSeek-R1-1.5B achieve strong accuracy with minimal resource consumption, making them suitable for practical deployment. CoT prompting further enhances reasoning performance, particularly for models with 3B parameters or more. With proper prompt design and modular deployment, lightweight LLMs can support secure, efficient, and interpretable InsurTech applications, enabling trustworthy AI-driven software systems in regulated domains. Hi Kuen Yu, Jacky W. Keung, Yicheng Sun, Yihan Liao, Richard Suen |
APSEC | 1 |
| 2025 | Beyond Log Parsers: A Scalable AI-Driven Framework for Efficient Log Anomaly Detection in Software EngineeringabstractLog anomaly detection is critical for ensuring software system reliability and security, yet challenges persist in log parser dependency, small-scale dataset applicability, and hyperparameter tuning efficiency. Existing methods over-rely on predefined log templates, leading to information loss and high computational overhead. Additionally, anomaly detection models often struggle with limited log data, and hyperparameter tuning remains computationally expensive in dynamic environments. In this paper, we empirically evaluate seven state-of-the-art anomaly detection models across varied software systems, assessing the necessity of log parsers and model performance on small-scale datasets. Furthermore, we propose SMAC-, an enhanced real-time hyperparameter optimization framework, integrating stochastic gradient descent (SGD) and adaptive learning to improve model adaptability and efficiency. Our experiments on six benchmark datasets demonstrate that SMAC-achieves an overall average F1-score improvement of 4.27%, a 27.55% reduction in hyperparameter tuning time compared to other models, and a 1.35% increase in F1-score when adapting to newly emerging logs, compared to its counterpart without SGD integration. These findings underscore the practical advantages of AI-driven log analysis, providing valuable insights into scalable, software-engineered anomaly detection. Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Shuo Liu 0020, Yihan Liao |
COMPSAC | 3 |
| 2025 | StuLAC: An Adaptive LLM-Driven Framework for Scalable Student Feedback Analysis in Software-Driven Educational SystemsabstractWith the growing scalability challenges in higher education, automated student feedback analysis has become crucial for course evaluation and pedagogical improvements. However, traditional methods struggle to handle mixed sentiments, adapt to evolving feedback trends, and maintain computational efficiency. To address these challenges, we propose StuLAC, a Software Engineering-driven framework that integrates Large Language Models (LLMs) with Adaptive Template-Based Caching (ATC). StuLAC employs hierarchical matching for fine-grained classification and dynamically updates feedback templates through context-aware cache refinement. Empirical results on 80,000 student feedback entries demonstrate that StuLAC-generated summaries improve overall quality by 10.5% compared to manually generated reports, while also achieving faster processing times. Additionally, StuLAC attains an 86.4% accuracy and an 86.24% F1-score in sentiment detection. StuLAC’s Feedback Summary Generation provides actionable insights that enhance data-driven decision-making in educational settings. These findings establish StuLAC as a scalable and adaptive solution for improving AI-driven educational feedback systems. Yicheng Sun, Hi Kuen Yu, Jacky W. Keung, Yuchen Cao 0006, Yihan Liao |
COMPSAC | 2 |
| 2025 | SemiRALD: A semi-supervised hybrid language model for robust Anomalous Log DetectionabstractDeep learning-based Anomalous Log Detection (DALD) tools are critical for software reliability, but current approaches face challenges, including information loss during log parsing, reliance on large labeled datasets, and fragility in low-resource scenarios. To overcome the above limitations, we propose SemiRALD, a semi-supervised learning-based robust ALD approach that leverages Large Language Model (LLM) for log parsing, enhancing both flexibility and accuracy. It utilizes a hybrid language model to repeatedly fit the samples with generate pseudo-labels, thereby training DALD models with limited resources and facilitating efficient anomaly detection tasks. In detail, SemiRALD utilizes ChatGPT and in-context learning for automated log parsing, thereby improving the log integrity during log parsing. Subsequently, it harnesses a semi-supervised learning framework and our proposed hybrid language model to remedy the performance degeneration caused by low-resource restriction in practice. Semi-supervised learning requires only a small amount of labeled data throughout the entire process, while the hybrid language model is built on the architecture of RoBERTa and an attention-based BiLSTM. Experiments on the HDFS and BGL datasets demonstrate that SemiRALD achieves an average F1-score improvement of 7.3% and 8.2%, respectively, over seven benchmark models. On small-scale datasets (0.1% of the original size), SemiRALD outperforms competitors by 31.4% and 46.0% in F1-score, respectively. Its consistent performance across diverse datasets highlights its generalizability and robustness. SemiRALD is capable of handling anomaly detection tasks in both large-scale and low-resource datasets, delivering significant advancements in anomaly log detection and offering robust, adaptable solutions to address prevalent challenges in the field of software reliability engineering. Yicheng Sun, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020, Hi Kuen Yu |
Inf. Softw. Technol. | 5 |
| 2024 | Unveiling Hidden Anomalies: Leveraging SMAC-LSTM for Enhanced Software Log AnalysisabstractSoftware logs are essential records generated during the functioning of software systems, aiding in the identification of irregularities and prevention of system failures. Recently, deep learning models have garnered significant interest among researchers due to their efficacy in detecting anomalies within software logs. This research paper constructs a novel dataset, consisting of three parts: two datasets derived from our software system, along with a publicly available dataset obtained from the LogHub platform. The extensive logs within the dataset undergo preprocessing to extract meaningful features. Furthermore, this study introduces a novel model named SMAC-LSTM, designed specifically for detecting anomalies in software logs. Sequential Model-based Algorithm Configuration (SMAC) is a suitable method for hyperparameter optimization and automated deep learning. SMAC-LSTM involves determining the optimal hyperparameter values for the LSTM model using the SMAC. Additionally, SMAC-LSTM combines the temporal dependency capturing ability of Long Short-Term Memory (LSTM) with a context-dependent mechanism achieved through a Bayesian optimization algorithm based on random forests. This fusion enhances the model's ability to detect subtle anomalies in time series data, which are frequently disregarded by con-ventional LSTM models. The thorough evaluation demonstrates the superior performance of SMAC-LSTM models compared to traditional deep learning models, showcasing significant enhance-ments in precision (98.63%), and recall (92.31%), with an F1-Score of 95.36%, outperforming all other models. These results underscore the potential of SMAC-LSTM in the realm of software log anomaly detection. Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Wenqiang Luo, Shuo Liu 0020 |
COMPSAC | 4 |