Shuo Liu 0020

dblp:07/6773-20 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-8877-3678ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 14 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving anomaly detection in software logs through hybrid language modeling and reduced reliance on parser
Yicheng Sun, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020, Hi Kuen Yu
Autom. Softw. Eng.4
2026 R2ComSync: improving code-comment synchronization with in-context learning and reranking
Zhen Yang 0022, Xiao Yu 0008, Jacky W. Keung, Shuo Liu 0020, Pak Yuen Patrick Chan, Yicheng Sun, Fengji Zhang
Empir. Softw. Eng.5
2026 An Empirical Study of Parameter-Efficient Fine-Tuning in Code Change Learning and Beyond
abstract
Compared to Full-Model Fine-Tuning (FMFT), Parameter-Efficient Fine-Tuning (PEFT) has demonstrated superior efficacy and efficiency in several code understanding tasks, owing to PEFT’s ability to alleviate the catastrophic forgetting issue of Pre-trained Language Models (PLMs) by updating only a small number of parameters. However, existing studies primarily involve static code comprehension, aligning with the pre-training paradigm of recent PLMs and facilitating knowledge transfer, but they do not account for dynamic code changes. Thus, it remains unclear whether PEFT outperforms FMFT in task-specific adaptation for code-change-related tasks.To address this question, we examine four prevalent PEFT methods (i.e., AT, LoRA, PT, and PreT) and compare their performance with FMFT across seven popular PLMs. In experiments, two widely studied code-change-related tasks, i.e., Just-In-Time Defect Prediction (JIT-DP) and Commit Message Generation (CMG) are involved, demonstrating that the four PEFT methods can surpass FMFT on JIT-DP but only exhibit comparable performances at best on CMG in common scenarios. While in cross-lingual and low-resource scenarios, they exhibit relative superiority. Afterward, a series of probing tasks from both static and dynamic perspectives are conducted in this paper, offering detailed explanations for the efficacy of PEFT and FMFT. Inspired by the distinctive advantages of PEFT and FMFT in their layer-wise probing results, we propose Pasta$k$, a self-adaPtive efficient layer-specific tuning framework for PLMs in code change learning, which combines FMFT and PEFT during the domain adaptation according to the guidance of probing results. Experiments in the CMG task demonstrate that Pasta$k$surpasses diverse PEFT methods in effectiveness. Even, Pasta$k$outperforms FMFT by 1.48%, 3.21%, and 1.87% at most in terms of BLEU, Meteor, and Rouge-L, while saving 26.26% and 20.65% in terms of training time and computational memory compared with FMFT.
Shuo Liu 0020, Jacky W. Keung, Zhi Jin 0001, Zhen Yang 0022, Fang Liu 0032, Hao Zhang 0145
IEEE Trans. Software Eng.1
2025 Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
abstract
The increasing demand for software development has driven interest in automating software engineering (SE) tasks using Large Language Models (LLMs). Recent efforts extend LLMs into multi-agent systems (MAS) that emulate collaborative development workflows, but these systems often fail due to three core deficiencies: under-specification, coordination misalignment, and inappropriate verification, arising from the absence of foundational SE structuring principles. This paper introduces Software Engineering Multi-Agent Protocol (SEMAP), a protocol-layer methodology that instantiates three core SE design principles for multi-agent LLMs: (1) explicit behavioral contract modeling, (2) structured messaging, and (3) lifecycleguided execution with verification, and is implemented atop Google’s Agent-to-Agent (A2A) infrastructure. Empirical evaluation using the Multi-Agent System Failure Taxonomy (MAST) framework demonstrates that SEMAP effectively reduces failures across different SE tasks. In code development, it achieves up to a $69.6 \%$ reduction in total failures for function-level development and $\mathbf{5 6 . 7 \%}$ for deployment-level development. For vulnerability detection, SEMAP reduces failure counts by up to $47.4 \%$ on Python tasks and $28.2 \%$ on $\mathrm{C} / \mathrm{C}++$ tasks.
Zhenyu Mao, Jacky W. Keung, Fengji Zhang, Shuo Liu 0020
APSEC4
2025 Beyond Log Parsers: A Scalable AI-Driven Framework for Efficient Log Anomaly Detection in Software Engineering
abstract
Log anomaly detection is critical for ensuring software system reliability and security, yet challenges persist in log parser dependency, small-scale dataset applicability, and hyperparameter tuning efficiency. Existing methods over-rely on predefined log templates, leading to information loss and high computational overhead. Additionally, anomaly detection models often struggle with limited log data, and hyperparameter tuning remains computationally expensive in dynamic environments. In this paper, we empirically evaluate seven state-of-the-art anomaly detection models across varied software systems, assessing the necessity of log parsers and model performance on small-scale datasets. Furthermore, we propose SMAC-, an enhanced real-time hyperparameter optimization framework, integrating stochastic gradient descent (SGD) and adaptive learning to improve model adaptability and efficiency. Our experiments on six benchmark datasets demonstrate that SMAC-achieves an overall average F1-score improvement of 4.27%, a 27.55% reduction in hyperparameter tuning time compared to other models, and a 1.35% increase in F1-score when adapting to newly emerging logs, compared to its counterpart without SGD integration. These findings underscore the practical advantages of AI-driven log analysis, providing valuable insights into scalable, software-engineered anomaly detection.
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Shuo Liu 0020, Yihan Liao
COMPSAC4
2025 Can Mamba Be Better? An Experimental Evaluation of Mamba in Code Intelligence
abstract
The Transformer architecture and its core attention mechanism form the foundation of Pre-trained Language Models (PLMs) and have driven their remarkable progress across a wide range of code intelligence tasks. However, the quadratic complexity inherent in the attention mechanism poses scalability challenges. Recently, sub-quadratic architectures such as Mamba and Mamba-2 have emerged as compelling alternatives to the Transformer. While they have shown promising results and attracted increasing academic interest, their effectiveness in code intelligence tasks has not yet been fully explored.To fill this gap, we present the first systematic empirical study of Mamba-based PLMs on three typical code tasks (i.e., code completion, code generation, and code clone detection), covering both the code comprehension and generation categories to delve into their effectiveness and efficiency. We first pre-train two Mamba-based PLMs on code based on Mamba and Mamba-2, respectively. Subsequently, we evaluate these four PLMs against typical Transformer-based PLMs (e.g., CodeGPT) with Full fine-Tuning (FT) and Parameter-Efficient Fine-Tuning (PEFT) settings, demonstrating the overall superiority of Mamba-based PLMs across all code tasks. Subsequent experiments involve the architecture analysis via pre-training from scratch to isolate the influence of the training corpora and low-resource analysis via deliberately limiting the fine-tuning data volume. All demonstrate the superiority of Mamba-based PLMs in both efficacy and efficiency. Finally, we also extend the sizes of PLMs to larger scales (7B at most) and make comparisons with more diverse PLMs/LLMs. Experimental results demonstrate that pre-training corpora and tasks also heavily affect the code modeling performance, apart from architectures. This work provides a comprehensive investigation into Mamba-based PLMs in the context of code intelligence, uncovering their strengths, limitations, and potential for future applications.
Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Zhenyu Mao, Yicheng Sun
ASE1
2025 Exploring continual learning in code intelligence with domain-wise distilled prompts
Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Fang Liu 0032, Fengji Zhang, Yicheng Sun
Inf. Softw. Technol.1
2025 SemiSMAC: A semi-supervised framework for log anomaly detection with automated hyperparameter tuning
abstract
Context: Logs generated during software operations are critical for system reliability and anomaly detection. However, their diversity, the scarcity of labeled data, and hyperparameter tuning challenges hinder traditional detection methods. Objective: This paper presents SemiSMAC, a novel semi-supervised framework that leverages the Large Language Model for log parsing and grouping, combined with Sequential Model-based Algorithm Configuration (SMAC) for hyperparameter optimization to enhance anomaly detection. Method: In this work, we leverage ChatGPT for log parsing and introduce a novel log grouping approach. This grouping process requires only a small number of labeled samples, which ChatGPT uses to generate pseudo-labels for the remaining data, thereby expanding the training set. Furthermore, SemiSMAC utilizes a Sequential Model-based Algorithm Configuration (SMAC) to automatically optimize the hyperparameters of the embedded models. This integration leads to consistent performance improvements, particularly in resource-constrained environments. Results: SemiSMAC-LSTM, which uses LSTM as the backbone of the SemiSMAC framework, demonstrates superior performance in experiments on four widely used datasets. It outperforms six benchmark models, including three supervised learning models. In low-resource scenarios, SemiSMAC-LSTM exhibits exceptional robustness, showcasing its effectiveness in handling challenging detection tasks. Conclusion: SemiSMAC demonstrates its potential to revolutionize anomaly detection in both large-scale and low-resource datasets. Its ability to deliver outstanding performance makes it a valuable tool for scalable and automated anomaly detection in real-world applications, paving the way for more reliable and scalable software engineering practices
Yicheng Sun, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020, Yihan Liao
Inf. Softw. Technol.4
2025 SemiRALD: A semi-supervised hybrid language model for robust Anomalous Log Detection
abstract
Deep learning-based Anomalous Log Detection (DALD) tools are critical for software reliability, but current approaches face challenges, including information loss during log parsing, reliance on large labeled datasets, and fragility in low-resource scenarios. To overcome the above limitations, we propose SemiRALD, a semi-supervised learning-based robust ALD approach that leverages Large Language Model (LLM) for log parsing, enhancing both flexibility and accuracy. It utilizes a hybrid language model to repeatedly fit the samples with generate pseudo-labels, thereby training DALD models with limited resources and facilitating efficient anomaly detection tasks. In detail, SemiRALD utilizes ChatGPT and in-context learning for automated log parsing, thereby improving the log integrity during log parsing. Subsequently, it harnesses a semi-supervised learning framework and our proposed hybrid language model to remedy the performance degeneration caused by low-resource restriction in practice. Semi-supervised learning requires only a small amount of labeled data throughout the entire process, while the hybrid language model is built on the architecture of RoBERTa and an attention-based BiLSTM. Experiments on the HDFS and BGL datasets demonstrate that SemiRALD achieves an average F1-score improvement of 7.3% and 8.2%, respectively, over seven benchmark models. On small-scale datasets (0.1% of the original size), SemiRALD outperforms competitors by 31.4% and 46.0% in F1-score, respectively. Its consistent performance across diverse datasets highlights its generalizability and robustness. SemiRALD is capable of handling anomaly detection tasks in both large-scale and low-resource datasets, delivering significant advancements in anomaly log detection and offering robust, adaptable solutions to address prevalent challenges in the field of software reliability engineering.
Yicheng Sun, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020, Hi Kuen Yu
Inf. Softw. Technol.4
2024 Unveiling Hidden Anomalies: Leveraging SMAC-LSTM for Enhanced Software Log Analysis
abstract
Software logs are essential records generated during the functioning of software systems, aiding in the identification of irregularities and prevention of system failures. Recently, deep learning models have garnered significant interest among researchers due to their efficacy in detecting anomalies within software logs. This research paper constructs a novel dataset, consisting of three parts: two datasets derived from our software system, along with a publicly available dataset obtained from the LogHub platform. The extensive logs within the dataset undergo preprocessing to extract meaningful features. Furthermore, this study introduces a novel model named SMAC-LSTM, designed specifically for detecting anomalies in software logs. Sequential Model-based Algorithm Configuration (SMAC) is a suitable method for hyperparameter optimization and automated deep learning. SMAC-LSTM involves determining the optimal hyperparameter values for the LSTM model using the SMAC. Additionally, SMAC-LSTM combines the temporal dependency capturing ability of Long Short-Term Memory (LSTM) with a context-dependent mechanism achieved through a Bayesian optimization algorithm based on random forests. This fusion enhances the model's ability to detect subtle anomalies in time series data, which are frequently disregarded by con-ventional LSTM models. The thorough evaluation demonstrates the superior performance of SMAC-LSTM models compared to traditional deep learning models, showcasing significant enhance-ments in precision (98.63%), and recall (92.31%), with an F1-Score of 95.36%, outperforming all other models. These results underscore the potential of SMAC-LSTM in the realm of software log anomaly detection.
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Wenqiang Luo, Shuo Liu 0020
COMPSAC6
2024 Delving into Parameter-Efficient Fine-Tuning in Code Change Learning: An Empirical Study
abstract
Compared to Full-Model Fine-Tuning (FMFT), Parameter Efficient Fine-Tuning (PEFT) has demonstrated superior performance and lower computational overhead in several code understanding tasks, such as code summarization and code search. This advantage can be attributed to PEFT's ability to alleviate the catastrophic forgetting issue of Pre-trained Language Models (PLMs) by updating only a small number of parameters. As a result, PEFT effectively harnesses the pre-trained general-purpose knowledge for downstream tasks. However, existing studies primarily involve static code comprehension, aligning with the pre-training paradigm of recent PLMs and facilitating knowledge transfer, but they do not account for dynamic code changes. Thus, it remains unclear whether PEFT outperforms FMFT in task-specific adaptation for code-change-related tasks. To address this question, we examine two prevalent PEFT methods, namely Adapter Tuning (AT) and Low-Rank Adaptation (LoRA), and compare their performance with FMFT on five popular PLMs. Specifically, we evaluate their performance on two widely-studied code-change-related tasks: Just-In-Time Defect Prediction (JIT-DP) and Commit Message Generation (CMG). The results demonstrate that both AT and LoRA achieve state-of-the-art (SOTA) results in JIT-DP and exhibit comparable performances in CMG when compared to FMFT and other SOTA approaches. Furthermore, AT and LoRA exhibit superiority in cross-lingual and low-resource scenarios. We also conduct three probing tasks to explain the efficacy of PEFT techniques on JIT-DP and CMG tasks from both static and dynamic perspectives. The study indicates that PEFT, particularly through the use of AT and LoRA, offers promising advantages in code-change-related tasks, surpassing FMFT in certain aspects. This research contributes to a deeper understanding of the capabilities of PEFT in leveraging pre-trained PLMs for dynamic code changes. The replication package is available at https://github.com/ishuoliu/PEFT4CC.
Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Fang Liu 0032, Qilin Zhou, Yihan Liao
SANER1
2024 Co-clustering for Federated Recommender System
abstract
As data privacy and security attract increasing attention, Federated Recommender System (FRS) offers a solution that strikes a balance between providing high-quality recommendations and preserving user privacy. However, the presence of statistical heterogeneity in FRS, commonly observed due to personalized decision-making patterns, can pose challenges. To address this issue and maximize the benefit of collaborative filtering (CF) in FRS, it is intuitive to consider clustering clients (users) as well as items into different groups and learning group-specific models. Existing methods either resort to client clustering via user representations-risking privacy leakage, or employ classical clustering strategies on item embeddings or gradients, which we found are plagued by the curse of dimensionality. In this paper, we delve into the inefficiencies of the K-Means method in client grouping, attributing failures due to the high dimensionality as well as data sparsity occurring in FRS, and propose CoFedRec, a novel Co-clustering Federated Recommendation mechanism, to address clients heterogeneity and enhance the collaborative filtering within the federated framework. Specifically, the server initially formulates an item membership from the client-provided item networks. Subsequently, clients are grouped regarding a specific item category picked from the item membership during each communication round, resulting in an intelligently aggregated group model. Meanwhile, to comprehensively capture the global inter-relationships among items, we incorporate an additional supervised contrastive learning term based on the server-side generated item membership into the local training phase for each client. Extensive experiments on four datasets are provided, which verify the effectiveness of the proposed CoFedRec.
Xinrui He, Shuo Liu 0020, Jacky W. Keung, Jingrui He
WWW2
2024 SimAC: simulating agile collaboration to generate acceptance criteria in user story elaboration
Yishu Li, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020
Autom. Softw. Eng.6
2024 Improving domain-specific neural code generation with few-shot meta-learning
Zhen Yang 0022, Jacky W. Keung, Zeyu Sun 0004, Yunfei Zhao 0003, Ge Li 0001, Zhi Jin 0001, Shuo Liu 0020, Yishu Li
Inf. Softw. Technol.7
2024 TerGEC: A graph enhanced contrastive approach for program termination analysis
Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Yihan Liao, Yishu Li
Sci. Comput. Program.1