VLDB 2026 Research / reviewers in the wild / expert
Yishu Li
dblp:247/2570
· DBLP profile ↗
19ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0003-4017-4294ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where Do Expectations Diverge? an Empirical Analysis of Software Engineering Graduate Preparedness in Industry
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Yihan Liao, Yishu Li |
COMPSAC | 5 |
| 2026 | Designing Psychologically Safe AI Tutors for Students: An Emotion-Aware Post-Hoc Intervention for LLM-Assisted Learning
Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Yihan Liao, Zhenyu Mao, Yishu Li |
COMPSAC | 6 |
| 2026 | R2Code: A Self-Reflective LLM Framework for Requirements-to-Code TraceabilityabstractAccurate requirement-to-code traceability is crucial for software maintenance. However, existing IR- and embedding-based methods are heavily dependent on lexical similarity, often yielding incomplete or inconsistent links across projects and languages and incurring high cost from long-context retrieval and prompting. This paper presents R2Code, an LLM-based semantic traceability framework designed to improve trace link accuracy while reducing inference cost. R2Code integrates three components: 1) a decomposition-enhanced Bidirectional Alignment Network (BAN) that aligns four-layer requirement semantics with corresponding code structures to support cross-level semantic matching; 2) a Self-Reflective Consistency Verification (SRCV) module that conducts explanation-guided consistency checking to calibrate link reliability; and 3) a Dynamic Context-Adaptive Retrieval (DCAR) mechanism that adjusts retrieval granularity and filters contexts using semantic-overlap weighting for efficient context utilization. Experiments on five public datasets spanning multiple domains and two programming languages demonstrate that R2Code consistently outperforms the strongest baselines, achieving an average F1 gain of 7.4%, while reducing token consumption by up to 41.7% through adaptive context control. Jacky W. Keung, Zhenyu Mao, Kehui Chen, Yishu Li |
COMPSAC | 6 |
| 2025 | Practitioners' Expectations on Log Anomaly DetectionabstractLog anomaly detection has become a common practice for software engineers to analyze software system behavior. Despite significant research efforts in log anomaly detection over the past decade, it remains unclear what are practitioners’ expectations on log anomaly detection and whether current research meets their needs. To fill this gap, we conduct an empirical study, surveying 312 practitioners from 36 countries about their expectations on log anomaly detection. In particular, we investigate various factors influencing practitioners’ willingness to adopt log anomaly detection tools. We then perform a literature review on log anomaly detection, focusing on publications in premier venues from 2015 to 2025, to compare practitioners’ needs with the current state of research. Based on this comparison, we highlight the directions for researchers to focus on to develop log anomaly detection techniques that better meet practitioners’ expectations. Yishu Li, Jacky W. Keung, Xiao Yu 0008, Huiqi Zou, Zhen Yang 0022, Federica Sarro, Earl T. Barr |
IEEE Trans. Software Eng. | 2 |
| 2025 | On the Influence of Data Resampling for Deep Learning-Based Log Anomaly Detection: Insights and RecommendationsabstractNumerous Deep Learning (DL)-based approaches have gained attention in software Log Anomaly Detection (LAD), yet class imbalance in training data remains a challenge, with anomalies often comprising less than 1% of datasets like Thunderbird. Existing DLLAD methods may underperform in severely imbalanced datasets. Although data resampling has proven effective in other software engineering tasks, it has not been explored in LAD. This study aims to fill this gap by providing an in-depth analysis of the impact of diverse data resampling methods on existing DLLAD approaches from two distinct perspectives. Firstly, we assess the performance of these DLLAD approaches across four datasets with different levels of class imbalance, and we explore the impact of resampling ratios of normal to abnormal data on DLLAD approaches. Secondly, we evaluate the effectiveness of the data resampling methods when utilizing optimal resampling ratios of normal to abnormal data. Our findings indicate that oversampling methods generally outperform undersampling and hybrid sampling methods. Data resampling on raw data yields superior results compared to data resampling in the feature space. These improvements are attributed to the increased attention given to important tokens. By exploring the resampling ratio of normal to abnormal data, we suggest generating more data for minority classes through oversampling while removing less data from majority classes through undersampling. In conclusion, our study provides valuable insights into the intricate relationship between data resampling methods and DLLAD. By addressing the challenge of class imbalance, researchers and practitioners can enhance DLLAD performance. Huiqi Zou, Pinjia He, Jacky W. Keung, Yishu Li, Xiao Yu 0008, Federica Sarro |
IEEE Trans. Software Eng. | 5 |
| 2024 | Enhancing the Transferability of Adversarial Attacks for End-to-End Autonomous Driving SystemsabstractAdversarial attacks play an important role in testing and enhancing the reliability of deep learning (DL) systems. Most existing attacks for DL-based autonomous driving systems (ADSs) demonstrate strong performance under the white-box setting but struggle with black-box transferability, while blackbox attacks are more practical in real-world scenarios as they operate without full model access. Numerous transferabilityenhancement techniques have been proposed in other fields (e.g., image classification), however, they remain unexplored for endtoend (E2E) ADSs. Our study fills the gap by conducting the first comprehensive empirical analysis of nine transferability-enhancement methods on E2E ADSs, covering two types: three input transformation enhancements and six attack objective enhancements. We evaluate their effectiveness on two datasets with four steering models. Our findings reveal that, out of nine enhancements, Resizing+ Translation delivers the best black-box transferability, producing up to 9.39° increase in MAE. Pred+Attn serves as the best objective enhancement, producing a maximum of 5.55° (white-box) and 6.21° (black-box) increase in MAE. Through attention heatmap visualizations, we discover that different models focus on similar regions when predicting, thereby enhancing the transferability of attention-based attacks. In conclusion, our study provides valuable results and insights into the transferability-enhancement techniques for E2E ADSs, which also serve as a robust benchmark for further advancements in the autonomous driving field. Jacky W. Keung, Yihan Liao, Yishu Li, Yicheng Sun |
APSEC | 5 |
| 2024 | Agile Requirements Engineering in a Distributed Environment: Experiences from the Software Industry During Unprecedented Global ChallengesabstractUnprecedented global challenges such as the COVID-19 pandemic necessitated a widespread transition to Work-From-Home (WFH) arrangements for project teams, posing significant challenges in conveying requirements within agile Requirements Engineering (RE). While numerous studies have examined the impact of transitioning work routines during the pandemic, limited research exists on the specific challenges of agile RE operating within the WFH context. Given the pervasive shift in the software development ecosystem worldwide, where WFH is projected to persist even in the post-COVID era, it is imperative to ascertain the challenges associated with WFH-based agile RE. During the pandemic, we collaborated with startups to conduct an industry-academia project. By adopting the methodology of action research, this study comprehensively analyzed agile RE practices and reported the key challenges encountered within the WFH context. To mitigate these challenges, several collaborative RE techniques were employed in three intervention cycles. Interviews were conducted to thoroughly analyze the results. This study also provides insights into collaborative RE techniques and valuable lessons learned. Considering the increasing prevalence of WFH as a working mode in the post-pandemic era, this study equips the community with practical strategies to navigate agile RE challenges and better prepare for unprecedented challenges in the future. Yishu Li, Jacky W. Keung, Kwabena Ebo Bennin, Zhen Yang 0022 |
COMPSAC | 1 |
| 2024 | LLM-Based Class Diagram Derivation from User Stories with Chain-of-Thought PromptingsabstractIn agile requirements engineering, user stories are the primary means of capturing project requirements. However, deriving conceptual models, such as class diagrams, from user stories requires significant manual effort. This paper explores the potential of leveraging Large Language Models (LLMs) and a tailored Chain-of- Thought (CoT) prompting technique to automate this task. We conducted a comprehensive preliminary study to investigate different prompting techniques applied to the task. The study involved comparing LLM-based approaches with guided and unguided human extraction to evaluate the effectiveness of the proposed LLM-based techniques. Our findings demonstrate that LLM-based approaches, particularly when combined with well-crafted few-shot prompts, outperform guided human extraction in identifying classes. However, we also identified areas of suboptimal performance through qualitative analysis. The proposed CoT prompting technique offers a promising pathway to automate the derivation of class diagrams in agile projects, reducing the reliance on manual effort. Our study contributes valuable insights and directions for future research in this field. Yishu Li, Jacky W. Keung, Chun Yong Chong, Yihan Liao |
COMPSAC | 1 |
| 2024 | Enhancing Valid Test Input Generation with Distribution Awareness for Deep Neural NetworksabstractComprehensive testing is important in improving the reliability of Deep Learning (DL)-based systems. Various Test Input Generators (TIGs) have been proposed to generate misbehavior-inducing test inputs. However, the lack of validity checking in TIGs often results in the generation of invalid inputs (i.e., out of the learned distribution), leading to unreliable testing. To save the effort of manually checking the validity and improve test efficiency, it is important to assess the effectiveness and reliability of automated validators. In this study, we comprehensively assess four automated Input Validators (IV s), Our findings show that the accuracy of IVs ranges from 49% to 77%. Distance-based IVs generally outperform reconstruction-based and density-based IVs for both classification and regression tasks. Based on the findings, we enhance existing testing frameworks by incorporating distribution awareness through joint optimization. The results demonstrate our framework leads to a 2 % to 10% increase in the number of valid inputs, which establishes our method as an effective technique for valid test input generation. Jacky W. Keung, Yan Xiao 0002, Yishu Li, Wing Kwong Chan |
COMPSAC | 6 |
| 2024 | MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane ReconstructionabstractThis paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and establishes a plane reconstruction pipeline based on monocular geometric cues, resulting in accurate, robust and scalable 3D plane detection and reconstruction in the wild. Specifically, we first leverage large-scale pre-trained neural networks to obtain the depth and surface normals from a single image. These monocular geometric cues are then incorporated into a proximity-guided RANSAC framework to sequentially fit each plane instance. We exploit effective 3D point proximity and model such proximity via a graph within RANSAC to guide the plane fitting from noisy monocular depths, followed by image-level multi-plane joint optimization to improve the consistency among all plane instances. We further design a simple but effective pipeline to extend this single-view solution to sparse-view 3D plane reconstruction. Extensive experiments on a list of datasets demonstrate our superior zero-shot generalizability over baselines, achieving state-of-the-art plane reconstruction performance in a transferring setting. Our code is available at https://github.com/thuzhaowang/MonoPlane. Wang Zhao 0001, Yishu Li, Sili Chen, Sharon X. Huang, Yong-Jin Liu 0001, Hengkai Guo |
IROS | 4 |
| 2024 | SimAC: simulating agile collaboration to generate acceptance criteria in user story elaboration
Yishu Li, Jacky W. Keung, Zhen Yang 0022, Shuo Liu 0020 |
Autom. Softw. Eng. | 1 |
| 2024 | Improving domain-specific neural code generation with few-shot meta-learning
Zhen Yang 0022, Jacky W. Keung, Zeyu Sun 0004, Yunfei Zhao 0003, Ge Li 0001, Zhi Jin 0001, Shuo Liu 0020, Yishu Li |
Inf. Softw. Technol. | 8 |
| 2024 | TerGEC: A graph enhanced contrastive approach for program termination analysis
Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Yihan Liao, Yishu Li |
Sci. Comput. Program. | 5 |
| 2024 | A Semisupervised Approach for Industrial Anomaly Detection via Self-Adaptive ClusteringabstractWith the rapid development of the Industrial Internet of Things, log-based anomaly detection has become vital for smart industrial construction that has prompted many researchers to contribute. To detect anomalies based on log data, semisupervised approaches stand out from supervised and unsupervised approaches because they only require a portion of labeled data and are relatively stable. However, the state-of-the-art semisupervised approaches still suffer from two main problems: manual parameter setting and unsatisfactory performance with high false positives. We propose AdaLog, an integrated semisupervised approach based on self-adaptive clustering, for industrial anomaly detection. In particular, the clustering step performs automatic label probability estimation by distinguishing 12 situations so that the label probability of each unlabeled data can be carefully calculated, leading to high accuracy. In addition, AdaLog employs a pretrained model to learn contextual information comprehensively and a transformer-based model to detect anomalies efficiently. To alleviate class imbalance, an undersampling method is incorporated. The results on three popular datasets demonstrate that AdaLog significantly outperforms three state-of-the-art semisupervised approaches by 17.8%–2489.8% on average in terms of F1-score, and is even superior to two supervised approaches in most cases with average improvements of 10.9%–23.8%. Jacky W. Keung, Pinjia He, Yan Xiao 0002, Xiao Yu 0008, Yishu Li |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | UniAda: Universal Adaptive Multiobjective Adversarial Attack for End-to-End Autonomous Driving SystemsabstractAdversarial attacks play a pivotal role in testing and improving the reliability of deep learning (DL) systems. Existing literature has demonstrated that subtle perturbations to the input can elicit erroneous outcomes, thereby substantially compromising the security of DL systems. This has emerged as a critical concern in the development of DL-based safety–critical systems like autonomous driving systems (ADSs). The focus of existing adversarial attack methods on end-to-end (E2E) ADSs has predominantly centered on misbehaviors of steering angle, which overlooks speed-related controls or imperceptible perturbations. To address these challenges, we introduce UniAda–a multiobjective white-box attack technique with a core function that revolves around crafting an image-agnostic adversarial perturbation capable of simultaneously influencing both steering and speed controls. UniAda capitalizes on an intricately designed multiobjective optimization function with the adaptive weighting scheme (AWS), enabling the concurrent optimization of diverse objectives. Validated with both simulated and real-world driving data, UniAda outperforms five benchmarks across two metrics, inducing steering and speed deviations from 3.54$^{\circ }$to 29$^{\circ }$and 11 to 22 km/h on average. This systematic approach establishes UniAda as a proven technique for adversarial attacks on modern DL-based E2E ADSs. Jacky W. Keung, Yan Xiao 0002, Yihan Liao, Yishu Li |
IEEE Trans. Reliab. | 5 |
| 2023 | Towards Requirements Engineering Activities for Machine Learning-Enabled FinTech ApplicationsabstractThe complexity required in the software development of machine learning (ML) applications introduces additional challenges to requirement engineering (RE) activities. RE researchers expressed concerns and the need for more discussions on RE for ML, requiring additional real-world case studies to evaluate RE activities for practical ML-enabled applications. This study aims to observe the RE activities for ML-enabled systems in a real-world context, taking action research in the ML-enabled FinTech project where the RE activities are being adjusted by engaging the data scientists to help and clarify ML-related requirements. This paper discussed the difficulties of RE activities from the perspectives of the data scientist and requirement engineer. Considering data and model relevance in developing the ML-enabled FinTech application, a RE framework iteratively made active changes according to the parameters is proposed, which includes the selected ML-related requirement characteristics to pursue and complete RE activities for ML-enabled application development. The feedback from the practitioners indicates that such practices address the difficulties of improving data quality and verifying model requirements in RE activities. The lessons learned by researchers and practitioners are also presented, which provides practical suggestions to the SE and RE communities with similar concerns in the related context. Yishu Li, Jacky W. Keung, Kwabena Ebo Bennin, Yangyang Huang |
APSEC | 1 |
| 2023 | AttSum: A Deep Attention-Based Summarization Model for Bug Report Title GenerationabstractConcise and precise bug report titles help software developers to capture the highlights of the bug report quickly. Unfortunately, it is common that bug reporters do not create high-quality bug report titles. Recent long short-term memory (LSTM)-based sequence-to-sequence models such as iTAPE were proposed to generate bug report titles automatically, but the text representation method and LSTM employed in such model are difficult to capture the accurate semantic information and draw the global dependencies among tokens effectively. This article proposes a deep attention-based summarization model (i.e.,AttSum) to generate high-quality bug report titles. Specifically, theAttSummodel employs the encoder.decoder framework, which utilizes the robustly optimized bidirectional-encoder-representations-from-transformers approach to encode the bug report bodies to capture contextual semantic information better, the stacked transformer decoder to automatically generate titles, and the copy mechanism to handle the rare token problem. To validate the effectiveness ofAttSum, we conduct automatic and manual evaluations on 333563 “$< body, title>$” pairs of bug reports and perform a practical analysis of its ability to improve low-quality titles. The result shows thatAttSumis superior to the state-of-the-art baselines by a substantial margin both on automatic evaluation metrics (e.g., by 3.4%–58.8% and 7.7%–42.3% in terms of recall-oriented understudy for gisting evaluation in F1 and bilingual evaluation understudy, separately) and three human-set modalities (e.g., by 1.9%–57.5%). Moreover, we analyze the impact of the training data size onAttSumand the results imply that our approach is robust enough to generate much better titles. Jacky W. Keung, Xiao Yu 0008, Huiqi Zou, Yishu Li |
IEEE Trans. Reliab. | 6 |
| 2022 | Bandwidth-Efficient Multi-video Prefetching for Short Video StreamingabstractApplications that allow sharing of user-created short videos exploded in popularity in recent years. A typical short video application allows a user to swipe away the current video being watched and start watching the next video in a video queue. Such user interface causes significant bandwidth waste if users frequently swipe a video away before finishing watching. Solutions to reduce bandwidth waste without impairing the Quality of Experience (QoE) are needed. Solving the problem requires adaptively prefetching of short video chunks, which is challenging as the download strategy needs to match unknown user viewing behavior and network conditions. In our work, we first formulate the problem of adaptive multi-video prefetching in short video streaming. Then, to facilitate the integration and comparison of researchers' algorithms towards solving the problem, we design and implement a discrete-event simulator, which we release as open source. Finally, based on the organization of the Short Video Streaming Grand Challenge at ACM Multimedia 2022, we analyze and summarize the algorithms of the contestants, with the hope of promoting the research community towards addressing this problem. Xutong Zuo, Yishu Li, Mohan Xu, Wei Tsang Ooi, Jiangchuan Liu, Junchen Jiang, Xinggong Zhang, Kai Zheng 0003, Yong Cui 0001 |
ACM Multimedia | 2 |
| 2022 | CASMS: Combining clustering with attention semantic model for identifying security bug reports
Jacky W. Keung, Zhen Yang 0022, Xiao Yu 0008, Yishu Li, Hao Zhang 0085 |
Inf. Softw. Technol. | 5 |