Qiaolin Qin

dblp:335/7896 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2026
0009-0004-9735-2965ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 7 first-author · 9 since 2021
YearPublicationVenuePosition
2026 A Story About Cohesion and Separation: Label-Free Metric for Log Parser Evaluation
Qiaolin Qin, Jianchen Zhao, Heng Li 0007, Weiyi Shang, Ettore Merlo
SANER1
2026 Plug it and Play on Logs: A configuration-free statistic-based log parser
Qiaolin Qin, Xingfang Wu, Heng Li 0007, Ettore Merlo
Empir. Softw. Eng.1
2026 Unsupervised, robust, and lightweight detection of data pattern anomalies and outliers
abstract
Context: As a current consensus, data quality strongly impacts the process of building software and AI systems. Hence, practitioners must detect the anomalies in data and repair these underlying problems. When dealing with big data in the industry, statistic-based unsupervised anomaly detectors come in handy since they do not require labels and are highly scalable. However, we noticed that these tools unsupervised, always require data-dependent parameters, which can largely affect the detection performance and are effort-consuming to configure. Objectives: In this work, we propose a fully unsupervised, statistic-based cell-level data anomaly detector, LUCARIO (Learning Unsupervised, Cell-level Anomaly-detector for Regex Incompatibilities and Outliers). Our approach aims to detect common cell-level data anomalies (pattern violations and outliers) without manual efforts in data annotations or parameter configurations, yet providing a robust performance for different data across diverse domains. Methods: According to previous studies, we categorized cell anomalies into two categories: pattern violations and outliers (categorical and numerical). We proposed three detection approaches based on heuristics and statistical theories to identify these anomalies. To evaluate LUCARIO’s effectiveness and usability, we conducted experiments on six open-source datasets and a real-life industrial dataset from our industrial partner CompanyX . Results: According to our experiment on six open-source datasets in various domains, LUCARIO can stably detect cell-level data issues (pattern violations and outliers) regardless of the dataset’s size and anomaly rate. LUCARIO reached an average F1 score of 0.54, higher than all baseline unsupervised anomaly detectors, including GPT-5 with few-shot prompting. Practitioners from CompanyX generally agree that LUCARIO can benefit their data quality by detecting critical data issues and providing reliable suggestions. Conclusion: The experimental results show that LUCARIO has the potential to improve the data used for both software and AI system construction in real-life applications, suggesting its practicality in data management.
Qiaolin Qin, Heng Li 0007, Ettore Merlo
Inf. Softw. Technol.1
2026 Prune bias from the root: Bias removal and fairness estimation by pruning sensitive attributes in pre-trained DNN models
abstract
Deep learning models (DNNs) are widely used in high-stakes decision-making domains, but they often inherit and amplify biases present in training data, leading to unfair predictions. Given this context, fairness estimation metrics and bias removal methods are required to select and enhance fair models. However, we found that existing metrics lack robustness in estimating multi-attribute group fairness. Further, existing post-processing bias removal methods often focus on group fairness and fail to address individual fairness or optimize along multiple sensitive attributes. In this study, we explore the effectiveness of attribute pruning (i.e., zeroing out sensitive attribute weights in a pre-trained DNN’s input layer) in both bias removal and multi-attribute group fairness estimation. To study attribute pruning’s impact on bias removal, we conducted experiments on 32 models and 4 widely used datasets, and compared its effect in single-attribute group bias removal and accuracy preservation with 3 baseline post-processing methods. We then leveraged 3 datasets with multiple sensitive attributes to demonstrate how to use attribute pruning for multi-attribute group fairness estimation. Single-attribute pruning can better preserve model accuracy than conventional post-processing methods in 23 out of 32 cases, and enforces individual fairness by design. However, since individual fairness and group fairness are fundamentally different objectives, attribute pruning’s effect on group fairness metrics is often inconsistent. We also extend our approach to a multi-attribute setting, demonstrating its potential for improving individual fairness jointly across sensitive attributes and for enabling multi-attribute fairness-aware model selection. Attribute pruning is a practical post-processing approach for enforcing individual fairness, with limited and data-dependent impact on group fairness. These limitations reflect the inherent trade-off between individual and group fairness objectives. In addition, attribute pruning provides a useful mechanism for bias estimation, particularly in multi-attribute contexts. We advocate for its adoption as a comparison baseline in fairness-aware AI development and encourage further exploration. • Prove the effectiveness of single-attribute pruning for bias removal on DNNs. • Propose the use of attribute pruning on multiple features. • Introduce a robust multi-attribute fairness estimation method. • Explicitly discuss the limits of algorithmic bias removal and provide suggestions.
Qiaolin Qin, Ettore Merlo
Inf. Softw. Technol.1
2025 Automated, Unsupervised, and Auto-Parameterized Inference of Data Patterns and Anomaly Detection
abstract
With the advent of data-centric and machine learning (ML) systems, data quality is playing an increasingly critical role for ensuring the overall quality of software systems. Data preparation, an essential step towards high data quality, is known to be a highly effort-intensive process. Although prior studies have dealt with one of the most impacting issues, data pattern violations, these studies usually require data-specific configurations (i.e., parameterized) or use carefully curated data as learning examples (i.e., supervised), relying on domain knowledge and deep understanding of the data, or demanding significant manual effort. In this paper, we introduce RIOLU: Regex Inferencer autO-parameterized Learning with Uncleaned data. RIOLU is fully automated, automatically parameterized, and does not need labeled samples. RIOLU can generate precise patterns from datasets in various domains, with a high F1 score of 97.2 %, exceeding the state-of-the-art baseline. In addition, according to our experiment on five datasets with anomalies, RIOLU can automatically estimate a data column's error rate, draw normal patterns, and predict anomalies from unlabeled data with higher performance (up to$\mathbf{8 0 0. 4 \%}$improvement in terms of F1) than the state-of-the-art baseline, even outperforming ChatGPT in terms of both accuracy (12.3 % higher F1) and efficiency (10 % less inference time). A variant of RIOLU, with user guidance, can further boost its precision, with up to$\mathbf{3 7. 4 \%}$improvement in terms of F1. Our evaluation in an industrial setting further demonstrates the practical benefits of RIOLU.
Qiaolin Qin, Heng Li 0007, Ettore Merlo, Maxime Lamothe
ICSE1
2025 Exploring the Potential of Llama Models in Automated Code Refinement: A Replication Study
abstract
Code reviews are an integral part of software development and have been recognized as a crucial practice for minimizing bugs and favouring higher code quality. They serve as an important checkpoint before committing code and play an essential role in knowledge transfer between developers. However, code reviews can be time-consuming and can stale the development of large software projects. In a recent study, Guo et al. assessed how ChatGPT3.5 can help the code review process. They evaluated the effectiveness of ChatGPT in automating the code refinement tasks, where developers recommend small changes in the submitted code. While Guo et al.'s study showed promising results, proprietary models like ChatGPT pose risks to data privacy and incur extra costs for software projects. In this study, we explore alternatives to ChatGPT in code refinement tasks by including two open-source, smaller-scale large language models: CodeLlama and Llama 2 (7B parameters). Our results show that, if properly tuned, the Llama models, particularly CodeLlama, can achieve reasonable performance, often comparable to ChatGPT in auto-mated code refinement. However, not all code refinement tasks are equally successful: tasks that require changing existing code (e.g., refactoring) are more manageable for models to automate than tasks that demand new code. Our study highlights the potential of open-source models for code refinement, offering cost-effective, privacy-conscious solutions for real-world software development.
Genevieve Caumartin, Qiaolin Qin, Sharon Chatragadda, Janmitsinh Panjrolia, Heng Li 0007, Diego Costa 0001
SANER2
2025 Preprocessing is All You Need: Boosting the Performance of Log Parsers with a General Preprocessing Framework
abstract
Log parsing has been a long-studied area in software engineering due to its importance in identifying dynamic vari-ables and constructing log templates. Prior work has proposed many statistic-based log parsers (e.g., Drain), which are highly efficient; they, unfortunately, met the bottleneck of parsing performance in comparison to semantic-based log parsers, which require labeling and more computational resources. Meanwhile, we noticed that previous studies mainly focused on parsing and often treated preprocessing as an ad hoc step (e.g., masking numbers). However, we argue that both preprocessing and parsing are essential for log parsers to identify dynamic variables: the lack of understanding of preprocessing may hinder the optimal use of parsers and future research. Therefore, our work studied existing log preprocessing approaches based on Loghub, a popular log parsing benchmark. We developed a general preprocessing framework with our findings and evaluated its impact on existing parsers. Our experiments show that the preprocessing framework significantly boosts the performance of four state-of-the-art statistic-based parsers. Drain, the best statistic-based parser, obtained improvements across all four parsing metrics (e.g., Fl score of template accuracy, FTA, increased by 108.9%). Compared to semantic-based parsers, it achieved a 28.3% improvement in grouping accuracy (GA), 38.1 % in FGA, and an 18.6% increase in FTA. Our work pioneers log preprocessing and provides a generalizable framework to enhance log parsing.
Qiaolin Qin, Roozbeh Aghili, Heng Li 0007, Ettore Merlo
SANER1
2025 Representation-based fairness evaluation and bias correction robustness assessment in neural networks
abstract
Context: While machine learning has achieved high predictive performance in many domains, decisions may still be biased and unfair regarding specific demographic groups characterized by sensitive attributes such as gender, age, or race. Objectives: In this paper, we introduce a novel approach to assess model fairness and bias correction robustness based on Computational Profile Distance (CPD) analysis with respect to sensitive attributes. Methods: To study model fairness, we quantify the model’s representation difference using the computational profile learned from different subgroups (e.g., male and female) on the individual and group level. To analyze the robustness of bias correction outcomes, we compare the correction suggestions provided based on confidence (i.e., softmax score) and likelihood (i.e., CPD). Results: To demonstrate the potential of the proposed approach, experiments have been performed using 24 models targeting 3 datasets used in previous fairness studies. Our experiments showed that computational profile distributions can effectively address model fairness from a representation perspective. Further, the experiments indicated that confidence-based bias correction decisions can vary largely from likelihood-based ones, and we should take both suggestions into account to obtain robust outcomes. Conclusion: Demonstrated with a set of experiments, our CPD-based approaches can help users build their trust in fairness assessment and bias mitigation of AI decisions, in ethically sensitive domains such as human resources, finance, health, and more.
Qiaolin Qin, Benjamin Djian, Ettore Merlo, Heng Li 0007, Sébastien Gambs
Inf. Softw. Technol.1
2024 Understanding Web Application Workloads and Their Applications: Systematic Literature Review and Characterization
abstract
Web applications, accessible via web browsers over the Internet, facilitate complex functionalities without local software installation. In the context of web applications, a workload refers to the number of user requests sent by users or applications to the underlying system. Existing studies have leveraged web application workloads to achieve various objectives, such as workload prediction and auto-scaling. However, these studies are conducted in an ad hoc manner, lacking a systematic understanding of the characteristics of web application workloads. In this study, we first conduct a systematic literature review to identify and analyze existing studies leveraging web application workloads. Our analysis sheds light on their workload utilization, analysis techniques, and high-level objectives. We further systematically analyze the characteristics of the web application workloads identified in the literature review. Our analysis centers on characterizing these workloads at two distinct temporal granularities: daily and weekly. We successfully identify and categorize three daily and three weekly patterns within the workloads. By providing a statistical characterization of these workload patterns, our study highlights the uniqueness of each pattern, paving the way for the development of realistic workload generation and resource provisioning techniques that can benefit a range of applications and research areas.
Roozbeh Aghili, Qiaolin Qin, Heng Li 0007, Foutse Khomh
ICSME2