Zexian Zhang

dblp:293/1125 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 The impact of feature selection and feature reduction techniques for code smell detection: A comprehensive empirical study
Zexian Zhang, Shuang Yin, Haoxuan Chen
Autom. Softw. Eng.1
2024 On the Relative Value of Feature Selection Techniques for Code Smell Detection
abstract
Machine/deep learning-based code smell detection aims to develop a classification model based on code smell features to predict the presence of code smell in new code instances. To ensure accurate detection, it is crucial to eliminate irrelevant or redundant features that may negatively impact performance. Previous studies have produced inconsistent findings about the impact of feature selection techniques for code smell detection, possibly because they examined only a limited number of different techniques. To address this gap, our study aims to provide a comprehensive analysis of feature selection techniques in code smell detection. We investigate 34 feature selection techniques with 7 classification models to build the code smell detection models on 6 code smell datasets. To assess these effects, we use 3 evaluation metrics, i.e., Precision, Recall, and F-measure, and compare the performance differences using the Scott-Knott effect size difference test and the McNemar's test. The results show that (1) Not all feature selection techniques significantly improve detection performance. The techniques with better performance are chi-square, probabilistic significance, information gain, and symmetrical uncertainty. (2) In general, probabilistic significance should be used as the “generic” feature selection technique because detection models using probabilistic significance can identify more of the same smelly instances compared to models using other methods. (3) The high-frequency features selected by the four highest-performing techniques, which are important for identifying the corresponding code smells, are different for each dataset.
Zexian Zhang, Shuang Yin, Haoxuan Chen
APSEC1
2024 Practitioners' Expectations on Code Smell Detection
abstract
Code smell detection can automatically identify code smells in software source code to help developers to improve code maintainability, readability, and overall code quality. Currently, a wide variety of code smell detection techniques/tools are proposed for practical use. However, it is unclear what practitioners expect for code smell detection tools and whether the existing research meets their needs. To fill the gap, we conduct an empirical study. We first interview 10 software development professionals and subsequently survey 310 software practitioners about their practices and expectations of code smell detection tools. In addition, we conduct an extensive literature review of code smell detection papers published in major publications from 2014 to 2024, and compare current research findings with practitioners' expectations. From this comparison, we highlight the direction in which researchers need to work to develop code smell detection techniques that are important to practitioners.
Zexian Zhang, Shuang Yin, Wenliu Wei, Jacky W. Keung
COMPSAC1
2024 What Makes a High-Quality Training Dataset for Large Language Models: A Practitioners' Perspective
abstract
Large Language Models (LLMs) have demonstrated remarkable performance in various application domains, largely due to their self-supervised pre-training on extensive high-quality text datasets. However, despite the importance of constructing such datasets, many leading LLMs lack documentation of their dataset construction and training procedures, leaving LLM practitioners with a limited understanding of what makes a high-quality training dataset for LLMs. To fill this gap, we initially identified 18 characteristics of high-quality LLM training datasets, as well as 10 potential data pre-processing methods and 6 data quality assessment methods, through detailed interviews with 13 experienced LLM professionals. We then surveyed 219 LLM practitioners from 23 countries across 5 continents. We asked our survey respondents to rate the importance of these characteristics, provide a rationale for their ratings, specify the key data pre-processing and data quality assessment methods they used, and highlight the challenges encountered during these processes. From our analysis, we identified 13 crucial characteristics of high-quality LLM datasets that receive a high rating, accompanied by key rationale provided by respondents. We also identified some widely-used data pre-processing and data quality assessment methods, along with 7 challenges encountered during these processes. Based on our findings, we discuss the implications for researchers and practitioners aiming to construct high-quality training datasets for optimizing LLMs.
Xiao Yu 0008, Zexian Zhang, Feifei Niu, Xing Hu 0008, Xin Xia 0001, John C. Grundy
ASE2
2024 Data preparation for Deep Learning based Code Smell Detection: A systematic literature review
Fengji Zhang, Zexian Zhang, Jacky W. Keung, Xiangru Tang, Zhen Yang 0022, Xiao Yu 0008
J. Syst. Softw.2