EDBT 2026 Demo / reviewers in the wild / expert
Qing Mi
dblp:180/3638
· DBLP profile ↗
34ranked-venue papers
19as first author
24since 2021 · last 2025
0000-0001-5063-3189ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 19 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Task Gradual Inference with a Single Encoder-Decoder Network for Automatic Portrait MattingabstractThis paper presents a multi-task gradual inference model, MTGINet, for automatic portrait matting. It handles the subtasks of automatic portrait matting, namely portrait-transition-background trimap segmentation and transition region matting, with a single encoder-decoder structure. First, we enrich the highest stage of features from the encoder with portrait shape context via a shape context aggregation (SCA) module for trimap segmentation. Then, we fuse the SCA-enhanced features with detailed clues from the encoder for transition-region-aware alpha matting. The gradual inference model naturally allows sufficient interaction between the subtasks via forward computation and backwards propagation during training, and therefore achieves high accuracy while maintaining low complexity. In addition, considering the discrepancies in feature requirements across subtasks, we adapt the features from the encoders before reusing them via a feature rectification module. In addition to the MTGINet model, we have constructed a new large-scale dataset, HPM-17K, for half-body portrait matting. It consists of 16,967 images with diverse backgrounds. Comparative experiments with existing deep models on the public P3M-10K dataset and our HPM-17K dataset demonstrate that the proposed model exhibits state-of-the-art performance. Wenbing Yang, Wei Ma 0008, Qing Mi, Hongbin Zha |
Comput. Vis. Media | 3 |
| 2025 | Towards Explainable Code Readability Classification With Graph Neural NetworksabstractABSTRACT Code readability is of central concern for developers, as a more readable code indicates higher maintainability, reusability, and portability. In recent years, many deep learning–based code readability classification methods have been proposed. Among them, a graph neural network (GNN)–based model has achieved the best performance in the field of code readability classification. However, it is still unclear what aspects of the model's input lead to its decisions, which hinders its practical use in the software industry. To improve the interpretability of existing code readability classification models and identify key code characteristics that drive their readability predictions, we propose an explanation framework with GNN explainers towards transparent and trustworthy code readability classification. First, we propose a simplified Abstract Syntax Tree (AST)–based code representation method, which transforms Java code snippets into ASTs and discards lower‐level nodes with limited information. Then, we retrain the state‐of‐the‐art GNN‐based model together with our simplified program graphs. Finally, we employ SubgraphX to explain the model's code readability predictions at the subgraph level and visualize the explanation results to further analyze what causes such predictions. The experimental results show that sequential logic, code comments, selection logic, and nested structure are the most influential code characteristics when classifying code snippets as readable or unreadable. Further investigations indicate the model's proficiency in capturing features related to complex logic structures and extensive data flows but point to its limitations in identifying readability issues associated with naming conventions and code formatting. The explainability analysis conducted in this research is the first step towards more transparent and reliable code readability classification. We believe that our findings are useful in providing constructive suggestions for developers to write more readable code and delimitating directions for future model improvement. Qing Mi, Zhiyou Xiao, Liyan Tao |
J. Softw. Evol. Process. | 1 |
| 2024 | WCL-CRC: A Weakly-Supervised Contrastive Learning Framework for Few-Shot Code Readability ClassificationabstractCode readability is a critical aspect of program comprehension and an important metric for overall software quality.Recently, many deep learning-based code readability classification models have been proposed.Although they reached state-ofthe-art classification results, we consider that their performance is limited due to the shortage of labeled data.To address this problem, we propose WCL-CRC, a two-stage weakly-supervised contrastive learning framework for few-shot code readability classification, which pre-trains existing models using contrastive learning techniques with a large amount of weakly-labeled data obtained from open-source repositories and then fine-tunes them with a few labeled data.We conduct experiments on Java and Python datasets.The results show that applying WCL-CRC to existing code readability classification models can improve their Accuracy by 0.21% to 2.43%, F-Measure by 0.52% to 2.66%, AUC by 0.14% to 1.58%, and MCC by 1.06% to 8.73%, indicating that our approach effectively learns readability-related information from a large amount of weakly-labeled data and significantly improves code readability classification performance in situations where labeled data is limited. Qing Mi, Yueyue Xi, Junyang Li 0002, Shijia Tang |
SEKE | 1 |
| 2024 | WCL-CRC: A Weakly-Supervised Contrastive Learning Framework for Few-Shot Code Readability Classification
Qing Mi, Yueyue Xi, Junyang Li 0002, Shijia Tang |
SEKE | 1 |
| 2024 | Towards Understanding the Distribution and Evolution of Code Smells: An Empirical Study on Java and Python Projects (S)abstractCode smells are usually associated with poor coding practices, design issues, and potential bugs.Studies on code smells are important.However, existing studies lack research on the distribution and evolution of code smells in different programming languages.To fill this gap, we collected 9 Java projects and 9 Python projects from GitHub and analyzed them with existing tools.We found that Java's code smell rate (the percentage of the number of code smells in total lines of code) of 11.11% is higher than Python's code smell rate of 1.31%."Magic Number", "Long Statement", and "Unutilized Abstraction" are the top three code smells in Java, while "Design Smells", "Common Weakness Enumeration", and "Unused Code" are the top three code smells in Python.The code smell aggregation level (the percentage of code smells in the top five modules in the total number of code smells) in Java is 44.07%while in Python is 59.18%, indicating that both languages are likely to accumulate code smells in certain modules.During the project's progress, the code smell rate remained steady in both languages, while the aggregation level showed a downward trend.Our findings can help developers understand common code smells and their trends in projects, thus guiding them to improve software quality. Qing Mi, Tengyuan Zhang, Keru Cai |
SEKE | 1 |
| 2024 | Dual-modal non-local context guided multi-stage fusion for indoor RGB-D semantic segmentation
Wei Ma 0008, Fangfang Liang, Qing Mi |
Expert Syst. Appl. | 4 |
| 2023 | Identifying Topics and Trends in DevOps: A Study of Stack Overflow PostsabstractDevOps (i.e., Development and Operations) is a growing concept, which aims to make building, testing, and releasing software faster, more frequent, and more reliable by automating the software delivery and architectural change process. Since DevOps is a relatively new concept, we seek to understand the hot topics in DevOps and the challenges encountered so far. We use data from Stack Overflow (SO), the largest developer question-and-answer platform, to understand the interests and difficulties of DevOps. First, we collected all DevOps-related posts from 2009 to 2022 based on SO tags. Then we used the Latent Dirichlet Allocation model to identify topics, and manually analyzed time trends and hot topics. The results indicate that: (1) DevOps-related issues can be divided into four categories: Container, Pipeline, Configuration, and Deployment; (2) The issues that get more attention are DevOps foundation issues, cloud deployment failure issues, Kubernetes technology usage issues, and DevOps complex environment issues, while the most difficult issues to solve are CI/CD-related issues; (3) DevOps has been introduced since 2009 and the number of questions on the SO platform is growing, especially between 2014 and 2020; (4) The percentage of DevOps-related questions with no accepted answers on the SO platform reached 56.3%. Our results show that problems are evident in terms of complex environments and the lack of experts. Also, developers need to pay more attention to basic theoretical knowledge when they are new to DevOps. Qing Mi, Qinghang Bao, Longjie Cui |
SEAA | 1 |
| 2023 | An Empirical Study of Coding Style Compliance on Stack OverflowabstractStack Overflow (SO) is one of the world's largest technical Q&A websites, in which many posts contain code snippets.However, these code snippets may not comply with coding style guidelines and result in the problem of low readability and maintainability.To provide a better understanding of this coding style compliance issue for SO users, we plan and conduct an empirical study on SO.Specifically, we collected over 400,000 code snippets from SO in three languages, namely Python, C/C++, and JavaScript.The posts are divided into two types (i.e., question and answer) and analyzed separately.We found that for the question-and answer-type posts, more than 90% and 60% of code snippets contain style violations.The most frequently found violation is syntax errors for Python and indentations for C/C++ and JavaScript.In addition, the results show that with more violations in code snippets, the "Score" of Python and C/C++ posts, the "FavoriteCount" of C/C++ questions, and the "CommentCount" of JavaScript questions tend to be lower.The findings of our research indicate that code snippets on SO do not have good coding style compliance.Users, especially programming beginners are supposed to be wary of the potential problems of reusing code snippets on SO. Qing Mi, Haotian Bai, Xiaozhou Wang, Xingyue Song |
SEKE | 1 |
| 2023 | View-relation constrained global representation learning for multi-view-based 3D object recognition
Ruchang Xu, Qing Mi, Wei Ma 0008, Hongbin Zha |
Appl. Intell. | 2 |
| 2023 | An improved unified domain adversarial category-wise alignment network for unsupervised cross-domain sentiment classification
Xibin Jia, Meng Zeng, Luo Wang, Qing Mi |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | A graph-based code representation method to improve code readability classification
Qing Mi, Han Weng, Qinghang Bao, Longjie Cui, Wei Ma 0008 |
Empir. Softw. Eng. | 1 |
| 2023 | DSC-MDE: Dual structural contexts for monocular depth estimation
Wubin Yan, Lijun Dong, Wei Ma 0008, Qing Mi, Hongbin Zha |
Knowl. Based Syst. | 4 |
| 2023 | What makes a readable code? A causal analysis methodabstractAbstract Context Code readability is one of the most important quality attributes for software source code. To investigate which features affect code readability, most existing studies rely on correlation‐based methods. However, spurious correlations (a mathematical relationship wherein two variables appear to be causal but are not) involved in correlation‐based methods may affect research conclusions. Objective In order to remove spurious correlations and obtain conclusions from the perspective of causation as to what makes a readable code, we propose a causal theory‐based approach to analyze the relationship between code features and code readability scores. Method First, we adopt the PC algorithm and additive noise models to construct the causal graph on the basis of the selected code features. Then, we use the linear regression algorithm based on the back‐door criterion to obtain the causal effect of different features on code readability. Result We conduct a set of experiments on readability data labeled by human annotators. The experimental results show that the average number of comments positively impacts code readability, with each additional unit increasing the code readability score by 0.799 points. Whereas the average number of assignments, identifiers, and periods have a negative impact, with each additional unit decreasing the code readability score by 0.528, 0.281, and 0.170 points respectively. Conclusion We believe that our findings will provide developers with a better understanding of the patterns behind code readability, and guide developers to optimize their code as the ultimate goal. Qing Mi, Zhi Cai, Xibin Jia |
Softw. Pract. Exp. | 1 |
| 2022 | ReINView: Re-interpreting Views for Multi-view 3D Object RecognitionabstractMulti-view-based 3D object recognition is important in robot-environment interaction. However, recent methods simply extract features from each view via convolutional neural networks (CNNs) and then fuse these features together to make predictions. These methods ignore the inherent ambiguities of each view caused due to 3D-2D projection. To address this problem, we propose a novel deep framework for multi-view-based 3D object recognition. Instead of fusing the multi-view features directly, we design a re-interpretation module (ReINView) to eliminate the ambiguities at each view. To achieve this, ReINView re-interprets view features patch by patch by using their context from nearby views, considering that local patches are generally co-visible at nearby viewpoints. Since contour shapes are essential for 3D object recognition as well, ReINView further performs view-level re-interpretation, in which we use all the views as context sources since the target contours to be re-interpreted are globally observable. The re-interpreted multi-view features can better reflect the 3D global and local structures of the object. Experiments on both ModelNet40 and ModelNet10 show that the proposed model outperforms state-of-the-art methods in 3D object recognition. Ruchang Xu, Wei Ma 0008, Qing Mi, Hongbin Zha |
IROS | 3 |
| 2022 | Rank Learning-Based Code Readability Assessment with Siamese Neural NetworksabstractAutomatically assessing code readability is a relatively new challenge that has attracted growing attention from the software engineering community. In this paper, we outline the idea to regard code readability assessment as a learning-to-rank task. Specifically, we design a pairwise ranking model with siamese neural networks, which takes as input a code pair and outputs their readability ranking order. We have evaluated our approach on three publicly available datasets. The result is promising, with an accuracy of 83.5%, a precision of 86.1%, a recall of 81.6%, an F-measure of 83.6% and an AUC of 83.4%. Qing Mi |
ASE | 1 |
| 2022 | How students choose names: A replication studyabstractNames of classes/methods/variables play an important role in code readability. To investigate how developers choose names, Feitelson et al. conducted an empirical survey and suggested a method to improve naming quality. We replicated their study, but limited the survey subjects to university students. Specifically, we conducted two experiments including 341 students from freshmen to seniors. The aim of the first experiment was to investigate the characteristics of the names given by students. The experimental results showed that the name length as well as the number of words contained in names increased with the grade and students have ambiguity in understanding variable names. The second experiment was to verify whether Feitelson et al.’s naming method can help improve the quality of the names given by students. The experimental results showed an improvement in naming quality for more than 67% of cases, which confirms the validity of the method for university students. Qing Mi, Bingnuo Chen |
ASE | 1 |
| 2022 | An Enhanced Data Augmentation Approach to Support Multi-Class Code Readability ClassificationabstractContext: Code readability plays a critical role in software maintenance and evolvement, where a metric for classifying code readability levels is both applicable and desired.However, most prior research has treated code readability classification as a binary classification task due to the lack of labeled data.Objective: To support the training of multi-class code readability classification models, we propose an enhanced data augmentation approach.Method: The approach includes the use of domainspecific data transformation and GAN-based data augmentation.By virtue of this augmentation approach, we could generate sufficient readability data and well train a multi-class code readability model.Result: A series of experiments are conducted to evaluate our augmentation approach.The experimental results show that a state-of-the-art multi-class code readability classification accuracy of 68.0% is reached with a significant improvement of 6.3% compared to only using the original data.Conclusion: As an innovative work of proposing multi-class code readability classification and an enhanced code readability data augmentation approach, our method is proved to be effective. Qing Mi, Yiqun Hao, Maran Wu, Liwei Ou |
SEKE | 1 |
| 2022 | Improving Multi-Class Code Readability Classification with An Enhanced Data Augmentation Approach (130)abstractBeing a critical factor affecting the maintainability and reusability of the software, code readability is growing crucial in modern software development, where a metric for classifying code readability levels is both applicable and desired. However, most prior research has treated code readability classification as a binary classification task due to the lack of labeled data. To support the training of multi-class code readability classification models, we propose an enhanced data augmentation approach that could be used to generate sufficient readability data and well train a multi-class code readability model. The approach includes the use of domain-specific data transformation and GAN-based data augmentation. We conduct a series of experiments to verify our augmentation approach and gain a state-of-the-art multi-class code readability classification performance with 69.5% Micro-F1, 54.0% Macro-F1 and 67.7% Macro-AUC. Compared to the results where no augmented data is used, the improvements on Micro-F1, Macro-F1 and Macro-AUC are significant with 6.9%, 11.3% and 11.2%, respectively. As an innovative work of proposing multi-class code readability classification and an enhanced code readability data augmentation approach, our method is proved to be effective. Qing Mi, Luo Wang, Lisha Hu, Liwei Ou |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2022 | Towards using visual, semantic and structural features to improve code readability classification
Qing Mi, Yiqun Hao, Liwei Ou, Wei Ma 0008 |
J. Syst. Softw. | 1 |
| 2022 | A Multimodality-Contribution-Aware TripNet for Histologic Grading of Hepatocellular CarcinomaabstractHepatocellular carcinoma (HCC) is a type of primary liver malignant tumor with a high recurrence rate and poor prognosis even undergoing resection or transplantation. Accurate discrimination of the histologic grades of HCC plays a critical role in the management and therapy of HCC patients. In this paper, we discuss a deep learning-based diagnostic model for HCC histologic grading with multimodal Magnetic Resonance Imaging (MRI) images to overcome the problem of limited well-annotated data and extract the discriminated fusion feature referring to the clinical diagnosis experience of radiologists. Accordingly, we propose a novel Multimodality-Contribution-Aware TripNet (MCAT) based on the metric learning and the attention-aware weighted multimodal fusion. The novelty of the method lies in the multimodality small-shot learning architecture designation and the multimodality adaptive weighted computing scheme. The comprehensive experiments are done on the clinic dataset with the well-annotation of lesion location by the professional radiologist. The experimental results show that our proposed MCAT is not only able to achieve acceptable quantitative measuring of HCC histologic grading based on the MRI sequences with small cases but also outperforms previous models in HCC histologic grading, reaching an accuracy of 84 percent, a sensitivity of 87 percent and precision of 89 percent. Xibin Jia, Qing Mi, Zhenghan Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | MAGAN: Multi-attention Generative Adversarial Networks for Text-to-Image Generation
Xibin Jia, Qing Mi |
PRCV (4) | 2 |
| 2021 | An unsupervised person re-identification approach based on cross-view distribution alignmentabstractAbstract Unsupervised clustering is a kind of popular solution for unsupervised person re‐identification (re‐ID). However, due to the influence of cross‐view differences, the results of clustering labels are not accurate. To solve this problem, an unsupervised re ID method based on cross‐view distributed alignment (CV‐DA) to reduce the influence of unsupervised cross‐view is proposed. Specifically, based on a popular unsupervised clustering method, density clustering DBSCAN is used to obtain pseudo labels. By calculating the similarity scores of images in the target domain and the source domain, the similarity distribution of different camera views is obtained and is aligned with the distribution with the consistency constraint of pseudo labels. The cross‐view distribution alignment constraint is used to guide the clustering process to obtain a more reliable pseudo label. The comprehensive comparative experiments are done in two public datasets, i.e. Market‐1501 and DukeMTMC‐reID. The comparative results show that the proposed method outperforms several state‐of‐the‐art approaches with mAP reaching 52.6% and rank1 71.1%. In order to prove the effectiveness of the proposed CV‐DA, the proposed constraint is added into two advanced re‐ID methods. The experimental results demonstrate that the mAP and rank increase by 0.5–2% after using the cross‐view distribution alignment constraint comparing with that of the associated original methods without using CV‐DA. Xibin Jia, Qing Mi |
IET Image Process. | 3 |
| 2021 | The effectiveness of data augmentation in code readability classification
Qing Mi, Yan Xiao 0002, Zhi Cai, Xibin Jia |
Inf. Softw. Technol. | 1 |
| 2021 | Siamese CNN-based rank learning for quality assessment of inpainted images
Xiangdong Meng, Wei Ma 0008, Chunhu Li, Qing Mi |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Improving bug localization with word embedding and enhanced convolutional neural networks
Yan Xiao 0002, Jacky W. Keung, Kwabena Ebo Bennin, Qing Mi |
Inf. Softw. Technol. | 4 |
| 2018 | An Inception Architecture-Based Model for Improving Code Readability ClassificationabstractThe process of classifying a piece of source code into a Readable or Unreadable class is referred to as Code Readability Classification. To build accurate classification models, existing studies focus on handcrafting features from different aspects that intuitively seem to correlate with code readability, and then exploring various machine learning algorithms based on the newly proposed features. On the contrary, our work opens up a new way to tackle the problem by using the technique of deep learning. Specifically, we propose IncepCRM, a novel model based on the Inception architecture that can learn multi-scale features automatically from source code with little manual intervention. We apply the information of human annotators as the auxiliary input for training IncepCRM and empirically verify the performance of IncepCRM on three publicly available datasets. The results show that: 1) Annotator information is beneficial for model performance as confirmed by robust statistical tests (i.e., the Brunner-Munzel test and Cliff's delta); 2) IncepCRM can achieve an improved accuracy against previously reported models across all datasets. The findings of our study confirm the feasibility and effectiveness of deep learning for code readability classification. Qing Mi, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah, Xiupei Mei |
EASE | 1 |
| 2018 | Bug Localization with Semantic and Structural Features using Convolutional Neural Network and Cascade ForestabstractBackground: Correctly localizing buggy files for bug reports together with their semantic and structural information is a crucial task, which would essentially improve the accuracy of bug localization techniques. Aims: To empirically evaluate and demonstrate the effects of both semantic and structural information in bug reports and source files on improving the performance of bug localization, we propose CNN_Forest involving convolutional neural network and ensemble of random forests that have excellent performance in the tasks of semantic parsing and structural information extraction. Method: We first employ convolutional neural network with multiple filters and an ensemble of random forests with multi-grained scanning to extract semantic and structural features from the word vectors derived from bug reports and source files. And a subsequent cascade forest (a cascade of ensembles of random forests) is used to further extract deeper features and observe the correlated relationships between bug reports and source files. CNNLForest is then empirically evaluated over 10,754 bug reports extracted from AspectJ, Eclipse UI, JDT, SWT, and Tomcat projects. Results: The experiments empirically demonstrate the significance of including semantic and structural information in bug localization, and further show that the proposed CNN_Forest achieves higher Mean Average Precision and Mean Reciprocal Rank measures than the best results of the four current state-of-the-art approaches (NPCNN, LR+WE, DNNLOC, and BugLocator). Conclusion: CNNLForest is capable of defining the correlated relationships between bug reports and source files, and we empirically show that semantic and structural information in bug reports and source files are crucial in improving bug localization. Yan Xiao 0002, Jacky W. Keung, Qing Mi, Kwabena Ebo Bennin |
EASE | 3 |
| 2018 | Not all bug reopens are negative: A case study on eclipse bug reports
Qing Mi, Jacky W. Keung, Yuqi Huo, Solomon Mensah |
Inf. Softw. Technol. | 1 |
| 2018 | Improving code readability classification using convolutional neural networks
Qing Mi, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah, Yujin Gao |
Inf. Softw. Technol. | 1 |
| 2018 | Machine translation-based bug localization technique for bridging lexical gap
Yan Xiao 0002, Jacky W. Keung, Kwabena Ebo Bennin, Qing Mi |
Inf. Softw. Technol. | 4 |
| 2018 | On the value of a prioritization scheme for resolving Self-admitted technical debt
Solomon Mensah, Jacky W. Keung, Jeffrey Svajlenko, Kwabena Ebo Bennin, Qing Mi |
J. Syst. Softw. | 5 |
| 2017 | Identifying Textual Features of High-Quality Questions: An Empirical Study on Stack OverflowabstractBackground: Stack Overflow (SO) is a programming-specific Q&A website that serves as a valuable repository of software engineering knowledge. For SO members, formulating a good question is the first step towards eliciting satisfactory responses. Aims: To guide SO members on how to make a good question, we conduct an empirical study using the publicly available Stack Overflow Data Dump for the period of 2008-2016. Method: We first choose 25 features along 5 dimensions to represent the textual characteristics that we are interested in. Making use of the Boruta algorithm, we then capture all features that are either strongly or weakly relevant to the question quality. Results: The results show that the number of tags and code snippets are the most discriminative features, whereas there is only a weak correlation between the question quality and the sentiment-related factors. Based on the empirical evidence, we provide useful and usable suggestions to SO members on how to optimize their questions. Conclusions: We consider that our findings will provide SO members with a better understanding of the patterns behind high-quality questions, this is to support effective and efficient utilization of Q&A websites as the ultimate goal. Qing Mi, Yujin Gao, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah |
APSEC | 1 |
| 2017 | Improving Bug Localization with an Enhanced Convolutional Neural NetworkabstractBackground: Localizing buggy files automatically speeds up the process of bug fixing so as to improve the efficiency and productivity of software quality teams. There are other useful semantic information available in bug reports and source code, but are mostly underutilized by existing bug localization approaches. Aims: We propose DeepLocator, a novel deep learning based model to improve the performance of bug localization by making full use of semantic information. Method: DeepLocator is composed of an enhanced CNN (Convolutional Neural Network) proposed in this study considering bug-fixing experience, together with a new rTF-IDuF method and pretrained word2vec technique. DeepLocator is then evaluated on over 18,500 bug reports extracted from AspectJ, Eclipse, JDT, SWT and Tomcat projects. Results: The experimental results show that DeepLocator achieves 9.77% to 26.65% higher Fmeasure than the conventional CNN and 3.8% higher MAP than a state-of-the-art method HyLoc using less computation time. Conclusion: DeepLocator is capable of automatically connecting bug reports to the corresponding buggy files and successfully achieves better performance based on a deep understanding of semantics in bug reports and source code. Yan Xiao 0002, Jacky W. Keung, Qing Mi, Kwabena Ebo Bennin |
APSEC | 3 |
| 2016 | An empirical analysis of reopened bugs based on open source projectsabstractBackground: Bug fixing is a long-term and time-consuming activity. A software bug experiences a typical life cycle from newly reported to finally closed by developers, but it could be reopened afterwards for further actions due to reasons such as unclear description given by the bug reporter and developer negligence. Bug reopening is neither desirable nor could be completely avoided in practice, and it is more likely to bring unnecessary workloads to already-busy developers. Aims: To the best of our knowledge, there has been a little previous work on software bug reopening. In order to further study in this area, we perform an empirical analysis to provide a comprehensive understanding of this special area. Method: Based on four open source projects from Eclipse product family, they are CDT, JDT, PDE and Platform, we first quantitatively analyze reopened bugs from perspectives of proportion, impacts and time distribution. After initial exploration on their characteristics, we then qualitatively summarize root causes for bug reopening, this is carried out by investigating developer discussions recorded in Eclipse Bugzilla. Results: Results show that 6%--10% of total bugs will lead to reopening eventually. Over 93% of reopened bugs place serious influence on the normal operation of the system being developed. Several key reasons for bug reopening have been identified in our empirical study. Conclusions: Although reopened bugs have significant impacts on both end users and developers, it is quite possible to reduce bug reopening rate through the adoption of appropriate methods, such as promoting effective and efficient communication among bug reporters and developers, which is supported by empirical evidence in this study. Qing Mi, Jacky W. Keung |
EASE | 1 |