Songtao Shang

dblp:165/9146 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0001-9437-2806ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2024 Software Defect Prediction Method Based on Clustering Ensemble Learning
abstract
The technique of software defect prediction aims to assess and predict potential defects in software projects and has made significant progress in recent years within software development. In previous studies, this technique largely relied on supervised learning methods, requiring a substantial amount of labeled historical defect data to train the models. However, obtaining these labeled data often demands significant time and resources. In contrast, software defect prediction based on unsupervised learning does not depend on known labeled data, eliminating the need for large‐scale data labeling, thereby saving considerable time and resources while providing a more flexible solution for ensuring software quality. This paper conducts software defect prediction using unsupervised learning methods on data from 16 projects across two public datasets (PROMISE and NASA). During the feature selection step, a chi‐squared sparse feature selection method is proposed. This feature selection strategy combines chi‐squared tests with sparse principal component analysis (SPCA). Specifically, the chi‐squared test is first used to filter out the most statistically significant features, and then the SPCA is applied to reduce the dimensionality of these significant features. In the clustering step, the dot product matrix and Pearson correlation coefficient (PCC) matrix are used to construct weighted adjacency matrices, and a clustering overlap method is proposed. This method integrates spectral clustering, Newman clustering, fluid clustering, and Clauset–Newman–Moore (CNM) clustering through ensemble learning. Experimental results indicate that, in the absence of labeled data, using the chi‐squared sparse method for feature selection demonstrates superior performance, and the proposed clustering overlap method outperforms or is comparable to the effectiveness of the four baseline clustering methods.
Hongwei Tao, Qiaoling Cao, Xiaoxu Niu, Zhenhao Geng, Songtao Shang
IET Softw.8
2024 Cross-Project Defect Prediction Using Transfer Learning with Long Short-Term Memory Networks
abstract
With the increasing number of software projects, within‐project defect prediction (WPDP) has already been unable to meet the demand, and cross‐project defect prediction (CPDP) is playing an increasingly significant role in the area of software engineering. The classic CPDP methods mainly concentrated on applying metric features to predict defects. However, these approaches failed to consider the rich semantic information, which usually contains the relationship between software defects and context. Since traditional methods are unable to exploit this characteristic, their performance is often unsatisfactory. In this paper, a transfer long short‐term memory (TLSTM) network model is first proposed. Transfer semantic features are extracted by adding a transfer learning algorithm to the long short‐term memory (LSTM) network. Then, the traditional metric features and semantic features are combined for CPDP. First, the abstract syntax trees (AST) are generated based on the source codes. Second, the AST node contents are converted into integer vectors as inputs to the TLSTM model. Then, the semantic features of the program can be extracted by TLSTM. On the other hand, transferable metric features are extracted by transfer component analysis (TCA). Finally, the semantic features and metric features are combined and input into the logical regression (LR) classifier for training. The presented TLSTM model performs better on the f ‐measure indicator than other machine and deep learning models, according to the outcomes of several open‐source projects of the PROMISE repository. The TLSTM model built with a single feature achieves 0.7% and 2.1% improvement on Log4j‐1.2 and Xalan‐2.7, respectively. When using combined features to train the prediction model, we call this model a transfer long short‐term memory for defect prediction (DPTLSTM). DPTLSTM achieves a 2.9% and 5% improvement on Synapse‐1.2 and Xerces‐1.4.4, respectively. Both prove the superiority of the proposed model on the CPDP task. This is because LSTM capture long‐term dependencies in sequence data and extract features that contain source code structure and context information. It can be concluded that: (1) the TLSTM model has the advantage of preserving information, which can better retain the semantic features related to software defects; (2) compared with the CPDP model trained with traditional metric features, the performance of the model can validly enhance by combining semantic features and metric features.
Hongwei Tao, Lianyou Fu, Qiaoling Cao, Xiaoxu Niu, Songtao Shang, Yang Xian
IET Softw.6
2024 A comparative study of software defect binomial classification prediction models based on machine learning
Hongwei Tao, Xiaoxu Niu, Lianyou Fu, Qiaoling Cao, Songtao Shang, Yang Xian
Softw. Qual. J.7
2018 An Improved Distributed File System Based on GPU Acceleration
abstract
HDFS is a popular distributed file system, widely used in many commercial fields, which can store TB, even PB level data. Fast data reading and writing is the most important problem for HDFS. However, with the volume of data increasing sharply, the traditional HDFS, built on the PC cluster platform, is no longer suitable for fast data reading and writing. GPU is a highly parallel computing unit. Its power of calculation, reading and writing is hundreds of times as fast as CPU. Hence, this paper proposes an improved distributed file system, which uses GPU as an accelerator. Firstly, the improved HDFS uses GPU instead of CPU response data reading and writing requests. Secondly, the improved HDFS uses GPU's cache as a buffer memory for data reading and writing. These two strategies significantly improve the performance of the distributed file system. The experimental results have proved the effectiveness of the improved algorithm.
Songtao Shang, Yong Gan, Huaiguang Wu
ICIS1
2016 A micro-video recommendation system based on big data
abstract
With the development of the Internet and social networking service, the micro-video is becoming more popular, especially for youngers. However, for many users, they spend a lot of time to get their favorite micro-videos from amounts videos on the Internet; for the micro-video producers, they do not know what kinds of viewers like their products. Therefore, this paper proposes a micro-video recommendation system. The recommendation algorithms are the core of this system. Traditional recommendation algorithms include content-based recommendation, collaboration recommendation algorithms, and so on. At the Bid Data times, the challenges what we meet are data scale, performance of computing, and other aspects. Thus, this paper improves the traditional recommendation algorithms, using the popular parallel computing framework to process the Big Data. Slope one recommendation algorithm is a parallel computing algorithm based on MapReduce and Hadoop framework which is a high performance parallel computing platform. The other aspect of this system is data visualization. Only an intuitive, accurate visualization interface, the viewers and producers can find what they need through the micro-video recommendation system.
Songtao Shang, Minyong Shi, Wenqian Shang, Zhiguo Hong
ICIS1
2016 A TV program recommendation system based on big data
abstract
With the development of science and technology, more people especially young teenagers do not want to pay more attention to traditional TV programs. Nowadays the challenge of traditional TV station is how to attract the audience's attention, so as to improve the audience rating of tradition TV programs. This paper proposes a recommendation system, which can improve audience rating. This system mainly contains three modules. Data gathering module is responsible for collecting audience rating data about TV programs on the Internet. Data mining module is responsible for analyzing the audience ration data, and finding interesting programs that the audiences want to watch. This program recommendation system is designed to improve audience rating, and catch the attention of audiences. The system is based on massive user data, and data mining algorithms to analyze the user's interests. Compared with traditional recommendation system, it is capable for Big Data and easier for TV station to recommend TV programs in which audiences are interested, as a way to adds vitality to the television industry.
Minyong Shi, Zhiguo Hong, Songtao Shang, Menghan Yan
ICIS4
2015 Research on public opinion based on Big Data
abstract
Public opinion is the people's response for social phenomena, issues, hot topics, attitudes, emotions, and so on. It reflects the focus problems of the current time of the society. By analyzing the public opinion, we can infer what will happen in the next time, and give better decision support for governments and businesses. Big Data technology is becoming a powerful data analyzing tools for massive data in recent years. Hadoop is an open source massive data processing platform based on Big Data. Mahout is a data mining algorithms' set based on Hadoop, which is designed for processing large-scale and complex data. In most instances, the public opinion information contains many text messages. For many traditional text mining algorithms, it is almost impossible to handle high dimensional data concerns large-volume and complex data sets. Hence, this paper uses Mahout text mining algorithms to process public opinion information.
Songtao Shang, Minyong Shi, Wenqian Shang, Zhiguo Hong
ICIS1