Man Wu

dblp:182/4227 · DBLP profile ↗
← Back
36ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0002-8459-2028ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semi-supervised consistency regularization for micro-video popularity prediction
Yongze Ji, Man Wu, Lilong Liu, Xiong Luo
Multim. Syst.2
2026 Mllm-guided hierarchical semantic ensemble with adaptive visual fusion for zero-shot composed video retrieval
Yiping Meng, Lilong Liu, Man Wu, Zhixiang Ding, Shengsheng Qian
Multim. Syst.3
2026 Minimal discriminative learning for open set recognition
Yiping Meng, Hairui Ren, Lilong Liu, Man Wu, Shengsheng Qian
Multim. Syst.4
2026 Multi-model cooperative denoising for robust cross-modal retrieval with noisy labels
Man Wu, Hengmiao Zhang, Xiong Luo
Multim. Syst.1
2026 Dyna-HemoNet: A lightweight and adaptive framework for accurate blood cell detection
Man Wu
Pattern Recognit. Lett.2
2026 FPSpike: A Fully Parallel and Reconfigurable Architecture for Accelerating Spiking Neural Networks With Structured Sparsity
abstract
This paper introduces FPSpike, a fully-parallel and reconfigurable architecture for accelerating spiking neural networks (SNNs) with structured sparsity. By introducing structured sparse synaptic connections, the neuron computation and weight storage costs are significantly reduced, while the wiring constraints of hardware implementation are alleviated, thus realizing a fully parallel SNN architecture on a single chip. Furthermore, FPSpike achieves high reconfigurability through local connections, allowing the network topology to be flexibly partitioned into multiple distinct hardware cores without redundancy, thereby enabling multi-task spatial parallel processing. In particular, a series of software-hardware co-designs are performed to improve the performance of FPSpike. Experimental results with FPGA-based implementation demonstrate that FPSpike achieves inference throughput improvements of$2.5\times \sim 43.9\times $and$3.2\times \sim 93.8\times $compared to Intel i7-12700K CPU and NVIDIA RTX3060Ti GPU, respectively. FPSpike also achieves performance speedups of$2.0\times \sim 20.2\times $and energy efficiency improvements of$1.3\times \sim 7.7\times $compared to the state-of-the-art FPGA-based SNN accelerators.
Yirong Kan, Man Wu, Yasuhiko Nakashima
IEEE Trans. Circuits Syst. I Regul. Pap.3
2026 DystoHD: an area-efficient hyperdimensional computing system with dynamic hypervector generation for memory-constrained devices
Yirong Kan, Man Wu, Yasuhiko Nakashima
J. Supercomput.3
2025 EATS: Energy-Aware Adaptive Topology Switching for NoCs
Man Wu, Shaswot Shresthamali, Yuan He 0002
ACM Great Lakes Symposium on VLSI1
2025 Improving few-shot object detection via mislabeling mitigation
Pei Yu, Gaocai Wang, Man Wu, Lili Wen
Neurocomputing3
2024 GKF-mQA: Generative Knowledge Fusion Based on Large Language Models for Enhancing Medical Question Answering
Xinbai Li, Man Wu
ADMA (4)2
2024 Probabilistic load forecasting based on quantile regression parallel CNN and BiGRU networks
Gaocai Wang, Xianfei Huang, Shuqiang Huang, Man Wu
Appl. Intell.5
2024 Quantitative evaluation of molecular generation performance of graph-based GANs
Jinli Zhang, Zongli Jiang, Man Wu, Chen Li 0027, Yoshihiro Yamanishi
Softw. Qual. J.4
2023 Training a General Spiking Neural Network with Improved Efficiency and Minimum Latency
Yunpeng Yao, Man Wu, Zheng Chen 0012
ACML2
2023 Mode Collapse Alleviation of Reinforcement Learning-based GANs in Drug Design
abstract
De novo drug design is a challenging task that involves understanding the principles of chemistry, chemical properties, and the rules that govern molecular interactions. Deep learning-based generative models, such as MolGAN, offer a promising approach for generating new molecules with the desired chemical properties from molecular graphs. Such models often combine a discrete generative adversarial network (GAN) and reinforcement learning (RL) to produce highly valid and novel molecules. However, the severe mode collapse problem leads to low performance. This study aims to alleviate and investigate the effect of multiple factors on mode collapse. We conducted experiments on different sampling methods, training epochs, and datasets of various volumes and evaluated the experimental results using performance metrics such as validity, uniqueness, novelty, and diversity. The experimental results demonstrate that noise sampling distributions, training epochs, and training data volumes affect performance. The experimental results provide a direction for mitigating the mode collapse problem for RL-based discrete GANs.
Zongli Jiang, Jinli Zhang, Man Wu, Chen Li 0027, Yoshihiro Yamanishi
BIBM4
2022 Automatic Sleep Staging via Frequency-Wise Spiking Neural Networks
abstract
Identifying sleep stages is a fundamental step for both early detection of disease and neuroscientific exploration. Automatic sleep staging is classic research for replacing the time-consuming gold-standard manual staging procedure. Recently, promising results have been achieved on automatic staging by extracting spatio-temporal features via deep neural networks from electroencephalogram (EEG). However, such methods fail to consistently yield good performance due to a missing piece in data representation: the dynamic fluctuations of neurons on top of EEG features that is non-trivial for automatic sleep staging task. This paper introduces a biomimicry spiking neural network (SNN) to map the aforementioned features serving for automatic sleep staging. Such SNNs are designed as an array of encoders that converts frequency-specific features into long-term spiking coding and then a popular ANN model is used as the staging machine by absorbing the stage-dependent spiking representation. For proof-of-concept, the performance of the proposed framework is demonstrated by introducing multiple sleep datasets. The experimental results showed that the proposed method achieved a competitive stage scoring performance, especially for Wake, N2, and N3, with higher Precision of 0.94, 0.87, and 0.86. Moreover, the ablation studies prove the SNN has the potential for extracting the neuron’s variation features.
Haohui Jia, Ziwei Yang 0002, Pei Gao, Man Wu, Chen Li 0027, Yirong Kan
BIBM4
2022 Temporal Adaptive Aggregation Network for Dynamic Graph Learning
abstract
Dynamic graphs are common in many applications, such as social networks with evolving nodes and edges over time. When handling such dynamics, existing approaches typically suffer from two limitations: (1) they primarily focus on network topology, without taking node class connections and temporal changes into consideration; and (2) the learning objective is primarily constrained by labeled nodes, which often result in over-smoothing and weak-generalization in representation learning, because labeled nodes are limited. In this paper, we propose a temporal adaptive aggregation network (TAAN) for dynamic graph learning. We consider a dynamic graph as a network with changing nodes and edges in temporal order. The temporal adaptive aggregation is to ensure that, for each node, the information aggregation is to consider neighbors from different classes, as well as their temporal order. For each snapshot of the dynamic network, data augmentation and consistency loss are combined to leverage labeled and unlabeled nodes to learn good node embedding. Meanwhile, in order to accommodate temporal changes of graphs, an incremental learning process is used to ensure that learning on each snapshot can inherit weights learned from previous time points, so graph learning can adapt to the dynamic graph environments. Experiments on real-world datasets validate the effectiveness of our approach.
Man Wu, Xingquan Zhu 0001
IEEE Big Data1
2022 SS-LRU: a smart segmented LRU caching
abstract
Many caching policies use machine learning to predict data reuse, but they ignore the impact of incorrect prediction on cache performance, especially for large-size objects. In this paper, we propose a smart segmented LRU (SS-LRU) replacement policy, which adopts a size-aware classifier designed for cache scenarios and considers the cache cost caused by misprediction. Besides, SS-LRU enhances the migration rules of segmented LRU (SLRU) and implements a smart caching with unequal priorities and segment sizes based on prediction and multiple access patterns. We conducted Extensive experiments under the real-world workloads to demonstrate the superiority of our approach over state-of-the-art caching policies.
Chunhua Li 0002, Man Wu, Ke Zhou 0001, Ji Zhang 0010, Yunqing Sun
DAC2
2022 Insulator Fault Diagnosis Based on Improved Transfer Learning from UAV Images
abstract
Insulator fault diagnosis is a daily but key task for the power transmission system. Long-term exposure to complex natural environment will cause different insulator defects. As a common defects, missing-cap defects of insulators will not only affect the structural strength of power insulators, but also bring a certain effect to the stable power transmission. With the rapid development of machine learning, some machine learning-based defect recognition methods have been proposed for fast and high-precision power inspection. However, the handcrafted features could not effectively express the aerial images against complex inspection environment to affect detection performance of the shallow learning algorithms. And the detection precision of deep learning algorithms will be affected by the unbalanced small-scale defects. Therefore, the fast and high-precision power inspection still faces a certain challenge in the smart grid. To address the above issues, fusion with the deep convolutional neural network (DCNN) and transfer learning, a novel fault diagnosis algorithm of power insulators is proposed to provide a fast and accurate power inspection scheme. To remove complex backgrounds, a fast insulator location algorithm based on the lightweight YOLOV4 model is proposed which is served for the following defect recognition. On the basis, to imitate human vision, a defect recognition algorithm is proposed based on multi-feature fusion. Meanwhile, to ensure the feature expression ability of transfer learning on power insulators, a novel optimization strategy of transfer learning is proposed to improve the recognition precision. Experiments show that the proposed method could acquire a good recognition performance than other recognition models.
Lei Yang 0053, Man Wu, Yanhong Liu 0001
SMC3
2022 MuGRA: A Scalable Multi-Grained Reconfigurable Accelerator Powered by Elastic Neural Network
abstract
A massive core computing architecture is developed for accelerating arbitrary calculations in fully parallel with high speed and low cost. The proposed architecture is reconfigurable in fine-grained (arbitrary functions), mid-grained (flexible function feature, accuracy, and number of operands), and coarse-grained (organization of cores). By implementing a large scale of novel bisection neural network (BNN) on hardware, the re-configuration is conducted by partitioning entire BNN into any specific pieces without redundancy. Each piece of BNN retrieves the arbitrary function approximately. By reconfiguring the BNN topology in software, we can easily adjust dimensions of the computing kernel without rewiring, and achieve a wide range of trade-offs between accuracy and efficiency in hardware. In this manner, the multi-grained reconfigurable accelerator (MuGRA) is achieved. Since MuGRA is flexible in all grained levels, various configurations for each validation are demonstrated with rich options of performance-cost matrix. From the FPGA implementation results, compared with other traditional function approximation methods, our method provides fewer parameter storage requirements. The comparison against related works proves that our accelerator effectively reduces the calculation latency with slight accuracy loss.
Yirong Kan, Man Wu, Yasuhiko Nakashima
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Dynamic Hand Gesture Recognition via Electromyographic Signal Based on Convolutional Neural Network
abstract
Dynamic gesture recognition is a typical human-computer interaction method owing to its great potential in practical applications. Currently, most of research work on gesture recognition has mainly focused on vision-based and surface electromyography (sEMG) methods. Compared to vision-based methods, the sequential sEMG signal can directly depict the muscle activity of different gestures which could lead to higher recognition efficiency. However, the effective feature design and selection of sEMG signal is still complicated since muscle fatigue and small electrode displacement will affect the recognition precision of sEMG signals. In this paper, a novel end-to-end dynamic gesture recognition method is developed. The raw sEMG signals are converted into an image form by using the time-frequency transformation method to obtain more comprehensive information for model training and test. And a recognition model based on Convolutional Neural Network (CNN) model is built for high-precision time-frequency image recognition. Experiments indicate that the proposed method could acquire distinguishing features from the pre-prossed images and the overall recognition accuracy on different gestures can reach up to 98.3%.
Shouan Song, Lei Yang 0053, Man Wu, Yanhong Liu 0001, Hongnian Yu
SMC3
2021 Progressive low-rank subspace alignment based on semi-supervised joint domain adaption for personalized emotion recognition
Junhai Luo, Man Wu, Yanping Chen 0009, Yang Yang 0112
Neurocomputing2
2021 DiaNet: An elastic neural network for effectively re-configurable implementation
abstract
An elastic neural network is developed and evolved towards effectively re-configurable hardware in fully parallel on chip. The original prototype of DiaNet is organized as a symmetrical bisection neural network, which is feasible to be partitioned into arbitrary pieces of neural networks (NNs) without redundancy. To prevent the depth explosion in implementing complex tasks (complicated pattern recognition for instance), the evolution of DiaNets is investigated in this work. By using the I/O layer integration technology which enables all neurons in the hidden layer of DiaNet to receive inputs, the number of layers is reduced to 8.8% of DiaNet prototype. In this manner, the DiaNet topology is feasible to implement complex NNs without the risk of depth explosion. Moreover, the skip connection technology is proposed to avoid the gradient vanishing due to deep learning, which is significant to DiaNets especially. Compared with the LeNet5 model as state-of-the-art, the evolved DiaNet topology achieves the parameter reduction of 90.86% for MNIST recognition with the negligible loss of accuracy. To reduce hardware utilization, the sensitivity to the decline of computational precision and bit-width is investigated to suggest the guideline for efficient hardware implementations. Finally, the effectiveness of DiaNet is verified by the proposed re-configurable architecture on FPGA with the power reduction of 10.8% compared to state-of-the-art implementations.
Man Wu, Yirong Kan, Tati Erlina, Yasuhiko Nakashima
Neurocomputing1
2021 OpenWGL: open-world graph learning for unseen class node classification
Man Wu, Shirui Pan, Xingquan Zhu 0001
Knowl. Inf. Syst.1
2021 An Attribute-Based Access Control Policy Retrieval Method Based on Binary Sequence
abstract
With the widespread application of new technologies, fine-grained authorization requires a large number of access control policies. However, the existing policy retrieval method applied to a large-scale policy environment has the problem of low retrieval efficiency. Therefore, this paper proposes an attribute access control policy retrieval method based on the binary sequence. This method uses binary identification and binary code to express access control requests and policies. When the policy is retrieved, the appropriate group is selected through the logical operation of the access control request and the policy binary identification. Within the group, the binary code of the access control request is matched with the binary code of all rules to find suitable rules, thereby reducing the number of matching attribute-value pairs in the rule and improving the efficiency of policy retrieval. Experimental results show that the policy retrieval method proposed in this paper has higher retrieval efficiency.
Ruijie Pan, Gaocai Wang, Man Wu
Secur. Commun. Networks3
2021 Learning Graph Neural Networks with Positive and Unlabeled Nodes
abstract
Graph neural networks (GNNs) are important tools for transductive learning tasks, such as node classification in graphs, due to their expressive power in capturing complex interdependency between nodes. To enable GNN learning, existing works typically assume that labeled nodes, from two or multiple classes, are provided, so that a discriminative classifier can be learned from the labeled data. In reality, this assumption might be too restrictive for applications, as users may only provide labels of interest in a single class for a small number of nodes. In addition, most GNN models only aggregate information from short distances ( e.g. , 1-hop neighbors) in each round, and fail to capture long-distance relationship in graphs. In this article, we propose a novel GNN framework, long-short distance aggregation networks, to overcome these limitations. By generating multiple graphs at different distance levels, based on the adjacency matrix, we develop a long-short distance attention model to model these graphs. The direct neighbors are captured via a short-distance attention mechanism, and neighbors with long distance are captured by a long-distance attention mechanism. Two novel risk estimators are further employed to aggregate long-short-distance networks, for PU learning and the loss is back-propagated for model learning. Experimental results on real-world datasets demonstrate the effectiveness of our algorithm.
Man Wu, Shirui Pan, Lan Du 0002, Xingquan Zhu 0001
ACM Trans. Knowl. Discov. Data1
2020 Improved Cubature Kalman Filter for Target Tracking in Underwater Wireless Sensor Networks
abstract
The underwater sensor network is currently a hot research field in academia and industry with many underwater applications, such as ocean monitoring, seismic monitoring, environment monitoring, and seabed exploration. Underwater target tracking is a critical component of ocean development. This paper studies the underwater target tracking problem of the wireless sensor network. The core technology of the target tracking algorithm is the filtering algorithm, which identifies the accuracy of the target tracking system. Nonlinear filtering is a hot issue in target tracking because feasible projects are mostly non-linear systems. The linearization method used in traditional Kalman filtering has serious shortcomings. Therefore, this paper presents the improved cubature Kalman filtering (ICKF) algorithm for underwater target tracking. There is uncertainty in the target movement, an adaptive forgetting factor is given into the cubature Kalman filtering algorithm to directly modify the error covariance to reduce the impact of uncertainties. Then, interactive multi-model technology is introduced to establish the IMMICKF algorithm with multiple states. Compared with other filtering algorithms, the new algorithm can effectively deal with non-linear target tracking problems and obtain better estimation accuracy. The numerical simulation is given to demonstrate the effectiveness of the IMMICKF algorithm.
Junhai Luo, Yanping Chen 0009, Man Wu, Yang Yang 0112
FUSION4
2020 OpenWGL: Open-World Graph Learning
abstract
In traditional graph learning tasks, such as node classification, learning is carried out in a closed-world setting where the number of classes and their training samples are provided to help train models, and the learning goal is to correctly classify unlabeled nodes into classes already known. In reality, due to limited labeling capability and dynamic evolving of networks, some nodes in the networks may not belong to any existing/seen classes, and therefore cannot be correctly classified by closed-world learning algorithms. In this paper, we propose a new open-world graph learning paradigm, where the learning goal is to not only classify nodes belonging to seen classes into correct groups, but also classify nodes not belonging to existing classes to an unseen class. The essential challenge of the open-world graph learning is that (1) unseen class has no labeled samples, and may exist in an arbitrary form different from existing seen classes; and (2) both graph feature learning and prediction should differentiate whether a node may belong to an existing/seen class or an unseen class. To tackle the challenges, we propose an uncertain node representation learning approach, using constrained variational graph autoencoder networks, where the label loss and class uncertainty loss constraints are used to ensure that the node representation learning are sensitive to unseen class. As a result, node embedding features are denoted by distributions, instead of deterministic feature vectors. By using a sampling process to generate multiple versions of feature vectors, we are able to test the certainty of a node belonging to seen classes, and automatically determine a threshold to reject nodes not belonging to seen classes as unseen class nodes. Experiments on real-world networks demonstrate the algorithm performance, comparing to baselines. Case studies and ablation analysis also show the rationale of our design for open-world graph learning.
Man Wu, Shirui Pan, Xingquan Zhu 0001
ICDM1
2020 Unsupervised Domain Adaptive Graph Convolutional Networks
abstract
Graph convolutional networks (GCNs) have achieved impressive success in many graph related analytics tasks. However, most GCNs only work in a single domain (graph) incapable of transferring knowledge from/to other domains (graphs), due to the challenges in both graph representation learning and domain adaptation over graph structures. In this paper, we present a novel approach, unsupervised domain adaptive graph convolutional networks (UDA-GCN), for domain adaptation learning for graphs. To enable effective graph representation learning, we first develop a dual graph convolutional network component, which jointly exploits local and global consistency for feature aggregation. An attention mechanism is further used to produce a unified representation for each node in different graphs. To facilitate knowledge transfer between graphs, we propose a domain adaptive learning module to optimize three different loss functions, namely source classifier loss, domain classifier loss, and target classifier loss as a whole, thus our model can differentiate class labels in the source domain, samples from different domains, the class labels from the target domain, respectively. Experimental results on real-world datasets in the node classification task validate the performance of our method, compared to state-of-the-art graph neural network algorithms.
Man Wu, Shirui Pan, Chuan Zhou 0001, Xiaojun Chang, Xingquan Zhu 0001
WWW1
2020 Using Live Video Streaming in Online Tutoring: Exploring Factors Affecting Social Interaction
abstract
The growth of live video streaming (LVS) technology provides new possibilities for online tutoring in that it accommodates a massive number of learners simultaneously. Questions still exist, however, about the extent to which new technology can support interactions between an instructor and a vast number of learners, as well as which factors would influence learners’ interactions with the instructor and peer learners. This study explored these questions by conducting a survey involving 189 senior high school students participating in online LVS tutoring. The results indicated that learner–instructor interaction dominated social interaction in the online tutoring environment with the current system design. This design may also contribute to the development of the perceived presence of peer learners with few direct information exchanges among peers. Social Connectedness and perceived enjoyment positively influenced learner–instructor interaction, whereas social fears and the social presence of the instructor negatively influenced learner–learner interaction.
Man Wu, Qin Gao
Int. J. Hum. Comput. Interact.1
2019 Long-short Distance Aggregation Networks for Positive Unlabeled Graph Learning
abstract
Graph neural nets are emerging tools to represent network nodes for classification. However, existing approaches typically suffer from two limitations: (1) they only aggregate information from short distance (e.g., 1-hop neighbors) each round and fail to capturelong distance relationship in graphs; (2) they require users to label data from several classes to facilitate the learning of discriminative models; whereas in reality, users may only provide labels of a small number of nodes in a single class. To overcome these limitations, this paper presents a novel long-short distance aggregation networks (\textttLSDAN ) for positive unlabeled (PU) graph learning. Our theme is to generate multiple graphs at different distances based on the adjacency matrix, and further develop a long-short distance attention model for these graphs. The short-distance attention mechanism is used to capture the importance of neighbor nodes to a target node. The long-distance attention mechanism is used to capture the propagation of information within a localized area of each node and help model weights of different graphs for node representation learning. A non-negative risk estimator is further employed, to aggregate long- short-distance networks, for PU learning using back-propagated loss modeling. Experiments on real-world datasets validate the effectiveness of our approach.
Man Wu, Shirui Pan, Lan Du 0002, Ivor W. Tsang, Xingquan Zhu 0001, Bo Du 0001
CIKM1
2019 An Optimal Bit Allocation Scheme for Cooperative Spectrum Sensing in Cognitive Radio Networks
Junhai Luo, Xiaoting He 0002, Man Wu, Yanping Chen 0009, Yang Yang 0005
FUSION3
2019 Domain-Adversarial Graph Neural Networks for Text Classification
abstract
Text classification, in cross-domain setting, is a challenging task. On the one hand, data from other domains are often useful to improve the learning on the target domain; on the other hand, domain variance and hierarchical structure of documents from words, key phrases, sentences, paragraphs, etc. make it difficult to align domains for effective learning. To date, existing cross-domain text classification methods mainly strive to minimize feature distribution differences between domains, and they typically suffer from three major limitations - (1) difficult to capture semantics in non-consecutive phrases and long-distance word dependency because of treating texts as word sequences, (2) neglect of hierarchical coarse-grained structures of document for feature learning, and (3) narrow focus of the domains at instance levels, without using domains as supervisions to improve text classification. This paper proposes an end-to-end, domain-adversarial graph neural networks (DAGNN), for cross-domain text classification. Our motivation is to model documents as graphs and use a domain-adversarial training principle to lean features from each graph (as well as learning the separation of domains) for effective text classification. At the instance level, DAGNN uses a graph to model each document, so that it can capture non-consecutive and long-distance semantics. At the feature level, DAGNN uses graphs from different domains to jointly train hierarchical graph neural networks in order to learn good features. At the learning level, DAGNN proposes a domain-adversarial principle such that the learned features not only optimally classify documents but also separates domains. Experiments on benchmark datasets demonstrate the effectiveness of our method in cross-domain classification tasks.
Man Wu, Shirui Pan, Xingquan Zhu 0001, Chuan Zhou 0001, Lei Pan 0002
ICDM1
2018 Revisting the Impact of Regression Models for Predicting the Number of Defects
abstract
Predicting the number of faults in software modules can be more helpful instead of predicting the modules being faulty or non-faulty.Chen et al. (SEKE 397-402, 2015) and Rathore et al. (Soft Computing 21: 7417-7434, 2017) empirically investigate the feasibility of some regression algorithms for predicting the number of defects.The experimental results showed that the decision tree regression algorithm performed best in terms of average absolute error (AAE), average relative error (ARE) and root mean square error (RMSE).However, they did not consider the imbalanced data distribution problem in defect datasets and employed improper performance measures for evaluating the regression models to evaluate the performance of models for predicting the number of defects.Hence, we revisit the impact of different regression algorithms for predicting the number of defects using Fault-Percentile-Average (FPA) as the performance measure.The experiments on 31 datasets from PROMISE repository show that the prediction performance of models for predicting the number of defects built by different regression algorithms are various, and the gradient boosting regression algorithm and the Bayesian ridge regression algorithm can achieve better performance. Keywords-predicting the number of defects; regression algorithm; data imbalance; Fault-Percentile-Average;
Man Wu, Sizhe Ye, Chunhua Li 0002, Ziyi Ma, Zhongwang Fu
SEKE1
2018 Cross-company defect prediction via semi-supervised clustering-based data filtering and MSTrA-based transfer learning
Xiao Yu 0008, Man Wu, Yiheng Jian, Kwabena Ebo Bennin, Mandi Fu, Chuanxiang Ma
Soft Comput.2
2017 A Reinforced Hungarian Algorithm for Task Allocation in Global Software Development
abstract
The allocation of software development tasks is a critical management activity in distributed development projects.One of the most important problem is to find the lowest-cost way to assign tasks in global software development, which can be solved by Hungarian algorithm.However, the original Hungarian algorithm only assume that a task can only be solved by one development site.The assumption is not agreed with the actual case where a software development task is usually be solved through a collaboration among several sites.To address such an issue, this paper proposes a reinforced Hungarian algorithm (RHA) for task assignment in global software development.RHA consists of three major stages.First, RHA transforms a n×m cost matrix into two n×n cost matrix by adding (2n-m) virtual development sites.Second, RHA performs the original Hungarian algorithm on the two n×n cost matrix to get the optimal assignment results.Finally, RHA removes the (2n-m) virtual development sites and gets the final optimal assignment result for m tasks.Simulation results indicate that RHA is a viable approach for the task assignment problem in global software development.1
Xiao Yu 0008, Man Wu, Xiangyang Jia
SEKE2
2017 Combing Data Filter and Data Sampling for Cross-Company Defect Prediction: An Empricial Study
abstract
Cross-company defect prediction (CCDP) is a practical way that trains a prediction model by exploiting one or multiple projects of a source company and then applies the model to target company.Unfortunately, larger irrelevant crosscompany (CC) data usually makes it difficult to build a prediction model with high performance.On the other hand, the CC data has the highly imbalanced nature between the defectiveprone and non-defective classes, which will degrade the performance of CCDP.To address such issues, this paper proposes an approach, in which data sampling is combined with data filter, to overcome these problems.Data sampling seeks a more balanced dataset through the addition or removal of instances, while data filter is a process of filtering out the irrelevant CC data so that the performance of CCDP models can be improved.We employ two data filtering methods called NN filter and DBSCAN filter combined with SMOTE (Synthetic Minority Oversampling Technique) and RUS (Random Under-Sampling).Eight different approaches would be produced when combing these four techniques: 1-NN filter performed prior to RUS; 2-NN filter performed after RUS; 3-NN filter performed prior to SMOTE; 4-NN filter performed after SMOTE; 5-DBSCAN filter performed prior to RUS; 6-DBSCAN filter performed after RUS; 7-DBSCAN filter performed prior to SMOTE; 8-DBSCAN filter performed after SMOTE.The empirical study was carried out on 15 publicly available project datasets.The experimental results demonstrate that NN filter performed prior to RUS (Approach 1) performs better than the other seven approaches.
Xiao Yu 0008, Man Wu, Mandi Fu
SEKE2